Micro bearing production detection data sharing method and system based on deep learning
Patent Information
- Application Number
- CN202610779967.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-02
AI Technical Summary
[0004]本发明解决的技术问题是:现有技术难以在微型轴承生产检测中,将高频时序信号与二维空间图像等异构多模态数据在统一的物理几何基准下进行精确的时空拓扑对齐;同时,现有技术难以在复杂的实际生产环境中,智能剥离由工艺变化或型号差异引起的“正常形态波动”与“真实的物理缺陷”;进而导致在工业物联网云边协同架构中进行数据共享时,难以在极低通信带宽占用的前提下,高保真地传输并重构出蕴含完整物理形貌和底层特征的多模态原始检测数据
[0014] The beneficial effects of this invention are as follows: This invention innovatively constructs a data sharing architecture that integrates "counterfactual normal mirror" and "defect deviation field," completely breaking through the bandwidth bottleneck of cross-node transmission of massive multimodal industrial inspection data. By extracting and distributing only highly compressed deviation field data, production context parameters, and mirror identifiers at the sending end, the receiving end can use a conditional variational generation network to deduce a baseline mirror under specific working conditions and superimpose it with the deviation field, thus restoring the original multimodal data with high fidelity and lossless performance with extremely low communication overhead. Simultaneously, a unified geometric coordinate system for the surface of a micro-bearing is constructed. By extracting the rotation angle phase, high-frequency time-domain signals are accurately projected into the spatial angle domain. Furthermore, a differential homeomorphic spatial transformation operator with local topology preservation constraints is introduced to ensure that heterogeneous data does not undergo local folding or tearing during continuous mapping, achieving pixel-level precise alignment of multi-source data on a real physical surface. Furthermore, this invention incorporates production context parameters as conditional constraints into the latent space for counterfactual deduction, intelligently separating "real physical defects" from "normal process morphology differences," significantly reducing the false alarm rate under complex operating conditions; and binds the deviation vector to the bearing geometry through deep spatial indexing, enabling the abstract defect features to be directly and structurally mapped back to their real physical locations, providing a high-value data foundation for refined quality traceability and visual diagnosis throughout the entire plant.
Smart Images

Figure CN122332616B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for sharing production and testing data of micro-bearings based on deep learning. Background Technology
[0002] In recent years, with the rapid development of the Industrial Internet and intelligent manufacturing, miniature bearings, as high-precision core components, face extremely high requirements for quality inspection during their production process. To comprehensively assess the surface defects and internal condition of miniature bearings, modern production lines typically employ multi-source heterogeneous sensors for joint inspection. This results in inspection data that is multimodal, high-dimensional, and massive. In cross-plant collaborative manufacturing or cloud-edge-device collaborative computing scenarios, the challenge lies in how to efficiently and accurately share this massive amount of inspection data to support remote verification and model iteration.
[0003] Currently, Chinese invention patent application CN118312874A discloses a deep learning-based bearing fault detection method. This method constructs a deep learning model, including multi-scale feature extraction (MFEM) and cross-domain collaborative attention (CDCAN). The model is trained to obtain optimal weights and biases, resulting in a deep learning model capable of achieving high accuracy in bearing fault diagnosis. The trained deep learning model is then used in rolling machinery fault diagnosis experiments. Multi-scale features across the time and frequency domains are extracted from the original vibration signals and then fused with the CDCAN mechanism. This helps the model provide more reliable diagnoses than state-of-the-art methods. Experiments were conducted on bearing and gearbox datasets to evaluate fault diagnosis performance. Extensive experimental results and comprehensive analysis demonstrate the advantages of the proposed CDCAN network in terms of diagnostic accuracy and adaptability. However, related technologies lack production context awareness and counterfactual reasoning capabilities, and lack a unified alignment mechanism for multimodal data in physical geometric space, making it impossible to achieve high-fidelity, low-bandwidth sharing of multimodal raw data. Summary of the Invention
[0004] The technical problem solved by this invention is that existing technologies struggle to accurately align high-frequency time-series signals with heterogeneous multimodal data such as two-dimensional spatial images under a unified physical geometric benchmark in the production and testing of miniature bearings. Furthermore, existing technologies struggle to intelligently separate "normal morphological fluctuations" from "real physical defects" caused by process variations or model differences in complex actual production environments. Consequently, when sharing data in an industrial IoT cloud-edge collaborative architecture, it is difficult to transmit and reconstruct multimodal raw test data containing complete physical morphology and underlying features with high fidelity under extremely low communication bandwidth requirements.
[0005] To address the aforementioned technical problems, the present invention provides the following technical solution: Firstly, a method for sharing micro-bearing production and testing data based on deep learning, comprising the following steps: Step S1: Obtain the original detection data and production context parameters of the miniature bearing, and use a multimodal joint semantic encoder to perform feature fusion and dimensionality reduction mapping to output the latent vector of the sample to be tested. Step S2: Using production context parameters as constraints, a reconstruction operation is performed in a unified latent space through a conditional variational generator network to obtain a counterfactual normal image. Based on a unified time reference and a unified spatial reference, the original detection data and the counterfactual normal image are aligned with each other to obtain aligned original detection data. Step S3: Calculate the deviation vector of the aligned original detection data relative to the counterfactual normal mirror image, and organize the deviation vector according to the spatial position index to obtain the defect deviation field; Step S4: Encapsulate the index identifiers and production context parameters of the defect deviation field, counterfactual normal image, and obtain shared data, and perform the distribution operation.
[0006] As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S1 specifically includes: Step S11: Divide the raw detection data according to the preset modal type, and perform synchronous alignment processing under a unified time reference and a unified spatial reference to obtain the normalized raw detection data; The original detection data includes bearing surface image data, bearing rotation time-domain signal, high-frequency acoustic feature data, bearing geometric deformation values, and geometric structural parameters of the miniature bearing; The geometric parameters of the miniature bearing include the inner ring size, outer ring size, raceway structure parameters, and axial structure parameters. Based on the normalized original test data, a unified micro-bearing surface geometric coordinate system is constructed according to the geometric structural parameters of the micro-bearing. The bearing surface image data and bearing geometric deformation values in the normalized original test data are mapped to the unified micro-bearing surface geometric coordinate system according to their physical spatial positions. The bearing rotation time domain signal and high-frequency acoustic feature data in the normalized original test data are extracted according to a unified time reference to extract the corresponding rotation angle phase, and are jointly mapped to the unified micro-bearing surface geometric coordinate system for unified representation by means of angular position spatial projection. Step S12: Input the normalized original detection data into the branch structure corresponding to the multimodal joint semantic encoder, perform feature fusion processing on the detection data of different modalities, and perform dimensionality reduction mapping in the unified latent space to obtain the initial latent vector. Step S13: Map the production context parameters to the unified latent space and concatenate them with the initial latent vector to obtain the context-related latent vector; The production context parameters include bearing model identifier, current production process node, detection modal attributes, and real-time process environment variables; Step S14: Perform consistency constraint processing on the context-related latent vector, calculate the differences between the feature components in the initial latent vector, and update the context-related latent vector according to the constraint calculation results to obtain the latent vector of the sample to be tested. As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S12 specifically includes: The normalized raw detection data are input into the corresponding branch structure of the multimodal joint semantic encoder according to the preset modality type; Specifically, the bearing surface image data is input into the residual convolutional network branch, the bearing rotation time-domain signal is input into the one-dimensional convolutional network branch, the high-frequency acoustic feature data is input into the spectral convolutional branch, and the bearing rotation time-domain signal is input into the fully connected mapping branch, which respectively obtain the initial feature vector of each mode. A unified dimension mapping is performed on the initial feature vectors to map them to a unified latent space, resulting in standardized feature vectors. The calculation formula for these standardized feature vectors is as follows: ; in, To standardize the feature vector, For the first Mapping matrix corresponding to each mode For the first The initial feature vectors corresponding to each mode. It is the bias vector; The standardized feature vectors are concatenated, and the standardized feature vectors corresponding to each modality are combined in a preset order to obtain the fused feature vector; The fusion feature vector is reduced in dimension and mapped to a unified latent space to obtain the initial latent vector. The initial latent vector is composed of feature components corresponding to different modes.
[0007] As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S2 specifically includes: Step S21: The latent vector of the sample to be tested, the production context parameters, and the preset defect-free baseline label are taken as joint inputs and fed into the conditional variational generation network. Reconstruction operations are performed within the unified latent space manifold to obtain the initial counterfactual latent vector. The processing logic is as follows: Let the latent vector of the sample to be tested be denoted as... The vector after mapping the production context parameters is denoted as The preset defect-free status label is denoted as Construct joint input vector It is represented as: ; By mapping the joint input vector to a probability distribution, a set of latent variable distribution parameters is obtained. Then, sampling operations are performed based on this set to obtain the sampled latent variables, calculated using the following formula: ; in, To sample latent variables, It is the mean vector. The standard deviation vector, Let be a random variable that follows a standard normal distribution; The sampled latent variables are input into the latent space to reconstruct the mapping function, thus obtaining the initial counterfactual latent vector; Step S22: Perform latent space constraint calculation on the initial counterfactual latent vector, introduce the production context parameters as constraints into the unified latent space manifold, and map and adjust the initial counterfactual latent vector in combination with the manifold distribution corresponding to the preset standard state label to obtain the constrained counterfactual latent vector. Step S23: Perform reverse mapping processing on the constrained counterfactual latent vectors according to the preset modality type, convert the latent space representation into multimodal detection data form, and obtain the counterfactual normal mirror; Step S24: Perform coordinate mapping calculations on the original detection data and the counterfactual normal mirror according to a unified time reference and a unified spatial reference, and establish a positional correspondence between the original detection data and the counterfactual normal mirror in a unified micro-bearing surface geometric coordinate system. Step S25: Under the unified geometric coordinate system of the micro-bearing surface, a differential homeomorphic space transformation operator is constructed based on the position correspondence and global mutual information value. A local topology preservation constraint term is added to the differential homeomorphic space transformation operator, and a continuous space mapping operation is performed on the original detection data to obtain the aligned original detection data.
[0008] As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S23 specifically includes: The constrained counterfactual latent vectors are input into the latent space projection layer corresponding to each mode, and the latent space sub-vectors corresponding to each mode are extracted from the constrained counterfactual latent vectors through linear projection operation; The latent space sub-vectors of each mode are subjected to inverse mapping processing. The latent space sub-vectors of each mode are input into the corresponding inverse mapping function to obtain the reconstructed detection data of each mode. The calculation formula is as follows: ; in, For the first Reconstruction detection data corresponding to each modality For the first The inverse mapping function corresponding to each mode The first one extracted from the constrained counterfactual latent vector. The latent space sub-vectors corresponding to each mode; The reconstructed detection data of each modality is organized according to the preset modality type, and the reconstructed detection data of each modality is arranged according to the data structure consistent with the normalized original detection data to obtain a multimodal reconstruction data set; The multimodal reconstructed data set is combined according to a unified time reference and a unified spatial reference to generate a counterfactual normal image with the same structure as the original detection data.
[0009] As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S25 specifically includes: Under the unified geometric coordinate system of the micro-bearing surface, the positional correspondence between the original detection data and the counterfactual normal mirror is obtained, and the global mutual information value between the original detection data and the counterfactual normal mirror is calculated to construct the corresponding position set; Based on positional correspondence and global mutual information, a differential homeomorphic spatial transformation operator is constructed in a unified geometric coordinate system of the micro-bearing surface. Local topology preservation constraints are added to the transformation objective function of the differential homeomorphic spatial transformation operator to perform continuous mapping calculations on the original spatial coordinates of the original detection data. The calculation formula is as follows: ; in, These are the mapped spatial coordinates. For differential homeomorphic space transformation operators, The original spatial coordinates of the original test data in the unified geometric coordinate system of the micro-bearing surface; The operator for transforming a differential homeomorphic space is subjected to continuous differentiability constraints, and its Jacobian matrix is subjected to non-singularity constraints based on local topological preservation constraints. Its expression is as follows: ; in, For the differential homeomorphic space transformation operator in the original space coordinates Jacobian matrix at the location, To find the determinant of the Jacobian matrix; Based on the mapped spatial coordinates, values are calculated in the original detection data according to the spatial index relationship. The values of the original detection data at the original spatial coordinates are then mapped back to the mapped spatial coordinates to obtain the aligned original detection data.
[0010] As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S3 specifically includes: Step S31: Under the unified geometric coordinate system of the miniature bearing surface, the aligned original detection data and the counterfactual normal mirror image are compared at corresponding positions to calculate the deviation vector, forming an initial deviation set. The calculation formula is as follows: ; in, For the set of deviation vectors, The original detection data after alignment. To be a counterfactual mirror image, This is the difference evaluation function corresponding to the preset modality type; When the preset mode type is bearing surface image data or bearing geometric deformation value, the difference evaluation function is the absolute difference function; When the preset mode type is a bearing rotation time-domain signal or high-frequency acoustic feature data, the difference evaluation function is the feature divergence function; Step S32 involves associating the set of deviation vectors according to a unified geometric coordinate system for the surface of the micro-bearing, binding each deviation vector to its corresponding mapped spatial coordinates to form a set of corresponding deviation vectors and mapped spatial coordinates. The processing logic is as follows: The spatial coordinate set of the aligned original detection data in the unified micro-bearing surface geometric coordinate system is obtained. The spatial coordinates are sorted according to the coordinate arrangement rules in the unified micro-bearing surface geometric coordinate system, and the sorted spatial coordinates are assigned corresponding index numbers to obtain the spatial position index. The spatial coordinate index is obtained by indexing each deviation vector in the deviation vector set according to its positional correspondence. According to the spatial coordinate index, each deviation vector is associated with the mapped spatial coordinates to establish a correspondence between the deviation vectors and the spatial coordinates. This correspondence is then arranged according to the spatial coordinate index order to obtain the set of correspondences between the deviation vectors and the mapped spatial coordinates. The calculation formula is as follows: ; in, This is the set of corresponding offset vectors and mapped spatial coordinates. To unify the geometric coordinate system of miniature bearing surfaces, the first A mapped spatial coordinate, The deviation vector in the deviation vector set corresponds to the mapped spatial coordinates. This refers to the spatial coordinate index number; Step S33: Under the unified geometric coordinate system of the micro bearing surface, the deviation vector set is divided into regions according to the preset spatial division method, and the deviation vectors are assigned to the corresponding regions according to the mapped spatial coordinates to obtain the regionalized deviation set. Generate the corresponding spatial location index based on the order of the mapped spatial coordinates; Step S34: Organize the regionalized deviation set according to spatial location index, arrange the deviation vectors in each region in a preset order, and construct the defect deviation field. The processing logic is as follows: Traverse the set of corresponding offset vectors and mapped spatial coordinates, extract all spatial coordinates, and sort the mapped spatial coordinates according to the spatial position index under the unified micro-bearing surface geometric coordinate system to obtain an ordered spatial coordinate sequence. The corresponding deviation vectors are synchronously rearranged based on the ordered spatial coordinate sequence to obtain an ordered deviation vector sequence. The ordered deviation vector sequence is combined with the corresponding spatial coordinate sequence and arranged according to the spatial position index to obtain an ordered correspondence sequence; The ordered correspondence sequence is organized according to a unified geometric coordinate system of the micro-bearing surface, and the correspondence between the deviation vector and the mapped spatial coordinates is uniformly represented to obtain the defect deviation field.
[0011] As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S4 specifically includes: Step S41: Analyze the defect deviation field, extract the correspondence between the deviation vector and the mapped spatial coordinates, and index the correspondence according to the spatial position index under the unified micro bearing surface geometric coordinate system to obtain the deviation field index data. Step S42: Extract the index identifier of the counterfactual normal image, and perform identifier mapping processing on the counterfactual normal image according to the unified time reference and unified spatial reference, and associate the index identifier of the counterfactual normal image with the spatial coordinate index in the defect deviation field. Step S43: Encode the production context parameters according to the preset parameter structure, and associate and calibrate the production context parameters according to the spatial location index, and associate the production context parameters with the spatial coordinate index in the defect deviation field. Step S44: Combine the off-field index data, the index identifier of the counterfactual normal mirror, and the production context parameters according to a unified data format, and organize the combined results in a structured manner to obtain shared data. Step S45: Sort the shared data according to the spatial location index, group it according to the production context parameters, and perform the distribution operation according to the preset distribution rules. The preset distribution rules include preset data splitting rules, preset field mapping rules, preset data frame generation rules, and preset sending order rules.
[0012] As a preferred embodiment of the deep learning-based micro-bearing production and inspection data sharing method described in this invention, step S45 specifically includes: According to the preset data splitting rules, the shared data is parsed and split according to the preset field structure. The index identifiers of defect deviation fields and counterfactual normal mirrors and production context parameters are extracted from the shared data. According to the preset field mapping rules, the index identifiers of the defect deviation field and the counterfactual normal image and the production context parameters are mapped according to the preset communication protocol and written to the corresponding positions in the order of the fields to generate transmission data frames. According to the preset data frame generation rules, the transmitted data frames are arranged in a preset order, and each transmitted data frame is numbered based on the sequence index to form a data transmission sequence; According to the preset sending order rules, the data sending sequence is written into the network transmission interface, and each data frame is sent frame by frame in the order of the sequence index, while the corresponding sequence index is recorded. At the receiving end, the received transmission data frames are rearranged according to the sequence index, and the index identifiers of the defect deviation field and the counterfactual normal image and the production context parameters are parsed according to the preset field structure. The index identifier of the counterfactual normal image is input into the conditional variational generation network to extract the counterfactual normal image, and the extracted counterfactual normal image is superimposed with the defect deviation field to restore the corresponding data content.
[0013] Secondly, a deep learning-based data sharing system for the production and testing of miniature bearings includes a data encoding module, a mirror alignment module, a deviation construction module, and a data distribution module. The data encoding module is used to acquire the original detection data and production context parameters of the miniature bearing, and to perform feature fusion and dimensionality reduction mapping using a multimodal joint semantic encoder to output the latent vector of the sample to be tested. The mirror alignment module is used to perform reconstruction operations in a unified latent space with production context parameters as constraints, through a conditional variational generation network, to obtain a counterfactual normal mirror. Based on a unified time reference and a unified spatial reference, the original detection data and the counterfactual normal mirror are subjected to co-position topology alignment processing to obtain aligned original detection data. The deviation construction module is used to calculate the deviation vector of the aligned original detection data relative to the counterfactual normal mirror, and organize the deviation vector according to the spatial position index to obtain the defect deviation field; The data distribution module is used to encapsulate the index identifiers and production context parameters of the defect deviation field and the counterfactual normal image to obtain shared data and perform distribution operations.
[0014] The beneficial effects of this invention are as follows: This invention innovatively constructs a data sharing architecture that integrates "counterfactual normal mirror" and "defect deviation field," completely breaking through the bandwidth bottleneck of cross-node transmission of massive multimodal industrial inspection data. By extracting and distributing only highly compressed deviation field data, production context parameters, and mirror identifiers at the sending end, the receiving end can use a conditional variational generation network to deduce a baseline mirror under specific working conditions and superimpose it with the deviation field, thus restoring the original multimodal data with high fidelity and lossless performance with extremely low communication overhead. Simultaneously, a unified geometric coordinate system for the surface of a micro-bearing is constructed. By extracting the rotation angle phase, high-frequency time-domain signals are accurately projected into the spatial angle domain. Furthermore, a differential homeomorphic spatial transformation operator with local topology preservation constraints is introduced to ensure that heterogeneous data does not undergo local folding or tearing during continuous mapping, achieving pixel-level precise alignment of multi-source data on a real physical surface. Furthermore, this invention incorporates production context parameters as conditional constraints into the latent space for counterfactual deduction, intelligently separating "real physical defects" from "normal process morphology differences," significantly reducing the false alarm rate under complex operating conditions; and binds the deviation vector to the bearing geometry through deep spatial indexing, enabling the abstract defect features to be directly and structurally mapped back to their real physical locations, providing a high-value data foundation for refined quality traceability and visual diagnosis throughout the entire plant. Attached Figure Description
[0015] Figure 1 A flowchart illustrating the steps of a deep learning-based method for sharing production and testing data of micro-bearings, as provided in one embodiment of the present invention. Figure 2 This is a schematic diagram of the basic process of a deep learning-based micro-bearing production and testing data sharing system provided in one embodiment of the present invention. Detailed Implementation
[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0017] Example 1, referring to Figure 1 This paper provides a method for sharing production and inspection data of micro-bearings based on deep learning, including the following steps: Step S1: Obtain the original detection data and production context parameters of the miniature bearing, and use a multimodal joint semantic encoder to perform feature fusion and dimensionality reduction mapping to output the latent vector of the sample to be tested. Step S2: Using production context parameters as constraints, a reconstruction operation is performed in a unified latent space through a conditional variational generator network to obtain a counterfactual normal image. Based on a unified time reference and a unified spatial reference, the original detection data and the counterfactual normal image are aligned with each other to obtain aligned original detection data. Step S3: Calculate the deviation vector of the aligned original detection data relative to the counterfactual normal mirror image, and organize the deviation vector according to the spatial position index to obtain the defect deviation field; Step S4: Encapsulate the index identifiers and production context parameters of the defect deviation field, counterfactual normal image, and obtain shared data, and perform the distribution operation.
[0018] In specific implementation, step S1 includes: Step S11: Divide the raw detection data according to the preset modal type, and perform synchronous alignment processing under a unified time reference and a unified spatial reference to obtain the normalized raw detection data; The raw test data includes bearing surface image data, bearing rotation time-domain signal, high-frequency acoustic feature data, bearing geometric deformation values, and geometric structural parameters of miniature bearings; The geometric parameters of miniature bearings include the inner ring size, outer ring size, raceway structure parameters, and axial structure parameters. Based on the normalized original test data, a unified micro-bearing surface geometric coordinate system is constructed according to the geometric structural parameters of the micro-bearing. The bearing surface image data and bearing geometric deformation values in the normalized original test data are mapped to the unified micro-bearing surface geometric coordinate system according to their physical spatial positions. The bearing rotation time domain signal and high-frequency acoustic feature data in the normalized original test data are extracted according to a unified time reference to extract the corresponding rotation angle phase, and are jointly mapped to the unified micro-bearing surface geometric coordinate system for unified representation by means of angular position spatial projection. Step S12: Input the normalized original detection data into the branch structure corresponding to the multimodal joint semantic encoder, perform feature fusion processing on the detection data of different modalities, and perform dimensionality reduction mapping in the unified latent space to obtain the initial latent vector. Step S13: Map the production context parameters to the unified latent space and concatenate them with the initial latent vector to obtain the context-related latent vector; Production context parameters include bearing model identifier, current production process node, detection modal attributes, and real-time process environment variables; Step S14: Perform consistency constraint processing on the context-related latent vector, perform constraint calculation on the differences between each feature component in the initial latent vector, and update the context-related latent vector according to the constraint calculation results to obtain the latent vector of the sample to be tested.
[0019] Specifically, the preset modal type refers to the natural classification of the original detection data in terms of physical form and acquisition method. This paper explicitly divides it into five categories: bearing surface image data, bearing rotation time-domain signal, high-frequency acoustic feature data, bearing geometric deformation values, and geometric structural parameters of miniature bearings. The reason for dividing it by modal type is that these types of data are completely different in terms of sampling principle, physical dimension, and semantic meaning. For example, image data naturally has a two-dimensional spatial structure, while time-domain signals are one-dimensional time series, and geometric deformation values are scalar fields distributed in physical space. If they are directly mixed together and input into the model, it is impossible to align the physical position and it is difficult to design a unified feature extraction structure.
[0020] After modal segmentation, the data from these different modes are synchronized in time and space to enable subsequent comparison and fusion based on physical location. A unified time reference is not simply clock alignment; rather, it addresses the time-domain signals and acoustic characteristics closely related to bearing rotation by extracting the rotation angle phase, converting the signals originally sampled in absolute time into a sequence organized by rotation angle, thereby eliminating misalignment issues caused by speed fluctuations. A unified spatial reference, on the other hand, constructs a standardized unified geometric coordinate system for the micro-bearing surface using the geometric parameters of the micro-bearing, including inner ring dimensions, outer ring dimensions, raceway structural parameters, and axial structural parameters. This coordinate system typically uses polar or cylindrical coordinates, with the bearing center as the origin, and uses radial radius, circumferential angle, and axial height to uniquely describe each physical location on the bearing surface. The physical spatial location refers to the actual spatial point on the bearing surface corresponding to the measured data, such as the bearing surface area corresponding to a pixel in an image, or the raceway position where a certain coordinate in the geometric deformation values is located.
[0021] The system maps different types of data to a unified coordinate system based on their physical spatial location: bearing surface image data and geometric deformation values inherently contain spatial coordinate information and can be directly mapped through coordinate transformation; while for bearing rotation time-domain signals and high-frequency acoustic feature data, the extracted rotation angle phase is used, with the angle corresponding to each sampling point as a coordinate component. Combined with the fixed installation position of the sensor on the bearing surface, the signal values are associated with the corresponding angular positions in the unified coordinate system through angular position spatial projection, thus achieving a unified representation of multimodal data under the same geometric coordinates. This system organizes detection data originally scattered across different time axes or different spatial reference systems into the same bearing surface geometric coordinate system, laying the foundation for subsequent point-by-point comparison and fusion.
[0022] The normalized raw detection data is fed into different branches of the multimodal joint semantic encoder. For example, image data goes through the residual convolutional network branch, time-domain signals go through the one-dimensional convolutional branch, acoustic features go through the spectral convolutional branch, and geometric deformation values go through the fully connected mapping branch, each extracting its corresponding initial feature vector. These feature vectors are then uniformly mapped to the same latent space, concatenated into a fused feature vector, and then dimensionality reduction is performed to obtain the initial latent vector. Here, the latent space can be understood as a compressed, highly semantically abstract vector space in which information from different modalities is represented as a set of feature components.
[0023] Production context parameters were introduced, including non-sensor information such as bearing model identifiers, current production process nodes, detection modal attributes, and real-time process environment variables. These parameters were encoded into fixed-length vectors, then mapped and projected to the same latent space dimension as the initial latent vectors, and then directly concatenated to obtain context-related latent vectors. This not only includes the semantics of the detection data itself but also incorporates the operating conditions and production circumstances at the time, enabling the model to distinguish between "deviations caused by defects" and "normal differences caused by process changes" when subsequently reconstructing a counterfactual normal image.
[0024] Consistency constraints are applied to the context-related latent vectors to ensure that the feature components of different modalities in the initial latent vectors maintain semantic consistency in the latent space, avoiding "semantic drift" between different modal features at the same physical location. A consistency loss function is introduced, such as calculating the distance between feature vectors of different modalities at the same spatial location. If the distance is too large, the latent vectors are updated through backpropagation, forcing the features of each modality to converge in the latent space. The latent vector of the test sample obtained after this constraint update becomes a compact feature representation that integrates multimodal detection information, is associated with the production context, and is semantically consistent and aligned. This provides high-quality input for subsequent construction of a counterfactual normal mirror using a conditional variational generative network.
[0025] In specific implementation, step S12 includes: The normalized raw detection data are input into the corresponding branch structure of the multimodal joint semantic encoder according to the preset modality type; Specifically, the bearing surface image data is input into the residual convolutional network branch, the bearing rotation time-domain signal is input into the one-dimensional convolutional network branch, the high-frequency acoustic feature data is input into the spectral convolutional branch, and the bearing rotation time-domain signal is input into the fully connected mapping branch, which respectively obtain the initial feature vector of each mode. A unified dimension mapping is performed on the initial feature vectors to map them to a unified latent space, resulting in standardized feature vectors. The calculation formula for these standardized feature vectors is as follows: ; in, To standardize the feature vector, For the first Mapping matrix corresponding to each mode For the first The initial feature vectors corresponding to each mode. It is the bias vector; The standardized feature vectors are concatenated, and the standardized feature vectors corresponding to each modality are combined in a preset order to obtain the fused feature vector; The fusion feature vector is reduced in dimension and mapped to a unified latent space to obtain the initial latent vector. The initial latent vector is composed of feature components corresponding to different modes.
[0026] Specifically, concatenating standardized feature vectors refers to directly combining standardized feature vectors from various modalities that have been mapped to a unified latent space but have not yet been fused, according to a predefined modal order, by concatenating them end-to-end along the feature dimension. For example, assuming the standardized feature vector from the image branch has a dimension of 128, the standardized feature vector from the temporal signal branch has a dimension of 64, the acoustic feature branch has a dimension of 64, and the geometric deformation numerical branch has a dimension of 32, then concatenating them in a preset order (e.g., image, temporal, acoustic, geometric deformation) will form a fused feature vector with a dimension of 128+64+64+32=288. The preset order is a fixed modal arrangement determined during the model design phase, ensuring that the initial feature vector positions are consistent each time they are concatenated, thus allowing subsequent dimensionality reduction mapping layers to learn stable and repeatable feature combination patterns. The concatenation operation itself does not involve any weighting or transformation; it is merely a physical combination of vectors along the dimension, bringing together features originally scattered at the ends of each modal branch into a unified tensor.
[0027] After concatenating the features to obtain the fused feature vector, a dimensionality reduction mapping is performed on it, compressing it from the concatenated high-dimensional space to a truly unified latent space. This latent space typically has a pre-defined dimension, much lower than the concatenated dimension, such as 64 or 128 dimensions. Dimensionality reduction is usually achieved through a fully connected layer. The weight matrix of this fully connected layer linearly transforms the high-dimensional fused feature vector to the low-dimensional target latent space, and a non-linear activation function is used to introduce some expressive power. The training objective of this fully connected layer is to eliminate redundancy between modalities while preserving as much key information as possible, and to organize the features of different modalities into a semantically consistent distribution in the latent space. The initial latent vector obtained after dimensionality reduction no longer belongs solely to a specific modality in each dimension, but rather represents abstract feature components resulting from the interaction and fusion of information from multiple modalities. This lays a solid foundation for subsequently introducing production context parameters and performing conditional variational reconstruction.
[0028] In specific implementation, step S2 includes: Step S21: The latent vector of the sample to be tested, the production context parameters, and the preset defect-free baseline label are taken as joint inputs and fed into the conditional variational generation network. Reconstruction operations are performed within the unified latent space manifold to obtain the initial counterfactual latent vector. The processing logic is as follows: Let the latent vector of the sample to be tested be denoted as... The vector after mapping the production context parameters is denoted as The preset defect-free status label is denoted as Construct joint input vector It is represented as: ; By mapping the joint input vector to a probability distribution, a set of latent variable distribution parameters is obtained. Then, sampling operations are performed based on this set to obtain the sampled latent variables, calculated using the following formula: ; in, To sample latent variables, It is the mean vector. The standard deviation vector, Let be a random variable that follows a standard normal distribution; The sampled latent variables are input into the latent space to reconstruct the mapping function, thus obtaining the initial counterfactual latent vector; Step S22: Perform latent space constraint calculation on the initial counterfactual latent vector, introduce the production context parameters as constraints into the unified latent space manifold, and map and adjust the initial counterfactual latent vector in combination with the manifold distribution corresponding to the preset standard state label to obtain the constrained counterfactual latent vector. Step S23: Perform reverse mapping processing on the constrained counterfactual latent vectors according to the preset modality type, convert the latent space representation into multimodal detection data form, and obtain the counterfactual normal mirror; Step S24: Perform coordinate mapping calculations on the original detection data and the counterfactual normal mirror according to a unified time reference and a unified spatial reference, and establish a positional correspondence between the original detection data and the counterfactual normal mirror in a unified micro-bearing surface geometric coordinate system. Step S25: Under the unified geometric coordinate system of the micro-bearing surface, a differential homeomorphic space transformation operator is constructed based on the position correspondence and global mutual information value. A local topology preservation constraint term is added to the differential homeomorphic space transformation operator, and a continuous space mapping operation is performed on the original detection data to obtain the aligned original detection data.
[0029] Specifically, the concepts of a pre-defined defect-free baseline label and a pre-defined standard state label are essentially the same thing, only with slightly different emphases in different steps. This label is a manually set semantic condition defined during the training phase, used to tell the generative network, "I need to generate a sample in a defect-free state." In conditional variational generative networks, this label is usually encoded as a vector, concatenated with the latent vector of the test sample and the production context parameters, serving as the input condition. Even if the production context parameters contain some variations related to the normal state (such as normal differences caused by different bearing models or different processes), the generative network can still reconstruct the "normal form" that the sample should have under the constraints of the current production conditions, with "defect-free" as the semantic goal, thereby separating the deviation caused by defects from the normal differences brought about by the process itself.
[0030] After constructing the joint input vector, it needs to be mapped to a probability distribution. Instead of directly inputting the joint input vector into the decoder, the encoder network first maps the joint input vector to a set of parameters representing the latent variable distribution: the mean vector and the standard deviation vector. These two parameters together define a Gaussian distribution in the latent space. The reason for not directly generating a deterministic latent vector is that the core idea of the variational autoencoder is to make the latent space continuous and locally smooth. Sampling enhances the diversity and robustness of the generated results. The sampling operation employs a reparameterization technique, ensuring that the gradient can propagate backwards while causing the latent vector obtained from each sampling to fluctuate around the mean vector. This results in a reasonable but slightly different counterfactual normal image after decoding.
[0031] After obtaining the sampled latent variables, they are input into the latent space reconstruction mapping function, which acts as the latent space transformation function in the conditional variational generation network. Through a multi-layer fully connected network, the sampled latent variables are nonlinearly transformed in a unified latent space to generate an initial counterfactual latent vector with defect-free semantics. The initial counterfactual latent vector is still in a unified latent space and contains all the feature information required to generate defect-free samples, but it has not yet undergone constraint adjustment and modality decoding.
[0032] The next step, S22, involves calculating latent space constraints on the initial counterfactual latent vector to ensure the generated result better reflects the "normal state that should be presented under current production conditions." These constraints include two aspects: first, production context parameters, which determine that the generated result must be consistent with the current bearing model, process node, and process environment; and second, the manifold distribution corresponding to the preset standard state label, which defines which region in the latent space should contain "defect-free samples." In practice, an additional constraint network or loss function is typically introduced. For example, adversarial constraints force the generated counterfactual latent vector to fall into the latent space manifold region corresponding to the standard state label. Simultaneously, the embedding vector of the production context parameters guides the generation path conditionally, ensuring that the final constrained counterfactual latent vector satisfies both the semantic requirement of being defect-free and matches the current production conditions.
[0033] After the constraints in the latent space are adjusted, the constraint counterfactual latent vectors are reverse mapped according to the preset modality type. That is, they are split into latent space sub-vectors corresponding to each modality, and then reconstructed into the actual detection data form through the decoder corresponding to each modality (such as deconvolution network, inverse spectral mapping, fully connected mapping). Finally, these reconstructed multimodal data are combined into a counterfactual normal image that is completely consistent with the original detection data structure.
[0034] Although the counterfactual normal image and the original test data semantically correspond to the "ideal normal state" and the "actual test state," their spatial positions within the unified geometric coordinate system of the micro-bearing surface may be offset or not perfectly matched. Therefore, it is necessary to establish a precise positional correspondence and perform continuous spatial mapping. The original test data and the counterfactual normal image are mapped using a unified time reference and a unified spatial reference. Here, the "unified time reference" for rotation-related signals is the rotation angle phase, and the "unified spatial reference" is the previously constructed bearing surface geometric coordinate system. Through these two references, the theoretical spatial position corresponding to each test data point in the counterfactual normal image can be found, establishing a preliminary point-to-point correspondence.
[0035] Based on these correspondences and global mutual information values, a differential homeomorphic spatial transformation operator is constructed. A differential homeomorphic transformation is a continuous, reversible, and topologically preserving spatial mapping that can smoothly correct deformations in images or spatial fields without disrupting local geometry. By incorporating local topology-preserving constraints, this transformation operator can continuously "distort" the original detection data to a state precisely aligned spatially with its counterfactual normal mirror image, ensuring pixel-level or point-to-point correspondence at every physical location on the bearing surface. The resulting "aligned original detection data" is spatially perfectly matched with its counterfactual normal mirror image, allowing for direct point-by-point calculation of the deviation vector, forming a defect deviation field that accurately reflects the spatial distribution and severity of defects.
[0036] In specific implementation, step S23 includes: The constrained counterfactual latent vectors are input into the latent space projection layer corresponding to each modality. Linear projection operations are then used to extract the latent space sub-vectors corresponding to each modality from the constrained counterfactual latent vectors. The calculation formula is as follows: ; in, The first one extracted from the constrained counterfactual latent vector. The latent space sub-vectors corresponding to each mode. For the first The first projection parameters corresponding to each mode To constrain the counterfactual latent vector, For the first The second projection parameters corresponding to each mode; The latent space sub-vectors of each mode are subjected to inverse mapping processing. The latent space sub-vectors of each mode are input into the corresponding inverse mapping function to obtain the reconstructed detection data of each mode. The calculation formula is as follows: ; in, For the first Reconstruction detection data corresponding to each modality For the first The inverse mapping function corresponding to each mode The first one extracted from the constrained counterfactual latent vector. The latent space sub-vectors corresponding to each mode; The reconstructed detection data of each modality is organized according to the preset modality type, and the reconstructed detection data of each modality is arranged according to the data structure consistent with the normalized original detection data to obtain a multimodal reconstruction data set; The multimodal reconstructed data set is combined according to a unified time reference and a unified spatial reference to generate a counterfactual normal image with the same structure as the original detection data.
[0037] Specifically, after dimensionality reduction mapping of the fused feature vector in step S12, the resulting initial latent vector resides in a unified latent space. Its dimensions are coupled feature components resulting from the fusion of multimodal information through linear and nonlinear mappings, no longer maintaining the interval structure of the original modal features during the splicing stage. Therefore, it lacks the conditions for component segmentation based on modal order. After constraint adjustment in steps S21 and S22, the resulting constrained counterfactual latent vector still resides in the unified latent space. Its feature distribution is the overall representation after multimodal semantic fusion, and its dimensions again do not correspond to independent features of a specific modality. Based on these characteristics, before performing modal inverse mapping, a corresponding latent space projection layer is constructed for each modality. The constrained counterfactual latent vector is input into the latent space projection layer corresponding to each modality. A linear mapping matrix is used to extract the feature subspace representation related to the target modality semantics from the unified latent space representation, resulting in the corresponding latent space subvector.
[0038] The latent space sub-vectors are vector representations obtained through projection operations in the unified latent space. They are used to characterize the semantic information of the corresponding modality and serve as inputs for subsequent modality decoding processes. The mapping parameters of the latent space projection layers of each modality are optimized through a multimodal joint training process, thereby enhancing the separability of different modal semantics in the unified latent space.
[0039] The inverse mapping process for the latent space sub-vectors of each modality involves inputting the latent space sub-vectors of each modality into a modality-specific inverse mapping function to reconstruct the detection data for that modality. The inverse mapping function is essentially the inverse process of each branch structure in the multimodal joint semantic encoder. For image modalities, the inverse mapping function is the inverse process of a residual convolutional network, typically using deconvolution or upsampling followed by convolution to progressively restore the latent space sub-vectors to the pixel matrix of the bearing surface image data. For time-domain signal modalities, the inverse mapping function is the inverse process of a one-dimensional convolutional network, using one-dimensional deconvolution to restore the latent space sub-vectors to a time-domain waveform sequence. For high-frequency acoustic feature data, the inverse mapping function is the inverse process of a spectral convolution branch, restoring the latent space sub-vectors to spectral features or time-domain acoustic signals. For geometric deformation values, the inverse mapping function is the inverse process of a fully connected mapping branch, mapping the latent space sub-vectors back to the original dimension of the geometric deformation values through a fully connected layer. The inverse mapping function for each modality is jointly optimized with the encoder during training to ensure that the reconstructed data retains the semantic information of the original data to the greatest extent.
[0040] The reconstructed detection data for each modality is organized according to a preset modality type. The reconstructed data for each modality still exists independently. For example, image data is a two-dimensional pixel matrix, the temporal signal is a one-dimensional array, and the geometric deformation value is a scalar field organized by spatial location; their data structures and physical meanings are different. Arranging the data according to a data structure consistent with the normalized original detection data means integrating these independent modal data according to the organization method of the normalized original detection data in step S11, forming a multimodal reconstructed data set. This data structure is usually a composite data structure containing multiple fields, such as a dictionary or a structured array, where each field corresponds to a modality. The data arrangement within each field (such as the width and height order of the image, the temporal order of the temporal signal, and the spatial index order of the geometric deformation value) is exactly the same as the normalized original detection data, ensuring that subsequent processing can access the data of each modality in a unified way without needing to re-parse the format.
[0041] Combining the multimodal reconstructed dataset according to a unified time and spatial reference transforms it into a complete counterfactual normal image that is structurally identical to the original detection data. This "combination" is not simply data packaging; it requires associating data from different modalities with the same spatiotemporal reference system at the physical semantic level. Image data and geometric deformation values already possess spatial coordinate attributes; during combination, it's sufficient to confirm they all reside within a unified geometric coordinate system of the micro-bearing surface. For time-domain signals and high-frequency acoustic feature data, which were originally mapped to the unified geometric coordinate system via angular spatial projection, combination requires restoring them from their angular domain representation to a form bound to spatial coordinates—that is, associating the signal value at each angular position with the corresponding geometric coordinates of the bearing surface. In this way, the resulting counterfactual normal image not only matches the original detection data in data structure but is also perfectly aligned in spatiotemporal reference, allowing subsequent steps to directly calculate point-by-point deviation vectors.
[0042] In specific implementation, step S25 includes: Under the unified geometric coordinate system of the micro-bearing surface, the positional correspondence between the original detection data and the counterfactual normal mirror is obtained, and the global mutual information value between the original detection data and the counterfactual normal mirror is calculated to construct the corresponding position set; Based on positional correspondence and global mutual information, a differential homeomorphic spatial transformation operator is constructed in a unified geometric coordinate system of the micro-bearing surface. Local topology preservation constraints are added to the transformation objective function of the differential homeomorphic spatial transformation operator to perform continuous mapping calculations on the original spatial coordinates of the original detection data. The calculation formula is as follows: ; in, These are the mapped spatial coordinates. For differential homeomorphic space transformation operators, The original spatial coordinates of the original test data in the unified geometric coordinate system of the micro-bearing surface; The operator for transforming a differential homeomorphic space is subjected to continuous differentiability constraints, and its Jacobian matrix is subjected to non-singularity constraints based on local topological preservation constraints. Its expression is as follows: ; in, For the differential homeomorphic space transformation operator in the original space coordinates Jacobian matrix at the location, To find the determinant of the Jacobian matrix; Based on the mapped spatial coordinates, values are calculated in the original detection data according to the spatial index relationship. The values of the original detection data at the original spatial coordinates are then mapped back to the mapped spatial coordinates to obtain the aligned original detection data.
[0043] Specifically, mutual information is an information-theoretic similarity measure used to assess the statistical dependence between two datasets. The mutual information value reaches its maximum when two images or spatial fields are geometrically well aligned. Calculating the global mutual information value between the original detection data and the counterfactual normal image in step S25 essentially provides a global similarity optimization target for subsequent spatial transformations. Specifically, the system treats the values at corresponding positions of the original detection data and the counterfactual normal image in a unified geometric coordinate system of the micro-bearing surface as two sets of random variables. A joint histogram is constructed to statistically analyze the probability distribution of different value pairs. Then, the mutual information value is calculated according to the definition formula, i.e., the marginal entropy of each variable and their joint entropy are calculated separately, and the sum of the marginal entropies is subtracted from the joint entropy to obtain the mutual information value. This global mutual information value is not directly used for the transformation but serves as one of the reference bases when constructing the differential homeomorphic transformation operator, helping to determine the initial direction and range of the transformation.
[0044] A differential homeomorphism is a continuously differentiable and invertible spatial mapping. Its most significant characteristic is that it preserves the topological structure of space; that is, originally adjacent points remain adjacent after the transformation, and originally disjoint regions do not overlap or tear apart. In actual construction, the system designs a transformation objective function based on the previously established positional correspondences and global mutual information values. This objective function typically includes three parts: first, a similarity metric, usually using mutual information or mean squared error to measure the degree of matching between the transformed original detection data and its counterfactual normal mirror image; second, a regularization term, used to constrain the smoothness of the transformation and prevent excessively drastic local deformation; and third, a local topology preservation constraint term. The specific form of the local topology preservation constraint term is usually to impose a constraint on the Jacobian matrix of the transformation, requiring its determinant to be greater than zero throughout the entire spatial domain, ensuring that the transformation is locally invertible and orientation-preserving near each point, without folding or reversal. After adding this constraint to the objective function, the system uses an iterative optimization algorithm to solve for the optimal transformation field that satisfies all constraints. This transformation field constitutes the differential homeomorphic spatial transformation operator, which can continuously map each original spatial coordinate of the original detection data in a unified coordinate system to obtain the mapped spatial coordinates.
[0045] Value calculation based on mapped spatial coordinates involves resampling the original detection data. Since the differential homeomorphic spatial transformation operator maps the original spatial coordinates to a new location, and the values of the original detection data at the original coordinates are known, interpolation is needed to determine the values at the new coordinates. Specifically, the system traverses the spatial grid containing the counterfactual normal image. For each target coordinate point on the grid, an inverse transformation is used to find the corresponding original coordinates in the original detection data. Then, according to spatial indexing relationships, interpolation calculations are performed on the neighboring points around the original coordinates in the original detection data to obtain the value at that target coordinate point. The interpolation method differs for different modalities: image data and geometric deformation values, which have continuous spatial distributions, typically use bilinear or trilinear interpolation; while time-domain signals and high-frequency acoustic feature data, after angular spatial projection, are essentially bound to angular coordinates. Therefore, interpolation needs to be performed using neighboring or linear interpolation along the angular dimension to ensure the integrity of the signal sequence. In this way, the values of the original detection data at all spatial locations are remapped to the new coordinates after mapping. The resulting aligned original detection data achieves precise point-to-point matching with the counterfactual normal mirror in space, thus providing a data foundation with a complete spatial correspondence for the next step of calculating the defect deviation field.
[0046] In specific implementation, step S3 includes: Step S31: Under the unified geometric coordinate system of the miniature bearing surface, the aligned original detection data and the counterfactual normal mirror image are compared at corresponding positions to calculate the deviation vector, forming an initial deviation set. The calculation formula is as follows: ; in, For the set of deviation vectors, The original detection data after alignment. To be a counterfactual mirror image, This is the difference evaluation function corresponding to the preset modality type; When the preset mode type is bearing surface image data or bearing geometric deformation value, the difference evaluation function is the absolute difference function; When the preset mode type is a bearing rotation time-domain signal or high-frequency acoustic feature data, the difference evaluation function is the feature divergence function; Step S32 involves associating the set of deviation vectors according to a unified geometric coordinate system for the surface of the micro-bearing, binding each deviation vector to its corresponding mapped spatial coordinates to form a set of corresponding deviation vectors and mapped spatial coordinates. The processing logic is as follows: The spatial coordinate set of the aligned original detection data in the unified micro-bearing surface geometric coordinate system is obtained. The spatial coordinates are sorted according to the coordinate arrangement rules in the unified micro-bearing surface geometric coordinate system, and the sorted spatial coordinates are assigned corresponding index numbers to obtain the spatial position index. The spatial coordinate index is obtained by indexing each deviation vector in the deviation vector set according to its positional correspondence. According to the spatial coordinate index, each deviation vector is associated with the mapped spatial coordinates to establish a correspondence between the deviation vectors and the spatial coordinates. This correspondence is then arranged according to the spatial coordinate index order to obtain the set of correspondences between the deviation vectors and the mapped spatial coordinates. The calculation formula is as follows: ; in, This is the set of corresponding offset vectors and mapped spatial coordinates. To unify the geometric coordinate system of miniature bearing surfaces, the first A mapped spatial coordinate, The deviation vector in the deviation vector set corresponds to the mapped spatial coordinates. This refers to the spatial coordinate index number; Step S33: Under the unified geometric coordinate system of the micro bearing surface, the deviation vector set is divided into regions according to the preset spatial division method, and the deviation vectors are assigned to the corresponding regions according to the mapped spatial coordinates to obtain the regionalized deviation set. Generate the corresponding spatial location index based on the order of the mapped spatial coordinates; Step S34: Organize the regionalized deviation set according to spatial location index, arrange the deviation vectors in each region in a preset order, and construct the defect deviation field. The processing logic is as follows: Traverse the set of corresponding offset vectors and mapped spatial coordinates, extract all spatial coordinates, and sort the mapped spatial coordinates according to the spatial position index under the unified micro-bearing surface geometric coordinate system to obtain an ordered spatial coordinate sequence. The corresponding deviation vectors are synchronously rearranged based on the ordered spatial coordinate sequence to obtain an ordered deviation vector sequence. The ordered deviation vector sequence is combined with the corresponding spatial coordinate sequence and arranged according to the spatial position index to obtain an ordered correspondence sequence; The ordered correspondence sequence is organized according to a unified geometric coordinate system of the micro-bearing surface, and the correspondence between the deviation vector and the mapped spatial coordinates is uniformly represented to obtain the defect deviation field.
[0047] Specifically, different modal data differ in their feature representation and statistical properties, therefore, different types of evaluation functions are used when assessing differences. When the modal type is bearing surface image data or bearing geometric deformation values, both types of data are essentially direct physical quantities in the spatial domain. The pixel values of the image data represent surface grayscale or reflectivity, while the geometric deformation values represent the actual deformation at that location. Both have the same physical meaning and dimensions as the values at the corresponding locations in the counterfactual normal mirror image, and therefore can be directly subjected to algebraic operations. The absolute difference function is most suitable in this case, that is, calculating the absolute value of the difference between the values of the aligned original detection data and the counterfactual normal mirror image at the same spatial location. This difference is used to characterize the degree of deviation between the actual state and the normal state. In some implementations, the difference function can also adopt a signed difference form, which retains the sign information of the difference to reflect the direction of deviation, thus making it suitable for application scenarios that analyze changing trends or directional features. This difference directly reflects the degree of deviation between the actual state and the normal state at that location, with a clear physical meaning and simple calculation. In some implementations, for the aforementioned spatial domain data, the difference function can also adopt different forms of numerical difference measurement methods, including signed difference forms or other equivalent transformation forms, to meet the needs of analyzing the directionality or trend of difference. However, the situation is completely different when the modal type is a bearing rotation time-domain signal or high-frequency acoustic feature data. These two types of data are essentially one-dimensional time series or spectral series, and the numerical values themselves do not have absolute spatial meaning. What truly reflects the bearing state is the structural information contained in the signal, such as waveform characteristics, spectral distribution, and harmonic components. If the absolute difference is calculated directly, due to factors such as phase drift and amplitude fluctuation, even a defect-free normal signal may produce a large numerical difference, causing a large number of false alarms. Extracting physically meaningful feature vectors from the signal, such as the root mean square, peak factor, kurtosis, and other statistical features of the time-domain signal, or the dominant frequency energy distribution and sideband features of the frequency-domain signal, and then calculating the divergence measure between these feature vectors, such as KL divergence, cosine distance, or Mahalanobis distance, can more robustly reflect the degree of anomaly in the signal and avoid misjudgment due to subtle changes in signal sampling. In some implementations, for time-domain signals or acoustic feature data, the difference evaluation function can also be calculated using a combination of various measurement methods, including combining the difference function with the feature divergence function to simultaneously characterize local numerical differences and overall distribution differences, thereby improving the comprehensiveness of the difference evaluation.
[0048] Each deviation vector is bound to its corresponding spatial coordinates to form an indexable set by establishing a unified spatial coordinate indexing system. The system first acquires all spatial coordinates of the aligned raw detection data in a unified geometric coordinate system for the micro-bearing surface. Then, it sorts the coordinates according to the rules of this coordinate system, which are explained in parentheses in the documentation: sorting by the first coordinate component, then by the second coordinate component if the first coordinate components are the same, and so on if a third coordinate component exists. This sorting method essentially defines a space-filling curve, mapping discrete points in two-dimensional or three-dimensional space to a one-dimensional ordered sequence, thus assigning a unique index number to each spatial coordinate. For image data, this sorting rule is equivalent to scanning by row and then by column; for the bearing surface in polar coordinates, the order is first radius, then angle, and then axial height. After sorting the spatial coordinates and assigning indices, the system indexes each deviation vector in the deviation vector set according to the positional correspondence. That is, each deviation vector finds the index number of its corresponding spatial coordinate, and then the deviation vectors are matched with the spatial coordinates one by one according to the order of the index numbers to form a corresponding set.
[0049] Preset spatial partitioning is a strategy for structurally dividing the geometric region of a bearing surface. It divides a continuous bearing surface into several physically meaningful sub-regions, facilitating subsequent analysis of defect distribution by region. Preset spatial partitioning is typically designed based on the bearing's geometric characteristics, such as dividing it according to raceway position into inner ring raceways, outer ring raceways, cage areas, and seal groove areas, or dividing it into sectors based on angles or strips based on radii. Different regions have different sensitivities and tolerances to defects. Organizing deviation vectors by region allows for more targeted defect analysis and facilitates differentiated processing based on regional importance during data sharing. When assigning deviation vectors to corresponding regions according to their mapped spatial coordinates, the system determines which preset region each spatial coordinate belongs to based on its geometric location, and then adds the deviation vector at that coordinate to the corresponding region's data set, forming a regionalized deviation set.
[0050] The regionalized deviation set is then finally organized and formatted to construct a complete defect deviation field. First, the corresponding set of deviation vectors and spatial coordinates is traversed, all spatial coordinates are extracted, and sorted according to spatial position indexing rules to obtain an ordered spatial coordinate sequence. Then, the deviation vectors are synchronously rearranged according to this sequence; that is, if the deviator in the ordered spatial coordinate sequence... The coordinates are , then the first The deviation vector should be the original and Bound This ensures that the deviation vectors and spatial coordinates are completely consistent in order. Next, the ordered sequence of deviation vectors is combined with the ordered sequence of spatial coordinates, arranged in index order to form an ordered correspondence sequence. The final step is to organize this ordered correspondence sequence according to a unified geometric coordinate system of the micro-bearing surface, representing the correspondence between the deviation vectors and spatial coordinates as a structured data field. In specific implementations, this unified representation is usually manifested as a multi-dimensional array or grid structure, where each element of the array corresponds to a spatial coordinate position, and the value of the element is the deviation vector at that position. For two-dimensional grids like image data, the defect deviation field is a two-dimensional matrix, with each element being a scalar or vector; for bearing surfaces in polar coordinates, a sparse structure or non-uniform grid may be used for storage. Regardless of the storage format, the final generated defect deviation field possesses two key characteristics: first, the spatial position and deviation value strictly correspond; second, the data structure is well-organized, facilitating serialization and transmission, thus preparing for subsequent data encapsulation and distribution.
[0051] In specific implementation, step S4 includes: Step S41: Analyze the defect deviation field, extract the correspondence between the deviation vector and the mapped spatial coordinates, and index the correspondence according to the spatial position index under the unified micro bearing surface geometric coordinate system to obtain the deviation field index data. Step S42: Extract the index identifier of the counterfactual normal image, and perform identifier mapping processing on the counterfactual normal image according to the unified time reference and unified spatial reference, and associate the index identifier of the counterfactual normal image with the spatial coordinate index in the defect deviation field. Step S43: Encode the production context parameters according to the preset parameter structure, and associate and calibrate the production context parameters according to the spatial location index, and associate the production context parameters with the spatial coordinate index in the defect deviation field. Step S44: Combine the off-field index data, the index identifier of the counterfactual normal mirror, and the production context parameters according to a unified data format, and organize the combined results in a structured manner to obtain shared data. Step S45: Sort the shared data according to the spatial location index, group it according to the production context parameters, and perform the distribution operation according to the preset distribution rules. The preset distribution rules include preset data splitting rules, preset field mapping rules, preset data frame generation rules, and preset sending order rules.
[0052] Specifically, the indexing and calibration mentioned in step S41, which involves indexing the corresponding relationships according to the spatial position index in the unified micro-bearing surface geometric coordinate system, essentially establishes a standardized spatial indexing system based on the defect deviation field. Since the defect deviation field itself is composed of the correspondence between deviation vectors and spatial coordinates, and the spatial coordinates have already been assigned unique index numbers according to the coordinate arrangement rules in step S32, the so-called index calibration here is to standardize and extract this existing indexing system to form deviation field index data. The deviation field index data is actually a structured directory that records the spatial coordinates corresponding to each spatial position index and the deviation vector at that position.
[0053] The counterfactual normal image itself is a multimodal data set, containing images, time-domain signals, acoustic features, and geometric deformation values, all of which need to be associated with spatial coordinate indices in the defect deviation field. Identifying and mapping the counterfactual normal image according to a unified time and spatial reference means mapping each data unit in the counterfactual normal image to a corresponding index in a unified micro-bearing surface geometric coordinate system, based on its spatial location and time phase. For image data and geometric deformation values, each pixel or measurement point has clear spatial coordinates and can be directly mapped to a spatial index. For time-domain signals and high-frequency acoustic feature data, it is necessary to first extract the rotation angle phase according to the unified time reference, then map the angle phase to a spatial location, and finally to a spatial index. After mapping, the index identifiers of the counterfactual normal image form an index structure that completely corresponds to the defect deviation field; that is, each spatial index is associated with both the deviation vector at that location and the corresponding normal image data.
[0054] The generation of counterfactual normal images is based on a conditional variational generative network. The generation process takes production context parameters as conditional inputs and generates corresponding counterfactual normal image data under the defect-free semantic constraints learned by the model. Index identifiers are used to locate and associate the generated results under a unified time and spatial reference.
[0055] Production context parameters include bearing model identifiers, current production process nodes, detection modal attributes, and real-time process environment variables. Some of these parameters are related to the entire sample, such as bearing model and process node, while others may be related to specific spatial locations, such as differences in process environments across different regions. Encoding these parameters according to a preset parameter structure involves converting them into numerical vectors or structured key-value pairs in a fixed format, and then associating them based on spatial location indices. For global parameters, the same parameter value can be associated with all spatial indices; for regional parameters, they are associated with spatial indices within the corresponding spatial region. This ensures that in subsequent data analysis and sharing, users can not only obtain information on the spatial distribution of defects but also understand the production environment conditions at the time the defect occurred.
[0056] The three types of data mentioned above are combined and processed according to a unified data format. This unified data format is typically a structured serialization format, such as JSON, Protocol Buffers, or a custom binary frame format with clearly defined fields. During the combination process, the data organization hierarchy needs to be determined, using spatial location indexes as the first level of organization. The deviation vectors, counterfactual normal image data fragments, and related production context parameters corresponding to each index are combined into a single data record. Then, all records corresponding to each index are arranged into a sequence according to the index order. Meta-information, such as data version number, coordinate system definition, modality type description, and total number of records, needs to be added to the data header to ensure correct parsing by the receiving end. After this structured processing, the originally scattered defect deviation fields, counterfactual normal images, and production context parameters are merged into a self-contained, independently transmittable shared data packet.
[0057] The preset distribution rules constitute a complete data transmission protocol. Preset data splitting rules define how shared data is divided into multiple data blocks before transmission. These rules consider network maximum transmission unit limitations, receiver buffering capacity, and transmission reliability requirements. The splitting rules specify the size boundaries of each data block, block numbering methods, and dependencies between blocks. Preset field mapping rules define how various fields in the shared data are mapped to the data frame structure of the transmission protocol. For example, spatial indices are mapped to address fields in the frame header, offset vectors are mapped to the payload data area, and production context parameters are mapped to extension fields. This ensures that the sender and receiver have a consistent understanding of the data's meaning. Preset data frame generation rules specify how the field-mapped data is assembled into data frames conforming to a specific communication protocol. These rules include frame start flags, checksum calculation methods, frame length encoding, and frame sequence number management. These rules guarantee the integrity and verifiability of data during physical transmission. The preset sending order rules define the sending order of each data frame. They may be sent in spatial index order, in priority grouping, or in an interleaved manner according to a certain scheduling strategy. At the same time, retransmission mechanisms and flow control strategies also need to be defined.
[0058] At the receiving end, data recovery is performed in reverse order of the processing rules: First, the received transmission data frames are rearranged according to the sequence index; then, the index identifiers of the defect deviation field and the counterfactual normal image, as well as the production context parameters, are obtained by parsing the preset field structure; subsequently, the conditional variational generation network is called based on the production context parameters to generate the counterfactual normal image, and the spatial position correspondence processing of the counterfactual normal image is performed according to the index identifier; finally, the counterfactual normal image and the defect deviation field are superimposed to obtain the corresponding detection data.
[0059] In specific implementation, step S45 includes: According to the preset data splitting rules, the shared data is parsed and split according to the preset field structure. The index identifiers of defect deviation fields and counterfactual normal mirrors and production context parameters are extracted from the shared data. According to the preset field mapping rules, the index identifiers of the defect deviation field and the counterfactual normal image and the production context parameters are mapped according to the preset communication protocol and written to the corresponding positions in the order of the fields to generate transmission data frames. According to the preset data frame generation rules, the transmitted data frames are arranged in a preset order to obtain a sequence index, and each transmitted data frame is numbered based on the sequence index to form a data transmission sequence; According to the preset sending order rules, the data sending sequence is written into the network transmission interface, and each data frame is sent frame by frame in the order of the sequence index, while the corresponding sequence index is recorded. At the receiving end, the received transmission data frames are rearranged according to the sequence index, and the index identifiers of the defect deviation field and the counterfactual normal image and the production context parameters are parsed according to the preset field structure. The index identifier of the counterfactual normal image is input into the conditional variational generation network to extract the counterfactual normal image, and the counterfactual normal image and the defect deviation field are superimposed to restore the corresponding data content.
[0060] Specifically, a predefined field structure is a pre-agreed data organization format that specifies the order, data type, length, and nesting relationship of the components of shared data in a serialized state. For example, a typical field structure might be defined as follows: the header field contains the data version number, timestamp, and sample ID; the index field contains the range and encoding method of the spatial location index; the body fields sequentially list the defect deviation field data block, the counterfactual normal mirror index identifier data block, and the production context parameter data block; and the tail field contains the checksum and end marker. The significance of this field structure lies in providing a unified standard for data splitting, transmission, and parsing, enabling the sender and receiver to accurately understand the boundaries and meaning of the data without additional metadata. Data splitting according to the predefined field structure means parsing the shared data into independent field units based on the field boundaries defined in this structure, and then extracting the defect deviation field, counterfactual normal mirror index identifier, and production context parameters respectively, preparing for subsequent field mapping.
[0061] The pre-defined communication protocol defines how these data fields are encapsulated into data frames transmitted over the network. Communication protocols typically include frame format definitions, field mapping rules, and transmission control mechanisms. At the frame format level, a typical data frame may consist of three parts: a frame header, a payload, and a frame trailer. The frame header contains the frame sequence number, frame type, and payload length; the payload carries the actual data content; and the frame trailer contains checksum information. Pre-defined field mapping rules, within the communication protocol framework, map the various fields extracted from shared data to specific locations within the data frame. For example, defect deviation fields are mapped to the payload area of multiple consecutive data frames; the index identifier of the counterfactual normal image is mapped to an extended field in the frame header; and production context parameters are mapped to specific types of control frames. This ensures that the mapped data meets the constraints of the communication protocol, such as payload size limits, alignment requirements, and checksum ranges, while also ensuring that the mapping relationship is reversible so that the receiving end can accurately reconstruct the data.
[0062] The pre-defined data frame generation rules further specify how to organize the mapped data into a complete data frame sequence. When generating transmission data frames, the first step is to determine the transmission order of each data frame. This order is typically arranged in ascending order of spatial location indices, because the defect deviation field itself is organized by spatial index. Sending them in this order allows the receiver to gradually reconstruct the spatial distribution of defects. For each data frame, the system assigns it a sequence index, a monotonically increasing number that reflects the data frame's position in the sequence and is also used for order reordering and packet loss detection at the receiver. The sequence index is usually placed in the frame header of the data frame and transmitted along with the payload data. After numbering, all data frames are arranged into a data transmission sequence according to their sequence index order.
[0063] The preset transmission order rules control how the data transmission sequence is written into the network transmission interface and actually transmitted. These rules typically include transmission rate control, retransmission mechanisms, and flow control. For example, the rules might stipulate that data is transmitted frame by frame in sequence index order, waiting for an acknowledgment from the receiver after each frame is transmitted, and retransmitting if no acknowledgment is received within a timeout period; alternatively, a sliding window mechanism might be used, allowing multiple frames to be transmitted consecutively within a window to improve transmission efficiency. During transmission, the transmitted sequence index is recorded simultaneously to track transmission progress and handle abnormal situations.
[0064] Because network transmission may experience latency jitter, packet loss retransmission, and multipath transmission, the order of data frames received at the receiving end may differ from the sending order. Reordering involves using the sequence index in the header of each data frame to rearrange the received frames according to their index order. A receive buffer is maintained; when a new data frame is received, it is inserted into the correct position in the buffer based on its sequence index. Simultaneously, it checks if a continuous sequence of frames has been completely received. Once a continuous sequence of frames is formed, it can be submitted to the upper layer for parsing and processing in sequence. This reordering mechanism ensures that even if the network transmission order is disordered, the ultimately restored data order is correct.
[0065] After the data frame is rearranged, the receiving end parses it according to a preset field structure, extracting the defect deviation field, the index identifier of the counterfactual normal image, and production context parameters. The index identifier of the counterfactual normal image is crucial, pointing to the corresponding image already stored or regenerable in the conditional variational generation network. The receiving end inputs this index identifier into the conditional variational generation network, extracting the complete counterfactual normal image through the network's forward computation. This image is then superimposed on the transmitted defect deviation field. The specific superposition method depends on the modal type: for image data and geometric deformation values, superposition is a direct pointwise addition or subtraction; for time-domain signals and acoustic feature data, superposition may involve signal reconstruction and feature fusion. After the superposition operation, the receiving end successfully recovers the original detection data content consistent with the sending end, completing the closed loop of the entire data sharing process.
[0066] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0067] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the protection scope of the present invention.
Claims
1. A method for sharing production and inspection data of micro-bearings based on deep learning, characterized in that, Includes the following steps: Step S1: Obtain the original detection data and production context parameters of the miniature bearing, and use a multimodal joint semantic encoder to perform feature fusion and dimensionality reduction mapping to output the latent vector of the sample to be tested. Step S2: Using production context parameters as constraints, a reconstruction operation is performed in a unified latent space through a conditional variational generator network to obtain a counterfactual normal image. Based on a unified time reference and a unified spatial reference, the original detection data and the counterfactual normal image are aligned with each other to obtain aligned original detection data. Step S3: Calculate the deviation vector of the aligned original detection data relative to the counterfactual normal mirror image, and organize the deviation vector according to the spatial location index to obtain the defect deviation field, specifically including: Step S31: Under the unified geometric coordinate system of the miniature bearing surface, the aligned original detection data and the counterfactual normal mirror image are compared at corresponding positions to calculate the deviation vector, forming an initial deviation set. The calculation formula is as follows: ; in, For the set of deviation vectors, The original detection data after alignment. To be a counterfactual mirror image, This is the difference evaluation function corresponding to the preset modality type; When the preset mode type is bearing surface image data or bearing geometric deformation value, the difference evaluation function is the absolute difference function; When the preset mode type is a bearing rotation time-domain signal or high-frequency acoustic feature data, the difference evaluation function is the feature divergence function; Step S32 involves associating the set of deviation vectors according to a unified geometric coordinate system for the surface of the micro-bearing, binding each deviation vector to its corresponding mapped spatial coordinates to form a set of corresponding deviation vectors and mapped spatial coordinates. The processing logic is as follows: The spatial coordinate set of the aligned original detection data in the unified micro-bearing surface geometric coordinate system is obtained. The spatial coordinates are sorted according to the coordinate arrangement rules in the unified micro-bearing surface geometric coordinate system, and the sorted spatial coordinates are assigned corresponding index numbers to obtain the spatial position index. The spatial coordinate index is obtained by indexing each deviation vector in the deviation vector set according to its positional correspondence. According to the spatial coordinate index, each deviation vector is associated with the mapped spatial coordinates to establish a correspondence between the deviation vectors and the spatial coordinates. This correspondence is then arranged according to the spatial coordinate index order to obtain the set of correspondences between the deviation vectors and the mapped spatial coordinates. The calculation formula is as follows: ; in, This is the set of corresponding offset vectors and mapped spatial coordinates. To unify the geometric coordinate system of miniature bearing surfaces, the first A mapped spatial coordinate, The deviation vector in the deviation vector set corresponds to the mapped spatial coordinates. This refers to the spatial coordinate index number; Step S33: Under the unified geometric coordinate system of the micro bearing surface, the deviation vector set is divided into regions according to the preset spatial division method, and the deviation vectors are assigned to the corresponding regions according to the mapped spatial coordinates to obtain the regionalized deviation set. Generate the corresponding spatial location index based on the order of the mapped spatial coordinates; Step S34: Organize the regionalized deviation set according to spatial location index, arrange the deviation vectors in each region in a preset order, and construct the defect deviation field. The processing logic is as follows: Traverse the set of corresponding offset vectors and mapped spatial coordinates, extract all spatial coordinates, and sort the mapped spatial coordinates according to the spatial position index under the unified micro-bearing surface geometric coordinate system to obtain an ordered spatial coordinate sequence. The corresponding deviation vectors are synchronously rearranged based on the ordered spatial coordinate sequence to obtain an ordered deviation vector sequence. The ordered deviation vector sequence is combined with the corresponding spatial coordinate sequence and arranged according to the spatial position index to obtain an ordered correspondence sequence; The ordered correspondence sequence is organized according to a unified geometric coordinate system of the micro-bearing surface, and the correspondence between the deviation vector and the mapped spatial coordinates is uniformly represented to obtain the defect deviation field; Step S4: Encapsulate the index identifiers and production context parameters of the defect deviation field, counterfactual normal image, and obtain shared data, and perform the distribution operation.
2. The method for sharing micro-bearing production and testing data based on deep learning as described in claim 1, characterized in that, Step S1 specifically includes: Step S11: Divide the raw detection data according to the preset modal type, and perform synchronous alignment processing under a unified time reference and a unified spatial reference to obtain the normalized raw detection data; The original detection data includes bearing surface image data, bearing rotation time-domain signal, high-frequency acoustic feature data, bearing geometric deformation values, and geometric structural parameters of the miniature bearing; The geometric parameters of the miniature bearing include the inner ring size, outer ring size, raceway structure parameters, and axial structure parameters. Based on the normalized original test data, a unified micro-bearing surface geometric coordinate system is constructed according to the geometric structural parameters of the micro-bearing. The bearing surface image data and bearing geometric deformation values in the normalized original test data are mapped to the unified micro-bearing surface geometric coordinate system according to their physical spatial positions. The bearing rotation time domain signal and high-frequency acoustic feature data in the normalized original test data are extracted according to a unified time reference to extract the corresponding rotation angle phase, and are jointly mapped to the unified micro-bearing surface geometric coordinate system for unified representation by means of angular position spatial projection. Step S12: Input the normalized original detection data into the branch structure corresponding to the multimodal joint semantic encoder, perform feature fusion processing on the detection data of different modalities, and perform dimensionality reduction mapping in the unified latent space to obtain the initial latent vector. Step S13: Map the production context parameters to the unified latent space and concatenate them with the initial latent vector to obtain the context-related latent vector; The production context parameters include bearing model identifier, current production process node, detection modal attributes, and real-time process environment variables; Step S14: Perform consistency constraint processing on the context-related latent vector, perform constraint calculation on the differences between each feature component in the initial latent vector, and update the context-related latent vector according to the constraint calculation results to obtain the latent vector of the sample to be tested.
3. The method for sharing micro-bearing production and testing data based on deep learning as described in claim 2, characterized in that, Step S12 specifically includes: The normalized raw detection data are input into the corresponding branch structure of the multimodal joint semantic encoder according to the preset modality type; Specifically, the bearing surface image data is input into the residual convolutional network branch, the bearing rotation time-domain signal is input into the one-dimensional convolutional network branch, the high-frequency acoustic feature data is input into the spectral convolutional branch, and the bearing rotation time-domain signal is input into the fully connected mapping branch, which respectively obtain the initial feature vector of each mode. A unified dimension mapping is performed on the initial feature vectors to map them to a unified latent space, resulting in standardized feature vectors. The calculation formula for these standardized feature vectors is as follows: ; in, To standardize the feature vector, For the first Mapping matrix corresponding to each mode For the first The initial feature vectors corresponding to each mode. It is the bias vector; The standardized feature vectors are concatenated, and the standardized feature vectors corresponding to each modality are combined in a preset order to obtain the fused feature vector; The fusion feature vector is reduced in dimension and mapped to a unified latent space to obtain the initial latent vector. The initial latent vector is composed of feature components corresponding to different modes.
4. The method for sharing micro-bearing production and testing data based on deep learning as described in claim 3, characterized in that, Step S2 specifically includes: Step S21: The latent vector of the sample to be tested, the production context parameters, and the preset defect-free baseline label are taken as joint inputs and fed into the conditional variational generation network. Reconstruction operations are performed within the unified latent space manifold to obtain the initial counterfactual latent vector. The processing logic is as follows: Let the latent vector of the sample to be tested be denoted as... The vector after mapping the production context parameters is denoted as The preset defect-free status label is denoted as Construct joint input vector It is represented as: ; By mapping the joint input vector to a probability distribution, a set of latent variable distribution parameters is obtained. Then, sampling operations are performed based on this set to obtain the sampled latent variables, calculated using the following formula: ; in, To sample latent variables, It is the mean vector. The standard deviation vector, Let be a random variable that follows a standard normal distribution; The sampled latent variables are input into the latent space to reconstruct the mapping function, thus obtaining the initial counterfactual latent vector; Step S22: Perform latent space constraint calculation on the initial counterfactual latent vector, introduce the production context parameters as constraints into the unified latent space manifold, and map and adjust the initial counterfactual latent vector in combination with the manifold distribution corresponding to the preset standard state label to obtain the constrained counterfactual latent vector. Step S23: Perform reverse mapping processing on the constrained counterfactual latent vectors according to the preset modality type, convert the latent space representation into multimodal detection data form, and obtain the counterfactual normal mirror; Step S24: Perform coordinate mapping calculations on the original detection data and the counterfactual normal mirror according to a unified time reference and a unified spatial reference, and establish a positional correspondence between the original detection data and the counterfactual normal mirror in a unified micro-bearing surface geometric coordinate system. Step S25: Under the unified geometric coordinate system of the micro-bearing surface, a differential homeomorphic space transformation operator is constructed based on the position correspondence and global mutual information value. A local topology preservation constraint term is added to the differential homeomorphic space transformation operator, and a continuous space mapping operation is performed on the original detection data to obtain the aligned original detection data.
5. The method for sharing micro-bearing production and testing data based on deep learning as described in claim 4, characterized in that, Step S23 specifically includes: The constrained counterfactual latent vectors are input into the latent space projection layer corresponding to each mode, and the latent space sub-vectors corresponding to each mode are extracted from the constrained counterfactual latent vectors through linear projection operation; The latent space sub-vectors of each mode are subjected to inverse mapping processing. The latent space sub-vectors of each mode are input into the corresponding inverse mapping function to obtain the reconstructed detection data of each mode. The calculation formula is as follows: ; in, For the first Reconstruction detection data corresponding to each modality For the first The inverse mapping function corresponding to each mode The first one extracted from the constrained counterfactual latent vector. The latent space sub-vectors corresponding to each mode; The reconstructed detection data of each modality is organized according to the preset modality type, and the reconstructed detection data of each modality is arranged according to the data structure consistent with the normalized original detection data to obtain a multimodal reconstruction data set; The multimodal reconstructed data set is combined according to a unified time reference and a unified spatial reference to generate a counterfactual normal image with the same structure as the original detection data.
6. The method for sharing micro-bearing production and testing data based on deep learning as described in claim 5, characterized in that, Step S25 specifically includes: Under the unified geometric coordinate system of the micro-bearing surface, the positional correspondence between the original detection data and the counterfactual normal mirror is obtained, and the global mutual information value between the original detection data and the counterfactual normal mirror is calculated to construct the corresponding position set; Based on positional correspondence and global mutual information, a differential homeomorphic spatial transformation operator is constructed in a unified geometric coordinate system of the micro-bearing surface. Local topology preservation constraints are added to the transformation objective function of the differential homeomorphic spatial transformation operator to perform continuous mapping calculations on the original spatial coordinates of the original detection data. The calculation formula is as follows: ; in, These are the mapped spatial coordinates. For differential homeomorphic space transformation operators, The original spatial coordinates of the original test data in the unified geometric coordinate system of the micro-bearing surface; The operator for the transformation of the differential homeomorphic space is subject to continuous differentiability constraints, and its Jacobian matrix is subject to nonsingularity constraints based on local topological preservation constraints. Its expression is as follows: ; in, For the differential homeomorphic space transformation operator in the original space coordinates Jacobian matrix at the location, To find the determinant of the Jacobian matrix; Based on the mapped spatial coordinates, values are calculated in the original detection data according to the spatial index relationship. The values of the original detection data at the original spatial coordinates are then mapped back to the mapped spatial coordinates to obtain the aligned original detection data.
7. The method for sharing micro-bearing production and testing data based on deep learning as described in claim 6, characterized in that, Step S4 specifically includes: Step S41: Analyze the defect deviation field, extract the correspondence between the deviation vector and the mapped spatial coordinates, and index the correspondence according to the spatial position index under the unified micro bearing surface geometric coordinate system to obtain the deviation field index data. Step S42: Extract the index identifier of the counterfactual normal image, and perform identifier mapping processing on the counterfactual normal image according to the unified time reference and unified spatial reference, and associate the index identifier of the counterfactual normal image with the spatial coordinate index in the defect deviation field. Step S43: Encode the production context parameters according to the preset parameter structure, and associate and calibrate the production context parameters according to the spatial location index, and associate the production context parameters with the spatial coordinate index in the defect deviation field. Step S44: Combine the off-field index data, the index identifier of the counterfactual normal mirror, and the production context parameters according to a unified data format, and organize the combined results in a structured manner to obtain shared data. Step S45: Sort the shared data according to the spatial location index, group it according to the production context parameters, and perform the distribution operation according to the preset distribution rules. The preset distribution rules include preset data splitting rules, preset field mapping rules, preset data frame generation rules, and preset sending order rules.
8. The method for sharing micro-bearing production and testing data based on deep learning as described in claim 7, characterized in that, Step S45 specifically includes: According to the preset data splitting rules, the shared data is parsed and split according to the preset field structure. The index identifiers of defect deviation fields and counterfactual normal mirrors and production context parameters are extracted from the shared data. According to the preset field mapping rules, the index identifiers of the defect deviation field and the counterfactual normal image and the production context parameters are mapped according to the preset communication protocol and written to the corresponding positions in the order of the fields to generate transmission data frames. According to the preset data frame generation rules, the transmitted data frames are arranged in a preset order, and each transmitted data frame is numbered based on the sequence index to form a data transmission sequence; According to the preset sending order rules, the data sending sequence is written into the network transmission interface, and each data frame is sent frame by frame in the order of the sequence index, while the corresponding sequence index is recorded. At the receiving end, the received transmission data frames are rearranged according to the sequence index, and the index identifiers of the defect deviation field and the counterfactual normal image and the production context parameters are parsed according to the preset field structure. The index identifier of the counterfactual normal image is input into the conditional variational generation network to extract the counterfactual normal image, and the extracted counterfactual normal image is superimposed with the defect deviation field to restore the corresponding data content.
9. A deep learning-based micro-bearing production and inspection data sharing system, applied in the deep learning-based micro-bearing production and inspection data sharing method as described in any one of claims 1-8, characterized in that, It includes a data encoding module, a mirror alignment module, an offset construction module, and a data distribution module; The data encoding module is used to acquire the original detection data and production context parameters of the miniature bearing, and to perform feature fusion and dimensionality reduction mapping using a multimodal joint semantic encoder to output the latent vector of the sample to be tested. The mirror alignment module is used to perform reconstruction operations in a unified latent space with production context parameters as constraints, through a conditional variational generation network, to obtain a counterfactual normal mirror. Based on a unified time reference and a unified spatial reference, the original detection data and the counterfactual normal mirror are subjected to co-position topology alignment processing to obtain aligned original detection data. The deviation construction module is used to calculate the deviation vector of the aligned original detection data relative to the counterfactual normal mirror, and organize the deviation vector according to the spatial position index to obtain the defect deviation field; The data distribution module is used to encapsulate the index identifiers and production context parameters of the defect deviation field and the counterfactual normal image to obtain shared data and perform distribution operations.
Citation Information
Patent Citations
Bearing fault detection method based on deep learning
CN118312874A
Bearing diagnosis method and system based on variable working condition multi-modal data fusion
CN118332488A
Image enhancement generation method based on face key point adaptive deformation
CN121962431A