Complex structural part scribed line segmentation method, device and system and storage medium

Through the training and denoising process of the deep diffusion model, the noise and lighting uneven problem of traditional methods in the segmentation of tile images of complex structural parts is solved, and high-precision tile segmentation is achieved, with strong adaptability.

CN120564003APending Publication Date: 2025-08-29BEIJING INFORMATION SCI & TECH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510936649.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

When traditional image segmentation methods process ticking images of complex structural parts, it is difficult to deal with noise, uneven lighting and complex backgrounds, resulting in inaccurate segmentation and difficult to achieve high-precision ticking recognition and segmentation.

Method used

The deep diffusion model is adopted, including the autoencoder module, the implicit spatial diffusion model and the conditional guidance module. Through the training and denoising process, the trunking segmentation results are generated with clear structure, accurate boundaries and consistent semantics.

Benefits of technology

It realizes high-precision segmentation of ticks for complex structural parts, improves the accuracy and consistency of segmentation, is highly adaptable, and can effectively deal with tick recognition in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564003A_ABST
    Figure CN120564003A_ABST
Patent Text Reader

Abstract

The invention discloses a complex structural member scribed line segmentation method, device and system, and a storage medium. The method comprises the following steps: S1, obtaining an original image; s2, training a depth diffusion model according to the original image; wherein the depth diffusion model comprises an auto-encoder module, an implicit spatial diffusion model and a condition guidance module; and S3, inputting a to-be-segmented image into the trained depth diffusion model to carry out complex structural part scribed line segmentation. By adopting the technical scheme of the invention, high-precision segmentation of the scribed line of the complex structural member is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and in particular relates to a complex structural part line segmentation method and device, system, and storage medium based on a diffusion segmentation model. Background Art

[0002] Among the many links in industrial production, the fine processing of products depends on the accurate analysis of industrial images. Among them, the high-precision segmentation and recognition of the markings on complex structural parts is an extremely critical link. Its segmentation accuracy directly affects the subsequent spatial position of the workpiece, thereby affecting the subsequent production process. From the perspective of image acquisition, the complex factors in the industrial environment bring many challenges to the acquisition of images of the markings on complex structural parts. Lighting conditions are often uneven. Strong light reflections or shadow areas can cause some areas of the markings on complex structural parts to be too bright or too dark, reducing the contrast of the markings on complex structural parts in the image and blurring the boundaries. The images collected at the industrial site are noisy, which further increases the difficulty of identifying and segmenting the markings on complex structural parts. In addition, complex structural parts made of different materials have different absorption and reflection characteristics of light, which also makes the markings on complex structural parts appear differently in the image, increasing the adaptability requirements of the segmentation algorithm.

[0003] Traditional image segmentation methods, such as threshold-based segmentation and edge detection algorithms, exhibit significant limitations when processing images of complex structural part inscriptions. Due to noise, uneven lighting, and complex background interference in complex structural part inscription images, accurate threshold selection is difficult, which can easily lead to inaccurate segmentation, broken or mis-segmented complex structural part inscriptions. Edge detection algorithms struggle to fully and accurately extract the edges of blurred or discontinuous complex structural part inscriptions. When faced with complex structural part inscriptions against complex backgrounds, traditional methods struggle to accurately distinguish the complex structural part inscriptions from the background. Summary of the Invention

[0004] Marking and alignment technology for complex structural parts is a key foundational process in mechanical manufacturing, used to accurately calibrate machining benchmarks. Its automation and intelligent development directly determine machining accuracy and production efficiency in the manufacturing field. The technical problem to be solved by this invention is to provide a method, device, system, and storage medium for marking and segmenting complex structural parts.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for dividing a complex structural part by engraving, comprising:

[0007] Step S1, obtaining an original image;

[0008] Step S2: training a deep diffusion model based on the original image; wherein the deep diffusion model includes: an autoencoder module, an implicit spatial diffusion model, and a conditional guidance module;

[0009] Step S3: input the image to be segmented into the trained depth diffusion model to perform line segmentation of complex structural parts.

[0010] Preferably, in step S2, in the first stage of training, the encoder-decoder structure of the autoencoder uses the true value labels of the inscriptions of complex structural parts to learn a high-quality mapping relationship from the image semantic label space to the latent variable space, and realizes the effective reconstruction of the image semantic label structure through the decoder; secondly, in the second stage of training, the implicit space diffusion model models the noise addition and denoising process of the diffusion model, and uses the denoising network of the implicit space diffusion model to learn and restore the original latent structure from the latent variables after noise perturbation under the conditional guidance of the original image semantic information through the conditional guidance module. The denoised latent variables are sent to the decoder part with frozen parameters in the autoencoder to generate a segmentation result with clear structure, accurate boundaries and consistent semantics.

[0011] Preferably, in step S3, under the guidance of the conditions of the image information to be segmented, starting from Gaussian noise, a denoising network of a trained implicit spatial diffusion model is used to obtain the segmentation results of the complex structural part lines after multiple denoising.

[0012] The present invention also provides a complex structural part scoring and segmenting device, comprising:

[0013] A first processing module is used to obtain an original image;

[0014] The second processing module trains a deep diffusion model based on the original image; wherein the deep diffusion model includes: an autoencoder module, an implicit spatial diffusion model and a conditional guidance module;

[0015] The third processing module is used to input the image to be segmented into the trained depth diffusion model to perform line segmentation of complex structural parts.

[0016] Preferably, the second processing module includes:

[0017] The first training unit is used to learn a high-quality mapping relationship from the image semantic label space to the latent variable space through the encoder-decoder structure of the autoencoder and the true value labels of the inscribed lines of complex structural parts, and to achieve effective reconstruction of the image semantic label structure through the decoder;

[0018] The second training unit is used to model the denoising and denoising processes of the diffusion model through the implicit spatial diffusion model. The denoising network of the implicit spatial diffusion model is used to learn and restore the original latent structure from the latent variables after noise perturbation under the conditional guidance of the semantic information of the original image through the conditional guidance module. The denoised latent variables are sent to the decoder part of the autoencoder with frozen parameters to generate a segmentation result with clear structure, accurate boundaries and consistent semantics.

[0019] Preferably, the third processing module is used to start from Gaussian noise under the guidance of the conditions of the image information to be segmented, use the denoising network of the trained implicit spatial diffusion model, and obtain the segmentation results of the complex structural parts after multiple denoising.

[0020] The present invention also provides a complex structural part line segmentation system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes the complex structural part line segmentation method when executed by the processor.

[0021] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is run, the method for dividing the complex structural part by engraving is executed.

[0022] The present invention first uses the encoder-decoder structure of the autoencoder in the first training phase to learn a high-quality mapping relationship from the image semantic label space to the latent variable space using the true value labels of the complex structural part's inscribed lines. The decoder then effectively reconstructs the image semantic label structure. Secondly, in the second training phase, the implicit spatial diffusion model models the denoising and de-noising processes of the diffusion model. The denoising network of the implicit spatial diffusion model, guided by the semantic information of the original image, learns to restore the original latent structure from the noise-perturbed latent variables. The denoised latent variables are fed into the decoder part of the autoencoder with frozen parameters, generating a segmentation result with clear structure, accurate boundaries, and consistent semantics. Finally, in the testing phase, guided by the information of the image to be segmented, starting with Gaussian noise, the trained implicit spatial diffusion model's denoising network is used to perform multiple denoising cycles to obtain the segmentation results of the complex structural part's inscribed lines. This invention can achieve high-precision segmentation of the complex structural part's inscribed lines. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0024] Figure 1 This is a flow chart of a method for dividing a complex structural part by marking lines according to an embodiment of the present invention;

[0025] Figure 2 Schematic diagram of the first training stage of the complex structural part automatic segmentation method based on the deep diffusion model;

[0026] Figure 3 Schematic diagram of the second training stage of the automatic segmentation method for complex structural parts based on the deep diffusion model. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0028] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] Example 1:

[0030] like Figure 1 、 2 As shown in FIG. 3 , an embodiment of the present invention provides a method for dividing a complex structural part by marking lines, comprising:

[0031] Step S1, obtaining an original image;

[0032] Step S2: training a deep diffusion model based on the original image; wherein the deep diffusion model includes: an autoencoder module, an implicit spatial diffusion model, and a conditional guidance module;

[0033] Step S3: input the image to be segmented into the trained depth diffusion model to perform line segmentation of complex structural parts.

[0034] As an implementation method of an embodiment of the present invention, in step S2, the first stage of training: semantic mapping and reconstruction of true value labels. The goal of the first stage is to learn a high-quality mapping relationship from the image semantic label space to the latent variable space, and to achieve effective reconstruction of the image semantic label structure through the decoder. Specifically, the true value label of the image is first input into the encoder part of the autoencoder module, and its structure and semantic features are extracted through multi-layer convolution and downsampling operations, and compressed into the latent space to obtain the encoded latent variable representation. Subsequently, the latent variables are sampled using the reparameterization technique, and the sampling results are restored to a reconstructed graph consistent with the original label semantic structure through the decoder. Among them, the encoding process can be expressed by the following formula:

[0035] z0=Encode(GT) (1)

[0036] Here, GT is the true value label of the input image, Encode() is the encoder process, and z0 is the encoded feature representation. In order to introduce generative modeling capabilities, it is usually assumed that the spatial feature representation z0 obtained after encoding obeys a certain probability distribution, and the latent space variable z0 obeys a mean of μ and a variance of σ. 2 Gaussian distribution, that is, z0∈N(μ,σ 2 ). It is sampled through the reparameterization technique to obtain the latent variable z. The core idea of ​​the reparameterization technique is to reconstruct the random sampling process into a differentiable deterministic transformation, so that the model can achieve end-to-end backpropagation during training. Specifically, by converting the operation of sampling from an arbitrary Gaussian distribution into sampling from a standard normal distribution, and introducing learnable mean and variance parameters for linear transformation, the problem of blocking gradient propagation caused by the sampling operation is effectively solved, and it is also ensured that the gradient information can be smoothly transmitted to the generative network, thereby improving the trainability and expressiveness of the generative model. The mathematical expression of the reparameterization technique is as follows:

[0037] z=μ+σ·ε,ε∈N(0,I) (2)

[0038] Where μ and σ are the mean and standard deviation of the encoder output, respectively, and ε is a random variable under the standard normal distribution. Subsequently, the sampled latent variable z is input to the decoder, which consists of multiple layers of upsampling and deconvolution operations, gradually restoring the low-dimensional representation to the original label image size and generating a predicted image. The reconstruction process is described as follows:

[0039] Y=Deconde(z) (3)Under the joint guidance of KL divergence loss, cross entropy loss and adversarial loss, this stage aims to reconstruct images whose semantic structure and spatial distribution are highly consistent with the true labels, providing structured and expressive initial latent variables for subsequent diffusion modeling.

[0040] As one implementation of an embodiment of the present invention, in step S2, the second phase of training involves diffusion model generation and conditional guidance. In this phase, the diffusion process takes over the semantic generation task. First, the ground-truth label of the labeled image is input into the encoder portion of a trained and frozen autoencoder module to obtain a latent variable representation. Subsequently, Gaussian noise is gradually added to the latent space to introduce random perturbations and construct diffusion trajectories. The denoising process is performed by a decoupled diffusion model, whose goal is to gradually recover the original latent structure from the noise-perturbed latent variables. Simultaneously, the original image is input into a conditional encoding network, which extracts high-level semantic and structural information from the image to serve as conditional guidance for the diffusion process. During this process, image conditional information is embedded into multiple layers of the denoising U-Net. Semantic fusion and guidance are achieved through the attention mechanism, ensuring that the conditional information is fully utilized during the denoising process, enhancing the conditional controllability of the diffusion process, and thus strengthening the semantic consistency and boundary accuracy of the final segmentation result. The denoising U-Net, in turn, predicts the image and noise components of the noisy latent variables separately by combining time step and image conditional information, and generates latent variables containing the conditional information from these components. The reverse process is implemented by a set of denoising U-Net structures with skip connections, which has powerful cross-scale modeling capabilities. At each time step t, the denoising U-Net network receives the input noisy latent variable z n Combined with the time condition t, the noise and corresponding image components that should be removed at the current moment are predicted, and the joint modeling of noise removal and semantic reconstruction is achieved. The formula is expressed as follows:

[0041] C pred ,n pred =UNet(z n ,t) (4)

[0042] Among them, C pred is the image estimation component, representing the semantic structure restored by the model at the current time step; n pred is the noise estimate component, representing the noise estimate that should be removed at the current time step. Together, they simulate the gradual restoration mechanism in the diffusion trajectory. The reverse process of the decoupled diffusion model guides the model to gradually regress from random perturbations to the original semantic distribution, achieving layer-by-layer perception and reconstruction of the complex structure of remote sensing images.

[0043] Finally, the denoised latent variables are fed into the decoder part of the autoencoder with frozen parameters to generate a segmentation map with clear structure, accurate boundaries and consistent semantics.

[0044] As an implementation method of an embodiment of the present invention, in step S3, under the guidance of the conditions of the image information to be segmented, starting from Gaussian noise, using the denoising network of the trained implicit space diffusion model, after multiple denoising (the number of denoising steps of the diffusion model determined in the training process), the final generated latent variable needs to be upsampled by the decoder of the autoencoder to restore it from the compressed latent space to the original image resolution, thereby obtaining the segmentation result of the complex structural part engraving. Among them, the diffusion path of the denoising U-Net contains two parallel branches, which respectively predict the noise component and image component in the image, and based on the principle of the decoupled diffusion model, by x t Generate an image x with a single noise in the region t-1 The decoder usually includes multiple upsampling layers and convolutional layers to gradually expand the spatial size and refine the details.

[0045] Example 2:

[0046] An embodiment of the present invention further provides a complex structural part scoring and segmenting device, comprising:

[0047] A first processing module is used to obtain an original image;

[0048] The second processing module trains a deep diffusion model based on the original image; wherein the deep diffusion model includes: an autoencoder module, an implicit spatial diffusion model and a conditional guidance module;

[0049] The third processing module is used to input the image to be segmented into the trained depth diffusion model to perform complex structural part line segmentation.

[0050] As an implementation manner of an embodiment of the present invention, the second processing module includes:

[0051] The first training unit is used to learn a high-quality mapping relationship from the image semantic label space to the latent variable space through the encoder-decoder structure of the autoencoder and the true value labels of the inscribed lines of complex structural parts, and to achieve effective reconstruction of the image semantic label structure through the decoder;

[0052] The second training unit is used to model the denoising and denoising processes of the diffusion model through the implicit spatial diffusion model. The denoising network of the implicit spatial diffusion model is used to learn and restore the original latent structure from the latent variables after noise perturbation under the conditional guidance of the semantic information of the original image through the conditional guidance module. The denoised latent variables are sent to the decoder part of the autoencoder with frozen parameters to generate a segmentation result with clear structure, accurate boundaries and consistent semantics.

[0053] As an implementation method of an embodiment of the present invention, the third processing module is used to start from Gaussian noise under the guidance of the conditions of the image information to be segmented, use the denoising network of the trained implicit spatial diffusion model, and obtain the segmentation results of the complex structural parts after multiple denoising.

[0054] Example 3:

[0055] An embodiment of the present invention further provides a complex structural part line segmentation system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a complex structural part line segmentation method when executed by the processor.

[0056] Example 4:

[0057] An embodiment of the present invention further provides a storage medium having a computer program stored thereon. When the computer program is run, the method for segmenting a complex structural part by marking lines is executed.

[0058] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for dividing complex structural parts by marking lines, characterized in that: include: Step S1, obtaining an original image; Step S2: training a deep diffusion model based on the original image; wherein the deep diffusion model includes: an autoencoder module, an implicit spatial diffusion model, and a conditional guidance module; Step S3: input the image to be segmented into the trained depth diffusion model to perform line segmentation of complex structural parts.

2. The complex structural part scoring method according to claim 1, wherein: In step S2, in the first stage of training, the encoder-decoder structure of the autoencoder uses the true value labels of the inscriptions of complex structural parts to learn a high-quality mapping relationship from the image semantic label space to the latent variable space, and realizes the effective reconstruction of the image semantic label structure through the decoder; secondly, in the second stage of training, the implicit space diffusion model models the denoising and denoising process of the diffusion model, and uses the denoising network of the implicit space diffusion model to learn and restore the original latent structure from the latent variables after noise perturbation under the conditional guidance of the original image semantic information through the conditional guidance module. The denoised latent variables are sent to the decoder part of the autoencoder with frozen parameters to generate a segmentation result with clear structure, accurate boundaries and consistent semantics.

3. The complex structural part scoring method according to claim 1, wherein: In step S3, under the guidance of the conditions of the image information to be segmented, starting from Gaussian noise, the denoising network of the trained implicit spatial diffusion model is used to obtain the segmentation results of the complex structural part lines after multiple denoising.

4. A complex structural part scoring and segmenting device, characterized in that: include: A first processing module is used to obtain an original image; The second processing module trains a deep diffusion model based on the original image; wherein the deep diffusion model includes: an autoencoder module, an implicit spatial diffusion model and a conditional guidance module; The third processing module is used to input the image to be segmented into the trained depth diffusion model to perform line segmentation of complex structural parts.

5. The complex structural part scoring and segmenting device according to claim 4, characterized in that: The second processing module includes: The first training unit is used to learn a high-quality mapping relationship from the image semantic label space to the latent variable space through the encoder-decoder structure of the autoencoder and the true value labels of the inscribed lines of complex structural parts, and to achieve effective reconstruction of the image semantic label structure through the decoder; The second training unit is used to model the denoising and denoising processes of the diffusion model through the implicit spatial diffusion model. The denoising network of the implicit spatial diffusion model is used to learn and restore the original latent structure from the latent variables after noise perturbation under the conditional guidance of the semantic information of the original image through the conditional guidance module. The denoised latent variables are sent to the decoder part of the autoencoder with frozen parameters to generate a segmentation result with clear structure, accurate boundaries and consistent semantics.

6. The complex structural part scoring and dividing device according to claim 5, characterized in that: The third processing module is used to obtain the segmentation results of the complex structural parts' inscribed lines by using the denoising network of the trained implicit spatial diffusion model starting from Gaussian noise under the guidance of the image information to be segmented.

7. A complex structural part scoring and segmentation system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the method for dividing a complex structural part by marking lines according to any one of claims 1 to 3 is executed.

8. A storage medium, characterized in that: The storage medium stores a computer program, which, when running, executes the complex structural component line segmentation method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Decorative surface material texture image generation method and system and medium

    CN118314243A

  • High-resolution remote sensing image semantic segmentation method based on diffusion model

    CN120673054A

  • Image segmentation method and system based on corrected diffusion model

    JP7663284B1