Automatic sapphire seed guiding system and method based on multi-stage visual algorithm model
The sapphire automatic crystal pulling system, based on a multi-stage visual algorithm model, solves the problems of low stability and automation in the sapphire crystal pulling process. It achieves phased differentiated control and automated closed-loop operation, thereby improving the stability and accuracy of crystal pulling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TDG YINXIA NEW MATERIAL CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-31
AI Technical Summary
The existing sapphire crystal pulling process mainly relies on manual operation, which has problems such as poor pulling stability, inability to adapt to the differentiated needs of multiple stages, and low level of automation, making it difficult to achieve consistent pulling and precise control.
An automated sapphire crystal-leading system based on a multi-stage visual algorithm model is adopted, including a crystal-leading visual image acquisition device, a detection and image processing system, and a three-stage visual judgment algorithm model. Through convolutional neural networks, dual-stream feature fusion networks, and an attention-based encoder-decoder structure, it achieves phased differentiated control and automated closed-loop operation.
It has improved the stability and consistency of the sapphire crystal pulling process, has the ability to process the entire process in stages with differentiation, completes automated closed-loop control, replaces manual operation, and improves the accuracy and automation level of crystal pulling.
Smart Images

Figure CN122484902A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sapphire crystal growth technology, and in particular to an automatic sapphire crystal pulling system and method based on a multi-stage visual algorithm model. Background Technology
[0002] Sapphire crystals possess excellent optical, thermal, and mechanical properties, making them promising candidates for applications in aerospace, defense, semiconductor lighting, and communication technologies. The growth process of sapphire crystals places extremely high technical demands on the seeding operation, and the quality of the seeding is the most critical factor determining the overall crystal quality.
[0003] Currently, the sapphire crystal seeding process in the industry mainly relies on manual operation. Operators observe the dissolution state of the seed crystal inside the furnace through an observation window and manually raise and lower the seed crystal based on experience and furnace temperature changes to complete the seeding operation. However, this manual seeding method has the following problems: First, the seeding stability is poor, and it relies too much on the operator's experience, making it difficult to guarantee the consistency of the seeding; second, it cannot adapt to the differentiated needs of the multi-stage seeding process. The requirements for identification accuracy and control vary at each stage of the seeding process, and the existing method, which uses uniform manual observation or simple visual inspection, cannot achieve precise control throughout the entire process; third, the level of automation is low, lacking a complete closed-loop system from image acquisition and status recognition to control execution. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides an automated sapphire crystal pulling system and method based on a multi-stage visual algorithm model. This system is adaptable to the entire crystal pulling process, possesses the capability for differentiated processing at different stages, and can automatically complete image acquisition, status recognition, judgment and decision-making, and execution control.
[0005] In a first aspect, embodiments of this application provide an automated sapphire crystal pulling system based on a multi-stage visual algorithm model, comprising: A crystal-leading visual image acquisition device is used to acquire real-time images of the crystal-leading process in a sapphire crystal growth furnace to obtain real-time crystal-leading images, and output the real-time crystal-leading images; The detection and image processing system is connected to the seed crystal visual image acquisition device. The detection and image processing system is used to receive the real-time image of the seed crystal, perform noise reduction processing on the real-time image of the seed crystal, identify the edge contour and bright and dark spots of the seed crystal from the noise-reduced image to obtain the seed crystal state feature information, and output the seed crystal state feature information. A three-stage visual judgment algorithm model is connected to the detection and image processing system. The three-stage visual judgment algorithm model includes a temperature testing stage model, a welding stage model, and a necking stage model. The three-stage visual judgment algorithm model is used to receive the seed crystal state feature information and send the seed crystal state feature information to one of the temperature testing stage model, the welding stage model, or the necking stage model according to the current crystal pulling stage. The temperature testing stage model is used to receive the seed crystal state feature information and use a convolutional neural network with residual connections to extract the seed crystal edge contour and determine the melting speed, thereby obtaining the temperature testing stage judgment result; the welding stage model is used to receive the seed crystal state feature information and use a dual-stream feature fusion network to monitor the changes and proportions of the melting area, thereby obtaining the welding stage judgment result; the necking stage model is used to receive the seed crystal state feature information and use an attention-based encoder-decoder structure to eliminate gas interference, thereby obtaining the necking stage judgment result. The three-stage visual judgment algorithm model is also used to generate operation instructions based on the judgment results of the temperature test stage, the judgment results of the welding stage, or the judgment results of the necking stage, and output the operation instructions. An execution control unit is connected to the three-stage visual judgment algorithm model. The execution control unit is used to receive the operation instructions and control the temperature of the sapphire crystal growth furnace and the seed crystal pulling speed according to the operation instructions.
[0006] According to some embodiments of the first aspect of this application, in the temperature testing phase model: The convolutional neural network includes multiple convolutional layers and multiple pooling layers, which are connected alternately in sequence. The convolutional kernels of the convolutional layers include two sizes: 3×3 and 5×5. The residual connection is set in the middle and later stages of the convolutional neural network to directly superimpose the output of the previous layer onto the input of the next layer. The model for the temperature testing phase uses a weighted cross-entropy loss function as its optimization objective, and the weights of the weighted cross-entropy loss function are dynamically adjusted according to the proportion of positive samples.
[0007] According to some embodiments of the first aspect of this application, in the welding stage model: The dual-stream feature fusion network includes a first sub-network and a second sub-network arranged in parallel, and a feature fusion layer connected to the output ends of the first sub-network and the second sub-network, respectively. The input end of the first sub-network is used to receive first-view image data, and the input end of the second sub-network is used to receive second-view image data. The first-view image data and the second-view image data are real-time images of the crystal pulling process captured by the crystal pulling visual image acquisition device from the observation window of the sapphire crystal growth furnace from two different angles. The first sub-network is used to extract high-level semantic features from the first viewpoint image data to obtain a first feature map and output the first feature map to the feature fusion layer; the second sub-network is used to extract high-level semantic features from the second viewpoint image data to obtain a second feature map and output the second feature map to the feature fusion layer. The feature fusion layer is used to merge the first feature map and the second feature map to obtain a fused feature map; The fusion stage model employs a multi-task learning framework, which simultaneously performs semantic segmentation of the fusion region, regression prediction of the fusion speed, and classification of the fusion ratio based on the fused feature map. The overall loss function of the multi-task learning framework is a weighted sum of the Dice coefficient loss function, the mean squared error loss function, and the binary cross-entropy loss function.
[0008] According to some embodiments of the first aspect of this application, in the necking stage model: The attention-based encoder-decoder structure includes an encoder and a decoder, with the output of the encoder connected to the input of the decoder. The encoder uses a VGG-16 network to extract multi-level features from the input image; The decoder includes a spatial attention module and a channel attention module. The spatial attention module is used to enhance the spatial location information of the feature map, and the channel attention module is used to enhance the channel dimension information of the feature map. The encoder-decoder structure also embeds multiple residual dense connection blocks, which directly transmit the features of the corresponding layer in the encoder to the corresponding layer in the decoder through skip connections; The loss function of the necking stage model is a weighted sum of the structural similarity loss function and the L1 norm loss function.
[0009] According to some embodiments of the first aspect of this application, the temperature testing stage model is trained from a dataset containing 10,000 labeled images, which cover the melting state of the seed crystal under different temperature conditions.
[0010] According to some embodiments of the first aspect of this application, the welding stage model is trained from a dataset containing 15,000 multi-view image pairs, each of which is labeled with a mask of the melting region, a numerical value of the melting speed, and a category label of the melting ratio.
[0011] According to some embodiments of the first aspect of this application, the necking stage model is trained from a dataset containing 8,000 pairs of images before and after gas disturbance.
[0012] According to some embodiments of the first aspect of this application, the crystal-guided visual image acquisition device includes an industrial high-definition camera, a reflective lens assembly, a camera support rod, a fixed base, and an adjustment mechanism. The adjustment mechanism includes a camera mounting base, a camera height adjustment base, a camera angle adjustment base, and a reflective lens moving base. The bottom of the camera support rod is rotatably mounted to one end of the fixed base via the camera angle adjustment base. The camera height adjustment base is slidably mounted on the camera support rod. The industrial high-definition camera is mounted on the camera height adjustment base via the camera mounting base. The reflective lens moving base is disposed at the other end of the fixed base, and the reflective lens assembly is mounted on the reflective lens moving base.
[0013] According to some embodiments of the first aspect of this application, the detection and image processing system includes: The image denoising module is used to receive the real-time image of the crystal, perform physical denoising on the real-time image of the crystal through an optical filter, and perform software denoising on the physically denoised image using a high-temperature radiation denoising algorithm and a dynamic blur compensation algorithm to obtain a denoised image. A dynamic detection module is connected to the image denoising module. The dynamic detection module is used to receive the denoised image, identify and depict the edge contour and bright and dark spots of the seed crystal from the denoised image, obtain the seed crystal state feature information, and output the seed crystal state feature information.
[0014] In a second aspect, embodiments of this application provide an automated sapphire crystal pulling method based on a multi-stage visual algorithm model, applied to the automated sapphire crystal pulling system based on a multi-stage visual algorithm model provided in the first aspect embodiment. The method includes: Real-time images of the crystal growth process inside the sapphire crystal growth furnace were acquired to obtain real-time images of the crystal growth process. The real-time image of the seed crystal is denoised, and the edge contour and bright and dark spots of the seed crystal are identified from the denoised image to obtain the seed crystal state feature information. Based on the current seed crystal stage, the seed crystal state characteristic information is sent to the corresponding visual judgment algorithm model: During the temperature testing phase, the seed crystal state feature information is sent to the temperature testing phase model. The temperature testing phase model uses a convolutional neural network with residual connections to extract the seed crystal edge contour and determine the melting rate, thereby obtaining the temperature testing phase judgment result. During the welding stage, the seed crystal state characteristic information is sent to the welding stage model, which uses a dual-flow feature fusion network to monitor the changes and proportions of the melting area and obtain the welding stage judgment result. During the necking stage, the seed crystal state characteristic information is sent to the necking stage model, which uses an attention-based encoder-decoder structure to eliminate gas interference and obtain the necking stage judgment result. An operation command is generated based on the judgment results of the temperature test stage, the welding stage, or the necking stage. Adjust the temperature of the sapphire crystal growth furnace and the seed crystal pulling speed according to the operation instructions until the crystal pulling is completed.
[0015] The beneficial effects of this application are reflected in: 1. By setting up a seed crystal visual image acquisition device and a detection and image processing system, the real-time automatic image acquisition, noise reduction processing and seed crystal state characteristic information of the seed crystal process are realized, which replaces manual observation and experience judgment, effectively avoids the inconsistency of seed crystals caused by human factors, and significantly improves the stability of seed crystals. 2. A three-stage visual judgment algorithm model was established, which includes a temperature testing stage model, a welding stage model, and a necking stage model. Based on the current crystal-leading stage, the seed crystal state feature information is sent to the corresponding model for processing. Specifically: the temperature testing stage model uses a convolutional neural network with residual connections to extract the seed crystal edge contour and determine the melting speed; the welding stage model uses a dual-stream feature fusion network to monitor the changes and proportions of the melting area; and the necking stage model uses an attention-based encoder-decoder structure to eliminate gas interference. This achieves phased, differentiated, and precise control over the entire crystal-leading process. 3. The crystal-leading visual image acquisition device, detection and image processing system, three-stage visual judgment algorithm model and execution control unit are connected in sequence to form a complete automated closed loop: image acquisition → processing and recognition → staged model judgment → generation of operation instructions → execution control unit to adjust temperature and lifting speed. The system can complete the entire process without manual intervention, realizing automated closed-loop control of sapphire crystal leading.
[0016] This application, through this configuration, can adapt to the entire crystal pulling process, has the ability to process in stages with differentiated characteristics, and can automatically complete image acquisition, status recognition, judgment and decision-making, and execution control. Attached Figure Description
[0017] Figure 1 This is a connection diagram of the sapphire automatic crystal pulling system based on a multi-stage visual algorithm model provided in the first aspect of this application; Figure 2 This is a schematic diagram of the structure of the crystal-guided visual image acquisition device provided in the first aspect embodiment of this application; Figure 3 This is a side view schematic diagram of the crystal-guided visual image acquisition device provided in the first aspect embodiment of this application; Figure 4 This is a schematic flowchart of the sapphire automatic crystal pulling method based on a multi-stage visual algorithm model provided in the second aspect embodiment of this application.
[0018] Figure label: 1. Camera support rod; 2. Camera height adjustment base; 3. Camera mounting base; 4. Camera angle adjustment base; 5. Angle adjustment component; 6. Industrial HD camera; 7. Support base; 8. Observation window fixing base; 9. Observation window end cap; 10. Sapphire crystal growth furnace observation window; 11. Reflective lens moving base; 12. Reflective rotating base; 13. Lens clamp; 14. Reflective lens. Detailed Implementation
[0019] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0020] The following description, in conjunction with the accompanying drawings, details a sapphire automatic crystal pulling system and method based on a multi-stage visual algorithm model provided in this application, through specific embodiments and application scenarios.
[0021] Reference Figure 1The first aspect of this application provides an automatic sapphire seeding system based on a multi-stage visual algorithm model, including a seeding visual image acquisition device, a detection and image processing system, a three-stage visual judgment algorithm model, and an execution control unit. The seeding visual image acquisition device acquires real-time images of the seeding process within the sapphire crystal growth furnace to obtain real-time seeding images and outputs them. The detection and image processing system is connected to the seeding visual image acquisition device. This system receives the real-time seeding images, performs noise reduction processing on them, and identifies the edge contours and bright / dark spots of the seed crystal from the noise-reduced images to obtain seed crystal state feature information, which is then output. The three-stage visual judgment algorithm model is connected to the detection and image processing system. This model includes a temperature testing stage model, a welding stage model, and a necking stage model. The three-stage visual judgment algorithm model receives the seed crystal state feature information and, based on the current seeding stage... The process involves sending seed crystal state feature information to one of three models: the temperature testing stage model, the welding stage model, or the necking stage model. The temperature testing stage model receives the seed crystal state feature information and uses a convolutional neural network with residual connections to extract the seed crystal edge contour and determine the melting rate, thus obtaining the temperature testing stage judgment result. The welding stage model receives the seed crystal state feature information and uses a dual-stream feature fusion network to monitor changes and proportions in the melting region, thus obtaining the welding stage judgment result. The necking stage model receives the seed crystal state feature information and uses an attention-based encoder-decoder structure to eliminate gas interference, thus obtaining the necking stage judgment result. The three-stage visual judgment algorithm model also generates operation instructions based on the temperature testing stage judgment result, the welding stage judgment result, or the necking stage judgment result, and outputs the operation instructions. The execution control unit is connected to the three-stage visual judgment algorithm model and receives the operation instructions, controlling the temperature of the sapphire crystal growth furnace and the seed crystal pulling speed according to the operation instructions.
[0022] It should be noted that the specific process for each stage is as follows: Temperature Testing Stage: By adjusting the heating power of the sapphire crystal growth furnace, the furnace temperature is gradually raised to near the melting point of the sapphire crystal. Simultaneously, the seed crystal rod is raised and lowered, and the seed crystal, fixed at the lower end of the rod, is slowly lowered until its lower end gently contacts the free surface of the sapphire melt. The goal of the temperature testing stage is to find suitable temperature conditions to prepare for subsequent welding operations. Welding Stage: After the temperature testing stage, the heating power of the sapphire crystal growth furnace is precisely adjusted to control the temperature at the contact surface between the lower end of the seed crystal and the melt to be slightly higher than the melting point of the sapphire crystal. This causes a small amount of melting on the surface of the seed crystal end. This process removes any contamination layer that may exist on the seed crystal surface, resulting in an atomically clean crystal surface. Subsequently, the heating power is stabilized to maintain the melt temperature at its melting point, ensuring sufficient wetting of the melt and the clean seed crystal surface, establishing favorable initial conditions for crystal growth. Necking stage: After the seed crystal and melt are successfully wetted, the seed crystal rod is pulled upwards at a very slow speed. At the same time, the heating power of the sapphire crystal growth furnace is finely adjusted so that the interface temperature between the seed crystal and the melt is slightly lower than the melting point of the sapphire crystal. At this time, the melt begins to grow epitaxially from the end of the seed crystal, forming a necking region with a small diameter. The necking stage can effectively eliminate dislocation defects that may exist in the seed crystal, providing a high-quality crystal seed for subsequent constant-diameter crystal growth.
[0023] Understandably, in the test phase model: the convolutional neural network includes multiple convolutional layers and multiple pooling layers, which are connected alternately in sequence. The convolutional kernels of the convolutional layers include two sizes: 3×3 and 5×5. The residual connection is set in the middle and later stages of the convolutional neural network to directly superimpose the output of the previous layer onto the input of the next layer. The test phase model uses the weighted cross-entropy loss function as the optimization objective, and the weighted cross-entropy loss function dynamically adjusts the weights according to the proportion of positive samples.
[0024] It should be noted that the convolutional kernels of the convolutional layers include two sizes: 3×3 and 5×5, both with a stride of 1, and "same" padding is used to maintain the feature map size. The first two convolutional layers of the convolutional neural network are each followed by a ReLU activation function to introduce non-linearity. A max-pooling layer with a 2×2 kernel and a stride of 2 is placed in the middle to later stages of the convolutional neural network to reduce the feature map resolution. Residual connections are placed in the middle to later stages of the convolutional neural network to directly superimpose the outputs of the previous layer onto the inputs of the next layer, thus avoiding gradient vanishing and improving training stability. During the trial phase, the model uses a weighted cross-entropy loss function as the optimization objective. The weighted cross-entropy loss function dynamically adjusts the weights according to the proportion of positive samples. The expression for the weighted cross-entropy loss function is as follows: ], in, This represents the value of the weighted cross-entropy loss function. Indicates the first A real label of 1 pixel. This indicates the model's prediction during the temperature testing phase. The probability value that a pixel belongs to the edge contour of the seed crystal. This represents the total number of pixels in the image; This represents the positive sample weighting factor, which is dynamically adjusted based on the proportion of seed crystal edge pixels in the image to alleviate the imbalance between positive and negative sample classes. By minimizing the weighted cross-entropy loss function, the model can progressively optimize the network parameters during the temperature testing phase, making the predicted seed crystal edge contour as close as possible to the true annotation.
[0025] In this embodiment, the temperature testing stage model adopts a deep learning structure based on convolutional neural networks to extract the edge contour of the seed crystal from the seed crystal state feature information and determine the melting rate. Specifically, the convolutional neural network consists of multiple convolutional layers and multiple pooling layers. Each convolutional layer is followed by a ReLU activation function to introduce non-linearity. The network first performs preliminary feature extraction on the image corresponding to the input seed crystal state feature information through two consecutive convolutional layers. The kernel size of the first convolutional layer is 3×3, and the kernel size of the second convolutional layer is 5×5, with a stride of 1. "Same" padding is used to keep the feature map size unchanged. Subsequently, a max pooling layer is set with a kernel size of 2×2 and a stride of 2 to reduce the feature map resolution and reduce the computational cost. In the middle and later stages of the convolutional neural network, residual connections are introduced to directly superimpose the output feature map of the previous layer onto the input feature map of the subsequent layer to avoid the gradient vanishing problem and improve the training stability and convergence speed of the model. At the end of the network, a global average pooling layer is set to compress the feature map into a one-dimensional vector. This one-dimensional vector is connected to a fully connected layer, which outputs the prediction result of the seed crystal edge contour and the judgment result of the melting speed.
[0026] Understandably, in the fusion stage model: the dual-stream feature fusion network includes a first sub-network and a second sub-network set in parallel, and feature fusion layers connected to the outputs of the first and second sub-networks respectively; the input of the first sub-network is used to receive first-view image data, and the input of the second sub-network is used to receive second-view image data. The first-view image data and the second-view image data are real-time images of the crystal pulling process captured from the observation window of the sapphire crystal growth furnace by the crystal pulling visual image acquisition device, taken from two different angles; the first sub-network is used to extract high-level semantic features from the first-view image data to obtain... The first feature map is output to the feature fusion layer. The second sub-network is used to extract high-level semantic features from the second-view image data to obtain the second feature map, which is then output to the feature fusion layer. The feature fusion layer merges the first and second feature maps to obtain a fused feature map. The fusion stage model adopts a multi-task learning framework to simultaneously perform semantic segmentation of the fused region, regression prediction of the fused speed, and classification of the fused ratio based on the fused feature map. The overall loss function of the multi-task learning framework is a weighted sum of the Dice coefficient loss function, the mean squared error loss function, and the binary cross-entropy loss function.
[0027] In this embodiment, both the first and second sub-networks adopt the ResNet-50 architecture to extract high-level semantic features from the first-view image data and the second-view image data, respectively, to obtain a first feature map and a second feature map. The first and second feature maps are then output to the feature fusion layer. The feature fusion layer merges the first and second feature maps by adding them element-wise to obtain a fused feature map. Subsequently, the fused feature map is input into a decoder consisting of multiple convolutional layers and fully connected layers to generate the final prediction result.
[0028] It should be noted that the feature fusion layer merges the first and second feature maps using element-wise addition to obtain a fused feature map; semantic segmentation uses the Dice coefficient loss function, regression prediction uses the mean squared error loss function, and classification uses the binary cross-entropy loss function; the overall loss function of the multi-task learning framework is a weighted sum of the Dice coefficient loss function, the mean squared error loss function, and the binary cross-entropy loss function. The expression for the overall loss function of the multi-task learning framework is as follows: + + , in, This represents the overall loss function value of the multi-task learning framework; , and This represents the adjustable weight coefficients for each subtask, used to balance the optimization objectives of the three tasks: semantic segmentation, regression prediction, and classification. This represents the Dice coefficient loss function. This represents the mean squared error loss function. This represents the binary cross-entropy loss function. Specifically, the expression for the Dice coefficient loss function is as follows: , This represents the set of pixels in the fusion region predicted by the model during the welding stage. This represents the set of pixels in the truly labeled melting region; the expression for the mean squared error loss function is as follows: , This represents the actual melting rate value. This represents the melting rate value predicted by the model during the welding stage. This represents the number of samples; the expression for the binary cross-entropy loss function is as follows: ],in, Indicates the number of samples. Indicates the first The true class label of each sample The model predicts the first stage of the welding process. The probability values of 3 samples belonging to the positive class.
[0029] Understandably, in the necking-stage model: the attention-based encoder-decoder structure includes an encoder and a decoder, with the encoder's output connected to the decoder's input; the encoder uses a VGG-16 network to extract multi-level features from the input image; the decoder includes a spatial attention module and a channel attention module, with the spatial attention module enhancing the spatial location information of the feature map and the channel attention module enhancing the channel dimension information of the feature map; the encoder-decoder structure also embeds multiple residual dense connection blocks, which directly transfer features from the corresponding layer in the encoder to the corresponding layer in the decoder through skip connections; the loss function of the necking-stage model is a weighted sum of the structural similarity loss function and the L1 norm loss function.
[0030] It should be noted that the encoder uses a pre-trained VGG-16 network as its backbone structure to progressively extract low- to high-level multi-layer features from the input image; the spatial attention module generates spatial weight maps to enhance information about key spatial locations in the feature maps, and the channel attention module generates channel weight vectors to enhance information about important channels in the feature maps; in the multiple residual dense connection blocks embedded in the encoder-decoder structure, each residual dense connection block directly transmits features from the corresponding layer in the encoder to the corresponding layer in the decoder through skip connections to recover details obscured by gas interference; the loss function of the necking stage model is a weighted sum of the structural similarity loss function and the L1 norm loss function, and the expression for the loss function of the necking stage model is as follows: + , in, This represents the overall loss function value during the necking phase, used to guide the model in removing gas interference and restoring a clear image during the necking phase. The structural similarity loss function is calculated based on the SSIM index and is used to measure the similarity between the de-interference image output by the model and the reference clear image in three dimensions: brightness, contrast, and structure. Denotes the L1 norm loss function. and The weight coefficients of the structural similarity loss function are represented. The weighting coefficients of the L1 norm loss function are represented. Let represent the loss function of the necking-down stage model. The expression for the structural similarity loss function is as follows: ,in, This represents the average pixel value of the reference image (the clear image), reflecting the image brightness; This represents the pixel mean of the predicted image (the de-interference image output by the model during the necking stage). This represents the pixel variance of the reference image. This represents the pixel variance of the predicted image. Covariance between the reference image and the predicted image and This represents the stability constant, used to avoid numerical instability caused by a zero denominator; the expression for the L1 norm loss function is as follows: ,in, Indicates the total number of pixels in the image. The first reference image is the clearest image. pixel value, The image predicted by the model during the necking stage is the first... Each pixel value.
[0031] Understandably, the model for the temperature testing phase was trained using a dataset containing 10,000 labeled images, which cover the melting state of the seed crystals under different temperature conditions.
[0032] It should be noted that the model for the temperature testing phase was trained using a dataset containing 10,000 labeled images. This dataset originated from high-definition image acquisition equipment in a real production environment, covering the melting state of seed crystals under different temperature conditions. Each image was manually labeled to ensure the accuracy of the seed crystal edge contours. During training, a stochastic gradient descent optimizer was used, with an initial learning rate set to 0.01 and a learning rate decay strategy of 0.9 every 10 training epochs to prevent the model from getting stuck in local optima in the later stages of training. The batch size was set to 32. To improve the model's generalization ability, various data augmentation techniques were employed, including random rotation ±10°, random cropping (cropping ratio from 80% to 100%), and horizontal flipping. After training with the above strategies, the model for the temperature testing phase can accurately extract the seed crystal edge contours from the seed crystal state feature information and determine the melting rate, outputting the temperature testing phase judgment results.
[0033] Understandably, the fusion stage model was trained on a dataset containing 15,000 multi-view image pairs, each of which was labeled with a mask for the fusion region, a numerical value for the fusion speed, and a category label for the fusion ratio.
[0034] It should be noted that during the construction of the dataset, each multi-view image pair was obtained from simultaneous dual-view acquisition in an actual production environment. The first and second view images correspond to furnace images acquired by two different angles of the reflecting mirror 14 in the crystal-guided vision image acquisition device. During the training of the welding stage model, the Adam optimizer was used with an initial learning rate of 0.0001, and a cosine annealing learning rate scheduling strategy was employed, with a batch size of 16. Simultaneously, online data augmentation methods were used, including real-time adjustment of the relative positional relationship between the two viewpoints during training. Specifically, this included random translation ±5 pixels and random rotation ±5° to increase data diversity and improve the model's robustness to viewpoint changes. After training using the above strategies, the welding stage model can simultaneously complete semantic segmentation of the melting region, regression prediction of melting speed, and classification of melting ratio based on the fused feature map, outputting the welding stage judgment result.
[0035] Understandably, the necking stage model was trained on a dataset containing 8,000 pairs of images before and after gas disturbance.
[0036] It should be noted that each image pair includes a clear image before gas interference and a degraded image after gas interference. During the dataset construction, the image pairs before and after gas interference were all derived from gas interference phenomena during the necking stage of sapphire crystal growth furnaces in actual production environments, and were manually screened to ensure image quality. During the training of the necking stage model, a stochastic gradient descent optimizer was used with an initial learning rate of 0.005, and a multinomial learning rate decay strategy was employed with a decay exponent of 0.9 and a batch size of 8. The expression for the multinomial learning rate decay strategy is as follows: , This represents the initial learning rate. Indicates the current iteration number. Indicates the total number of iterations. The decay exponent is represented by [insert value here]. To improve the model's robustness to gas interference, various data augmentation techniques are employed, including adding Gaussian noise, salt-and-pepper noise, and random blurring. Furthermore, adversarial training methods are used to generate synthetic data simulating gas interference to expand the training set. Specifically, a generative adversarial network is used to learn the distribution of real gas interference, generating additional pairs of interfered images to be added to the training set, thereby enhancing the model's generalization ability to unknown interference. After training with the above strategies, the necking-stage model can effectively eliminate gas interference, maintain the stability of the measurement data, and output the necking-stage judgment result.
[0037] Reference Figures 2 to 3 It is understood that the crystal-guided vision image acquisition device includes an industrial high-definition camera 6, a reflective lens assembly, a camera support rod 1, a support base 7, and an adjustment mechanism. The adjustment mechanism includes a camera mounting base 3, a camera height adjustment base 2, a camera angle adjustment base 4, and a reflective lens moving base 11. The bottom of the camera support rod 1 is rotatably mounted to one end of the support base 7 via the camera angle adjustment base 4. The camera height adjustment base 2 is slidably mounted on the camera support rod 1. The industrial high-definition camera 6 is mounted on the camera height adjustment base 2 via the camera mounting base 3. The reflective lens moving base 11 is located at the other end of the support base 7, and the reflective lens assembly is mounted on the reflective lens moving base 11.
[0038] It should be noted that the bottom of the camera support rod 1 is rotatably mounted to one end of the support base 7 via the camera angle adjustment base 4. The camera height adjustment base 2 is slidably mounted on the camera support rod 1. The industrial high-definition camera 6 is mounted on the camera height adjustment base 2 via the camera mounting base 3. By sliding the camera height adjustment base 2 along the camera support rod 1, the height of the industrial high-definition camera 6 can be adjusted. By using the angle adjustment component 5 (e.g., angle adjustment nut) provided on the camera angle adjustment base 4, the camera support rod 1 can be tilted forward or backward along the camera angle adjustment base 4, thereby achieving the optimal image acquisition angle of the industrial high-definition camera 6. The reflective lens assembly is mounted on the reflective lens moving base 11 and is used to reflect the image inside the sapphire crystal growth furnace observation window 10 to the industrial high-definition camera 6.
[0039] In some embodiments, two camera angle adjustment bases 4 are provided, and the number of camera support rods 1 corresponds to the number of camera angle adjustment bases 4. Each camera angle adjustment base 4 is provided with an angle adjustment nut 5. The two camera angle adjustment bases 4 are symmetrically arranged on both sides of one end of the support base 7. The camera support rods 1 are mounted on the support base 7 through the corresponding camera angle adjustment bases 4. The two ends of the camera height adjustment base 2 are slidably mounted on the two camera support rods 1 respectively. The camera mounting base 3 is mounted on the camera height adjustment base 2 and is located between the two camera support rods 1.
[0040] In some embodiments, the reflective lens assembly includes a reflective lens 14 and a lens clamp 13. The observation window end cover 9 of the sapphire crystal growth furnace is disposed at the other end of the support base 7 via the observation window fixing base 8. The reflective lens 14 is fixed on the lens clamp 13, and the lens clamp 13 is mounted on the reflective lens moving base 11. The reflective lens moving base 11 is mounted on the observation window end cover 9 of the sapphire crystal growth furnace and is used to adjust the horizontal position and rotation angle of the reflective lens 14. The reflective lens moving base 11 includes a reflective lens moving base body and a reflective rotating base 12. The reflective rotating base 12 is used to rotate the reflective lens 14 360 degrees around the fixed point as the center, so as to reflect the image inside the observation window 10 of the sapphire crystal growth furnace to the industrial high-definition camera 6.
[0041] Understandably, the detection and image processing system includes an image denoising module and a dynamic detection module. The image denoising module receives real-time images of the seed crystal, performs physical denoising on the real-time images of the seed crystal using an optical filter, and then uses a high-temperature radiation denoising algorithm and a dynamic blur compensation algorithm to perform software denoising on the physically denoised images to obtain denoised images. The dynamic detection module is connected to the image denoising module and is used to receive the denoised images, identify and depict the edge contours and bright and dark spots of the seed crystal from the denoised images to obtain the seed crystal state feature information, and output the seed crystal state feature information.
[0042] It should be noted that the optical filters include polarizing filters and infrared cutoff filters, used to eliminate stray light interference and suppress high-temperature infrared radiation, respectively. The high-temperature radiation denoising algorithm is based on an image gray-level distribution statistical model to suppress thermal noise in images under high-temperature conditions. The dynamic blur compensation algorithm uses Wiener filtering or blind deconvolution methods to compensate for motion blur caused by airflow disturbances in the furnace or camera shake. The dynamic detection module uses a deep learning-based semantic segmentation network. The output of the semantic segmentation network includes the seed crystal region mask, the coordinates of the seed crystal edge curve, and the position and gray-level values of bright and dark spots.
[0043] Sapphire crystals possess excellent optical, thermal, and mechanical properties, making them promising candidates for applications in aerospace, defense, semiconductor lighting, and communication technologies. The growth process of sapphire crystals places extremely high technical demands on the seeding operation, and the quality of the seeding is the most critical factor determining the overall crystal quality.
[0044] Currently, the sapphire crystal seeding process in the industry mainly relies on manual operation. Operators observe the dissolution state of the seed crystal inside the furnace through an observation window and manually raise and lower the seed crystal based on experience and furnace temperature changes to complete the seeding operation. However, this manual seeding method has the following problems: First, the seeding stability is poor, and it relies too much on the operator's experience, making it difficult to guarantee the consistency of the seeding; second, it cannot adapt to the differentiated needs of the multi-stage seeding process. The requirements for identification accuracy and control vary at each stage of the seeding process, and the existing method, which uses uniform manual observation or simple visual inspection, cannot achieve precise control throughout the entire process; third, the level of automation is low, lacking a complete closed-loop system from image acquisition and status recognition to control execution.
[0045] Therefore, the sapphire automatic seeding system based on a multi-stage visual algorithm model provided in this application embodiment 1. Acquires real-time images of the seeding process in the sapphire crystal growth furnace to obtain real-time seeding images; performs noise reduction processing on the real-time seeding images, and identifies the edge contours and bright / dark spots of the seed crystal from the noise-reduced images to obtain seed crystal state feature information; according to the current seeding stage, the seed crystal state feature information is sent to the corresponding visual judgment algorithm model: in the temperature testing stage, the seed crystal state feature information is sent to the temperature testing stage model, which uses a convolutional neural network containing residual connections to extract the edge contours of the seed crystal and determine the melting rate. The process involves obtaining the temperature testing stage judgment result; during the welding stage, the seed crystal state characteristic information is sent to the welding stage model, which uses a dual-flow feature fusion network to monitor the changes and proportions of the melting region, thus obtaining the welding stage judgment result; during the necking stage, the seed crystal state characteristic information is sent to the necking stage model, which uses an attention-based encoder-decoder structure to eliminate gas interference, thus obtaining the necking stage judgment result; operation instructions are generated based on the temperature testing stage judgment result, the welding stage judgment result, or the necking stage judgment result; and the temperature of the sapphire crystal growth furnace and the seed crystal pulling speed are adjusted according to the operation instructions until crystal pulling is completed. This application, through this setup, can adapt to the entire crystal pulling process, possess staged differentiated processing capabilities, and automatically complete image acquisition, state recognition, judgment decision-making, and execution control.
[0046] To verify the technical effect of the sapphire automatic crystal pulling system based on the multi-stage visual algorithm model in this embodiment, a comparative experiment was conducted. The experiment was carried out on the same type of sapphire crystal growth furnace. The crystal pulling operation was performed using the system, the traditional manual crystal pulling method, and the single visual method, and the crystal pulling success rate, crystal defect rate, and crystal quality consistency were compared and evaluated. (1) Comparison of crystal pulling success rate The experimental results show that the crystal pulling success rate of the system is significantly higher than that of the traditional manual crystal pulling method and the single visual method. It is believed that the improvement of the crystal pulling success rate of the system is mainly due to the precise control of the process parameters of each stage by the multi-stage visual algorithm model. In the temperature test stage, the system can quickly adjust the thermal field distribution and the position of the seed crystal by accurately extracting the edge contour of the seed crystal through the temperature test stage model (including convolutional neural network with residual connections) and combined with real-time melting speed monitoring, thereby ensuring the stability of the crystal pulling conditions. In the welding stage, the application of the welding stage model (dual-flow feature fusion network) effectively improves the accuracy of monitoring the changes and proportions of the melting area, and further optimizes the success rate of the crystal pulling operation. In contrast, traditional artificial crystal pulling methods rely on the operator's experience and judgment, which can easily lead to crystal pulling failure due to subjective factors; single vision methods lack multi-stage collaborative optimization capabilities and are difficult to cope with complex crystal pulling environments. (2) Comparison of crystal defect rates: Crystal defect rate is an important indicator for measuring the quality of sapphire crystals. Experimental results show that after adopting this system, the crystal defect rate is significantly reduced compared with traditional artificial crystal pulling methods. It is believed that this improvement is mainly attributed to the real-time monitoring and precise adjustment of melt state and process parameters in the crystal pulling process of this system: In the welding stage, this system uses the welding stage model to dynamically track the changes in the melting area, and optimizes the difference between the model prediction value and the true value through the loss function of the multi-task learning framework, thereby achieving precise control of melting speed and ratio, effectively reducing crystal defects caused by melt inhomogeneity or overheating; In the necking stage, this system eliminates the influence of gas interference on imaging through the necking stage model (an encoder-decoder structure based on an attention mechanism), ensuring the stability of measurement data and further reducing the defect rate. Traditional artificial crystal pulling methods lack real-time monitoring capabilities and cannot promptly detect and correct abnormalities during the crystal pulling process; single vision methods have insufficient algorithm robustness and anti-interference ability, making it difficult to maintain stable performance in complex environments.(3) Verification of crystal quality consistency: Crystal quality consistency is an important standard for evaluating the production quality of sapphire crystals. Experimental results show that this system has significant advantages in improving crystal quality consistency, mainly reflected in its multi-stage collaborative optimization capability and real-time adjustment mechanism: In the temperature testing stage, the model accurately measures the seed crystal outline and melting degree, ensuring the uniformity of the thermal field distribution and the accuracy of the seed crystal position; In the welding stage, the model further optimizes the stability of process parameters by dynamically monitoring the changes in melting area and speed; In the necking stage, the model eliminates the influence of gas interference on imaging, maintains high-precision measurement data, and thus improves the consistency of the crystal growth process. (4) Verification of Adaptability to Complex Environments: During the growth of sapphire crystals, complex environmental factors such as temperature changes and gas interference pose challenges to the performance of the visual algorithm. To verify the environmental adaptability of this system, the following verifications were conducted: At the hardware level, the crystal-growing visual image acquisition device of this system adopts a high-resolution, high-temperature resistant industrial camera and optimized optical lens components, combined with local cooling technology, to ensure that the equipment can still work stably under extreme temperature conditions; at the software level, the detection and image processing system designed an image preprocessing algorithm based on multi-frame fusion and background subtraction, which effectively eliminated the noise impact caused by gas interference; to address the image distortion problem caused by temperature changes, an adaptive correction model was used to dynamically compensate for real-time images. Experimental results show that the system after the above comprehensive optimization of hardware and software can still maintain high image recognition accuracy and measurement stability when facing severe temperature fluctuations and strong gas interference. (5) Real-time performance verification of the algorithm: The sapphire crystal pulling process has extremely high real-time requirements for the visual algorithm. This system adopts a series of strategies in terms of network structure optimization and computing resource allocation: by simplifying the number of convolutional layers, reducing the number of feature map channels, and introducing lightweight modules (such as depthwise separable convolution), the inference time of the model is significantly reduced. For example, in the edge contour extraction task during the temperature testing stage, the improved lightweight network architecture is used to replace the traditional deep network model, which can significantly shorten the processing time of a single frame image while maintaining high recognition accuracy. The sapphire automatic crystal pulling system based on the multi-stage visual algorithm model in this embodiment has achieved significant technical effects in terms of crystal pulling success rate, crystal defect rate reduction, and crystal quality consistency compared with traditional manual crystal pulling methods and single visual algorithms. In addition, by constructing a large-scale image database (covering multi-view images of the temperature testing stage, welding stage, and images before and after gas interference in the necking stage, etc.) and optimizing the loss function, the generalization ability and robustness of the model are further enhanced.
[0047] Reference Figure 4 , Figure 4This illustration shows an automatic crystal pulling method for sapphire based on a multi-stage visual algorithm model, provided by a second aspect embodiment of this application. Applied to the automatic crystal pulling system for sapphire based on a multi-stage visual algorithm model provided by a first aspect embodiment, the method includes the following steps: Step S1: Acquire real-time images of the crystal growth process inside the sapphire crystal growth furnace to obtain real-time images of crystal growth.
[0048] In this step, the industrial high-definition camera 6 and the reflective lens assembly in the crystal-growing visual image acquisition device are used to acquire images of the entire crystal-growing process in real time from the observation window of the sapphire crystal growth furnace. The acquisition frame rate is 10 frames / second, and the exposure time is dynamically adjusted from 1 millisecond to 10 milliseconds according to the brightness inside the furnace.
[0049] Step S2: Denoise the real-time image of the seed crystal and identify the edge contour and bright and dark spots of the seed crystal from the denoised image to obtain the seed crystal state feature information.
[0050] In this step, physical noise reduction is performed using optical filters, followed by software noise reduction using high-temperature radiation noise reduction algorithm and dynamic blur compensation algorithm to obtain a denoised image. Then, a deep learning semantic segmentation network is used to identify the edge contour curves and the position, area, and gray value of bright and dark spots of the seed crystal from the denoised image to form seed crystal state feature information.
[0051] Step S3: Based on the current seeding stage, the seed crystal state feature information is sent to the corresponding visual judgment algorithm model: In the temperature testing stage, the seed crystal state feature information is sent to the temperature testing stage model, which uses a convolutional neural network with residual connections to extract the seed crystal edge contour and judge the melting speed, thus obtaining the temperature testing stage judgment result; In the welding stage, the seed crystal state feature information is sent to the welding stage model, which uses a dual-flow feature fusion network to monitor the changes and proportions of the melting area, thus obtaining the welding stage judgment result; In the necking stage, the seed crystal state feature information is sent to the necking stage model, which uses an attention-based encoder-decoder structure to eliminate gas interference, thus obtaining the necking stage judgment result.
[0052] Step S4: Generate operation instructions based on the results of the temperature test stage, the welding stage, or the necking stage.
[0053] In this step, the operating instructions include temperature adjustment instructions and pulling speed adjustment instructions. The temperature adjustment instructions are used to control the heating power of the sapphire crystal growth furnace, and the pulling speed adjustment instructions are used to control the lifting and lowering speed of the seed crystal rod.
[0054] Step S5: Adjust the temperature of the sapphire crystal growth furnace and the pulling speed of the seed crystal according to the operation instructions until the crystal pulling is completed.
[0055] It should be noted that after the execution control unit receives the operation command, it adjusts the heating power through the PID controller to change the temperature inside the furnace, and at the same time drives the seed crystal rod to rise and fall at a specified speed through the servo motor, repeating steps S1 to S5 until the conditions for crystal development are met.
[0056] In the sapphire crystal growth process, the seed crystal pulling operation is the core step that determines the crystal quality. Currently, the industry mainly uses the manual seed crystal pulling method: operators observe the dissolution state of the seed crystal in the furnace through an observation window and manually control the raising and lowering of the seed crystal based on experience.
[0057] However, this manual crystal-pulling method has the following problems: First, manual crystal pulling relies excessively on the operator's experience and judgment. The varying skill levels and working conditions of different operators make it difficult to guarantee the consistency and repeatability of the crystal-pulling operation, resulting in significant fluctuations in crystal quality and a generally low success rate. Second, the requirements for image recognition accuracy and control strategies differ at each stage of the crystal-pulling process. Existing methods use uniform manual observation or simple visual inspection, failing to differentiate the processing based on the technical characteristics of different stages. This leads to decreased recognition accuracy in complex environments such as high-temperature radiation and gas interference, making it difficult to achieve precise control throughout the entire process. Third, there is a lack of a complete workflow that can automatically complete image acquisition, status recognition, stage judgment, command generation, and execution adjustment. The crystal-pulling operation still requires manual intervention, failing to achieve an automated closed loop from image acquisition to temperature and pulling speed adjustment.
[0058] Therefore, how to provide an automated sapphire crystal pulling method that can adapt to the entire crystal pulling process, has the ability to process differentiated data in stages, and can automatically complete image acquisition, status recognition, judgment and decision-making, and execution adjustment is a technical problem that urgently needs to be solved by those skilled in the art.
[0059] Based on this, the second aspect of this application provides an automated sapphire seeding method based on a multi-stage visual algorithm model. 1. By acquiring real-time images of the seeding process within the sapphire crystal growth furnace, obtaining real-time seeding images, and performing noise reduction processing on these images, the edge contours and bright / dark spots of the seed crystal are identified from the noise-reduced images to obtain seed crystal state feature information. This achieves automated image acquisition and state recognition of the seeding process, replacing manual observation and experience-based judgment, effectively eliminating inconsistencies in seeding caused by human factors, significantly improving the stability and repeatability of the seeding operation, thereby increasing the seeding success rate. 2. Different models are used in the temperature testing stage, welding stage, and necking stage: in the temperature testing stage, a convolutional neural network with residual connections is used to extract the edge contours of the seed crystal and determine the melting speed; in the welding stage, a dual-flow feature fusion network is used to monitor the changes and proportions of the melting area; in the necking stage, an encoder-decoder structure based on an attention mechanism is used to eliminate gas interference. Dedicated algorithm models are used for differentiated processing based on the technical characteristics of the three different stages of seeding, achieving precise control throughout the entire process. 3. Based on the results of the temperature test stage, welding stage, or necking stage, operation instructions are generated. The temperature of the sapphire crystal growth furnace and the pulling speed of the seed crystal are adjusted according to the operation instructions until the crystal pulling is completed, forming a complete automated closed-loop control process: data acquisition → noise reduction and identification → stage call → model judgment → instruction generation → adjustment and execution. This method can automatically complete the entire crystal pulling process without manual intervention, realizing closed-loop control from image acquisition to temperature and pulling speed adjustment, and significantly improving the automation level of the crystal pulling process.
[0060] This application, through this method, is able to adapt to the entire crystal pulling process, possess the ability to perform differentiated processing in stages, and automatically complete image acquisition, status recognition, judgment and decision-making, and execution adjustment.
[0061] In the description of the embodiments of this application, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," "third," or "fourth" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0062] In the description of the embodiments of this application, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
Claims
1. A sapphire automatic seeding system based on a multi-stage vision algorithm model, characterized in that, include: A crystal-leading visual image acquisition device is used to acquire real-time images of the crystal-leading process in a sapphire crystal growth furnace to obtain real-time crystal-leading images, and output the real-time crystal-leading images; The detection and image processing system is connected to the seed crystal visual image acquisition device. The detection and image processing system is used to receive the real-time image of the seed crystal, perform noise reduction processing on the real-time image of the seed crystal, identify the edge contour and bright and dark spots of the seed crystal from the noise-reduced image to obtain the seed crystal state feature information, and output the seed crystal state feature information. A three-stage visual judgment algorithm model is connected to the detection and image processing system. The three-stage visual judgment algorithm model includes a temperature testing stage model, a welding stage model, and a necking stage model. The three-stage visual judgment algorithm model is used to receive the seed crystal state feature information and send the seed crystal state feature information to one of the temperature testing stage model, the welding stage model, or the necking stage model according to the current crystal pulling stage. The temperature testing stage model is used to receive the seed crystal state feature information and use a convolutional neural network with residual connections to extract the seed crystal edge contour and determine the melting speed, thereby obtaining the temperature testing stage judgment result; the welding stage model is used to receive the seed crystal state feature information and use a dual-stream feature fusion network to monitor the changes and proportions of the melting area, thereby obtaining the welding stage judgment result; the necking stage model is used to receive the seed crystal state feature information and use an attention-based encoder-decoder structure to eliminate gas interference, thereby obtaining the necking stage judgment result. The three-stage visual judgment algorithm model is also used to generate operation instructions based on the judgment results of the temperature test stage, the judgment results of the welding stage, or the judgment results of the necking stage, and output the operation instructions. An execution control unit is connected to the three-stage visual judgment algorithm model. The execution control unit is used to receive the operation instructions and control the temperature of the sapphire crystal growth furnace and the seed crystal pulling speed according to the operation instructions.
2. The automatic sapphire seed crystal induction system based on a multi-stage visual algorithm model according to claim 1, characterized in that, In the temperature testing phase model: The convolutional neural network includes multiple convolutional layers and multiple pooling layers, which are connected alternately in sequence. The convolutional kernels of the convolutional layers include two sizes: 3×3 and 5×5. The residual connection is set in the middle and later stages of the convolutional neural network to directly superimpose the output of the previous layer onto the input of the next layer. The model for the temperature testing phase uses a weighted cross-entropy loss function as its optimization objective, and the weights of the weighted cross-entropy loss function are dynamically adjusted according to the proportion of positive samples.
3. The automatic sapphire seed crystal induction system based on a multi-stage visual algorithm model according to claim 1, characterized in that, In the welding stage model: The dual-stream feature fusion network includes a first sub-network and a second sub-network arranged in parallel, and a feature fusion layer connected to the output ends of the first sub-network and the second sub-network, respectively. The input end of the first sub-network is used to receive first-view image data, and the input end of the second sub-network is used to receive second-view image data. The first-view image data and the second-view image data are real-time images of the crystal pulling process captured by the crystal pulling visual image acquisition device from the observation window of the sapphire crystal growth furnace from two different angles. The first sub-network is used to extract high-level semantic features from the first viewpoint image data to obtain a first feature map and output the first feature map to the feature fusion layer; the second sub-network is used to extract high-level semantic features from the second viewpoint image data to obtain a second feature map and output the second feature map to the feature fusion layer. The feature fusion layer is used to merge the first feature map and the second feature map to obtain a fused feature map; The fusion stage model employs a multi-task learning framework, which simultaneously performs semantic segmentation of the fusion region, regression prediction of the fusion speed, and classification of the fusion ratio based on the fused feature map. The overall loss function of the multi-task learning framework is a weighted sum of the Dice coefficient loss function, the mean squared error loss function, and the binary cross-entropy loss function.
4. The automatic sapphire seed crystal picking system based on a multi-stage visual algorithm model according to claim 1, wherein, In the necking stage model: The attention-based encoder-decoder structure includes an encoder and a decoder, with the output of the encoder connected to the input of the decoder. The encoder uses a VGG-16 network to extract multi-level features from the input image; The decoder includes a spatial attention module and a channel attention module. The spatial attention module is used to enhance the spatial location information of the feature map, and the channel attention module is used to enhance the channel dimension information of the feature map. The encoder-decoder structure also embeds multiple residual dense connection blocks, which directly transmit the features of the corresponding layer in the encoder to the corresponding layer in the decoder through skip connections; The loss function of the necking stage model is a weighted sum of the structural similarity loss function and the L1 norm loss function.
5. The automatic sapphire seed crystal induction system based on a multi-stage visual algorithm model according to claim 1, wherein, The temperature testing stage model was trained using a dataset containing 10,000 labeled images, which cover the melting state of seed crystals under different temperature conditions.
6. The sapphire automatic crystal pulling system based on a multi-stage visual algorithm model according to claim 1, characterized in that, The welding stage model was trained on a dataset containing 15,000 multi-view image pairs, each of which was labeled with a mask of the melting area, a numerical value of the melting speed, and a category label of the melting ratio.
7. The sapphire automatic crystal pulling system based on a multi-stage visual algorithm model according to claim 1, characterized in that, The necking stage model was trained using a dataset containing 8,000 pairs of images before and after gas disturbance.
8. The sapphire automatic crystal pulling system based on a multi-stage visual algorithm model according to claim 1, characterized in that, The crystal-guided visual image acquisition device includes an industrial high-definition camera, a reflective lens assembly, a camera support rod, a fixed base, and an adjustment mechanism. The adjustment mechanism includes a camera mounting base, a camera height adjustment base, a camera angle adjustment base, and a reflective lens moving base. The bottom of the camera support rod is rotatably mounted to one end of the fixed base via the camera angle adjustment base. The camera height adjustment base is slidably mounted on the camera support rod. The industrial high-definition camera is mounted on the camera height adjustment base via the camera mounting base. The reflective lens moving base is located at the other end of the fixed base, and the reflective lens assembly is mounted on the reflective lens moving base.
9. The sapphire automatic crystal pulling system based on a multi-stage visual algorithm model according to claim 1, characterized in that, The detection and image processing system includes: The image denoising module is used to receive the real-time image of the crystal, perform physical denoising on the real-time image of the crystal through an optical filter, and perform software denoising on the physically denoised image using a high-temperature radiation denoising algorithm and a dynamic blur compensation algorithm to obtain a denoised image. A dynamic detection module is connected to the image denoising module. The dynamic detection module is used to receive the denoised image, identify and depict the edge contour and bright and dark spots of the seed crystal from the denoised image, obtain the seed crystal state feature information, and output the seed crystal state feature information.
10. A method for automatic crystal pulling of sapphire based on a multi-stage visual algorithm model, characterized in that, The method, applied to the sapphire automated crystal pulling system based on a multi-stage visual algorithm model as described in any one of claims 1 to 7, comprises: Real-time images of the crystal growth process inside the sapphire crystal growth furnace were acquired to obtain real-time images of the crystal growth process. The real-time image of the seed crystal is denoised, and the edge contour and bright and dark spots of the seed crystal are identified from the denoised image to obtain the seed crystal state feature information. Based on the current seed crystal stage, the seed crystal state characteristic information is sent to the corresponding visual judgment algorithm model: During the temperature testing phase, the seed crystal state feature information is sent to the temperature testing phase model. The temperature testing phase model uses a convolutional neural network with residual connections to extract the seed crystal edge contour and determine the melting rate, thereby obtaining the temperature testing phase judgment result. During the welding stage, the seed crystal state characteristic information is sent to the welding stage model, which uses a dual-flow feature fusion network to monitor the changes and proportions of the melting area and obtain the welding stage judgment result. During the necking stage, the seed crystal state characteristic information is sent to the necking stage model, which uses an attention-based encoder-decoder structure to eliminate gas interference and obtain the necking stage judgment result. An operation command is generated based on the judgment results of the temperature test stage, the welding stage, or the necking stage. Adjust the temperature of the sapphire crystal growth furnace and the seed crystal pulling speed according to the operation instructions until the crystal pulling is completed.