Automatic alignment method for long nozzle and steel ladle collector nozzle based on multi-source visual fusion
By using multi-source vision fusion technology and combining 3D and 2D industrial cameras with image segmentation models, automatic alignment of the long water inlet and the ladle water outlet was achieved, which solved the safety risks caused by manual operation and improved alignment accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-14
AI Technical Summary
In the continuous casting process of steel smelting, the alignment of the long nozzle and the ladle nozzle mainly relies on manual operation, which poses safety risks such as high temperature and molten steel splashing, and lacks visual perception capabilities.
A multi-source vision fusion method is adopted, which uses a 3D industrial camera set at the end of the four-axis gripper to collect three-dimensional point cloud data, combined with a 2D high-resolution industrial camera to collect high-resolution images, and uses a pre-trained image segmentation model to perform precise positioning, so as to achieve automatic alignment of the long nozzle and the ladle outlet.
It enables automatic alignment of the long nozzle and ladle outlet, reducing the work safety risks for smelting workers, improving alignment accuracy and efficiency, and avoiding the safety hazards of manual operation.
Smart Images

Figure CN121847742A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the fields of casting production and computer technology, and specifically to an automatic alignment method for long nozzles and ladle nozzles based on multi-source visual fusion. Background Technology
[0002] In the continuous casting process of steel smelting, the ladle nozzle and the long nozzle are two closely related and collaborative key components. They must be precisely matched to achieve stable pouring without leakage or air intake. Currently, a four-axis manipulator is typically used manually to align the long nozzle with the ladle nozzle to complete the molten steel pouring operation. However, the pouring process often involves risks such as high temperatures and molten steel splashing, thus increasing the safety risks for smelting workers. Summary of the Invention
[0003] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion later. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0004] Some embodiments of this disclosure propose an automatic alignment method for long water inlets and ladle water outlets based on multi-source visual fusion to solve the technical problems mentioned in the background section above.
[0005] In a first aspect, some embodiments of this disclosure provide an automatic alignment method for a ladle nozzle and a steel ladle outlet based on multi-source visual fusion. The method includes: in response to a four-axis gripper moving to a preset first observation position, acquiring three-dimensional point cloud data of the bottom of the steel ladle using a first camera, wherein the first camera is a 3D industrial camera mounted at the end of the four-axis gripper; determining first position information based on the three-dimensional point cloud data, wherein the first position information represents the three-dimensional region where the coarsely located ladle outlet is located; performing coordinate transformation on the first position information to obtain a second observation position; and at the second observation position, through... The second camera acquires a high-resolution image of the ladle drain outlet, wherein the second camera is a 2D high-resolution industrial camera installed at the end of the quadcopter. Based on a pre-trained image segmentation model and the high-resolution image, second position information is determined, wherein the second position information represents the sub-pixel position of the ladle drain outlet. Based on the first and second position information, third position information is determined, wherein the third position information represents the position of the ladle drain outlet in a three-dimensional coordinate system. Based on the third position information, the quadcopter is controlled to automatically align the lifted long nozzle with the ladle drain outlet.
[0006] Secondly, some embodiments of this disclosure provide an automatic alignment device for a long water inlet and a ladle outlet based on multi-source visual fusion. The device includes: a first acquisition unit configured to acquire three-dimensional point cloud data of the bottom of the ladle via a first camera in response to a four-axis gripper moving to a preset first observation position; wherein the first camera is a 3D industrial camera disposed at the end of the four-axis gripper; a first determination unit configured to determine first position information based on the three-dimensional point cloud data, wherein the first position information represents the three-dimensional region where the ladle outlet is located for coarse positioning; a coordinate transformation unit configured to perform coordinate transformation on the first position information to obtain a second observation position; and a second acquisition unit configured to acquire data at the second observation position. The system comprises: a position determination unit, configured to acquire a high-resolution image of the ladle drain outlet using a second camera (a 2D high-resolution industrial camera mounted at the end of the quadcopter); a second determination unit, configured to determine second position information based on a pre-trained image segmentation model and the high-resolution image, wherein the second position information represents the sub-pixel position of the ladle drain outlet; a third determination unit, configured to determine third position information based on the first and second position information, wherein the third position information represents the position of the ladle drain outlet in a three-dimensional coordinate system; and a control unit, configured to control the quadcopter to automatically align the lifted long nozzle with the ladle drain outlet based on the third position information.
[0007] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0008] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0009] The above embodiments of this disclosure have the following beneficial effects: Through the automatic alignment method for the long nozzle and ladle outlet based on multi-source visual fusion in some embodiments of this disclosure, automatic alignment of the long nozzle and ladle outlet is achieved, greatly reducing the work safety risks for smelting workers. Specifically, firstly, in response to the four-axis gripper moving to a preset first observation position, three-dimensional point cloud data of the bottom of the ladle is acquired by a first camera, wherein the first camera is a 3D industrial camera installed at the end of the four-axis gripper. In practice, existing four-axis grippers are mainly manually controlled and lack visual perception capabilities; therefore, this disclosure achieves the acquisition of three-dimensional point cloud data through a 3D industrial camera at the end of the four-axis gripper. Secondly, based on the three-dimensional point cloud data, first position information is determined, wherein the first position information represents the three-dimensional region where the ladle outlet is located for coarse positioning. This, combined with the three-dimensional point cloud data, achieves coarse positioning of the ladle outlet. Next, coordinate transformation is performed on the first position information to obtain a second observation position. This allows for the adjustment of the observation position of the 2D high-resolution industrial camera, ensuring that the second camera can acquire high-resolution images of the ladle outlet from a full-view perspective. Further, based on a pre-trained image segmentation model and the aforementioned high-resolution images, second position information is determined, representing the sub-pixel level position of the ladle outlet. Additionally, based on the first and second position information, third position information is determined, representing the position of the ladle outlet in a three-dimensional coordinate system. This achieves further precise positioning of the ladle outlet. Finally, based on the third position information, the four-axis gripper is controlled to automatically align the lifted long nozzle with the ladle outlet. In summary, this method achieves automatic alignment of the long nozzle and the ladle outlet, significantly reducing the safety risks for smelting workers. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a flowchart of some embodiments of the automatic alignment method for long water inlet and ladle water outlet based on multi-source visual fusion according to the present disclosure; Figure 2 This is a schematic diagram of a 3D point cloud data. Figure 3 This is a structural schematic diagram of some embodiments of the automatic alignment device for long water inlet and ladle water outlet based on multi-source visual fusion according to the present disclosure. Figure 4This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0012] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0013] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0014] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0015] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0016] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0017] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] refer to Figure 1 The flowchart 100 illustrates some embodiments of the automatic alignment method for long nozzles and ladle outlets based on multi-source visual fusion according to this disclosure. This automatic alignment method for long nozzles and ladle outlets based on multi-source visual fusion includes the following steps: Step 101: In response to the quadcopter moving to the preset first observation position, three-dimensional point cloud data of the bottom of the ladle is acquired by the first camera.
[0019] In some embodiments, the execution subject (e.g., a computing device) of the automatic alignment method for the long nozzle and ladle outlet based on multi-source visual fusion can acquire three-dimensional point cloud data of the bottom of the ladle through a first camera in response to the movement of the four-axis gripper to a preset first observation position.
[0020] The first observation position is a pre-set position facing the bottom of the ladle. The four-axis holder is a multi-axis control device used to support the long nozzle and connect it to the ladle's outlet. The first camera is a 3D industrial camera located at the end of the four-axis holder.
[0021] In practice, firstly, the aforementioned actuator can control the quadcopter to move to the first observation position. Then, when the quadcopter moves to the first observation position, it triggers the first camera located at the end of the quadcopter to perform three-dimensional point cloud data acquisition towards the bottom of the ladle.
[0022] As an example, see Figure 2 The diagram shown is a schematic representation of the 3D point cloud data. Figure 2 The image shows 3D point cloud data captured by the first camera facing the bottom of the ladle.
[0023] It should be noted that the aforementioned computing devices can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed on the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.
[0024] Step 102: Determine the first location information based on the 3D point cloud data.
[0025] In some embodiments, the aforementioned execution entity may determine the first location information based on three-dimensional point cloud data.
[0026] The aforementioned first location information represents the three-dimensional region where the ladle drain outlet is located, based on coarse positioning. The first location information may include: the coordinates of the corner points of the three-dimensional region and the coordinates of the center point of the ladle drain outlet's mating surface.
[0027] In practice, since the first observation position is a preset position, the relative position between the ladle drain and the first observation position is fixed. Furthermore, the ladle drain is a hardware structure with a fixed appearance. Therefore, firstly, the appearance features of the ladle drain can be pre-collected at the first observation position. Then, through feature matching, the 3D region of the 3D point cloud data that matches the appearance features is used as the first location information.
[0028] In some optional implementations of certain embodiments, the execution entity determines the first location information based on the aforementioned three-dimensional point cloud data, including: Step S1: Perform planar projection on the above three-dimensional point cloud data to obtain first planar point cloud data, second planar point cloud data, third planar point cloud data and fourth planar point cloud data.
[0029] The point cloud data of the first plane is obtained by projecting it onto the first plane. The point cloud data of the second plane is obtained by projecting it onto the second plane. The point cloud data of the third plane is obtained by projecting it onto the third plane. The point cloud data of the fourth plane is obtained by projecting it onto the fourth plane. The first plane and the second plane are perpendicular to each other. The first plane and the third plane are perpendicular to each other. The second plane and the third plane are perpendicular to each other. The fourth plane is perpendicular to the first plane and forms a preset plane angle with the second plane. Assuming that the coordinate system of the 3D point cloud data is a spatial rectangular coordinate system (o-XYZ), the first plane can be the XoY plane in the spatial rectangular coordinate system. The second plane can be the XoZ plane in the spatial rectangular coordinate system. The third plane can be the YoZ plane in the spatial rectangular coordinate system. The preset plane angle between the fourth plane and the second plane is 30 degrees.
[0030] In practice, the 3D point clouds in 3D point cloud data represent their corresponding positions using 3D coordinates. Therefore, by using planar projection, the 3D point clouds can be projected onto a first plane, a second plane, a third plane, and a fourth plane respectively, thus converting the 3D point cloud data into 2D data. This method reduces data processing complexity and allows for parallel background feature matching of the point cloud data from the first, second, third, and fourth planes, thereby improving matching speed.
[0031] Step S2: Perform background feature matching on the first planar point cloud data, the second planar point cloud data, the third planar point cloud data, and the fourth planar point cloud data respectively to determine the first background point cloud data, the second background point cloud data, the third background point cloud data, and the fourth background point cloud data.
[0032] Specifically, the first background point cloud data represents the two-dimensional point cloud contained in the first planar point cloud data that matches background features. The second background point cloud data represents the two-dimensional point cloud contained in the second planar point cloud data that matches background features. The third background point cloud data represents the two-dimensional point cloud contained in the third planar point cloud data that matches background features. The fourth background point cloud data represents the two-dimensional point cloud contained in the fourth planar point cloud data that matches background features.
[0033] In practice, since the first observation position is a preset position, the relative position between the ladle outlet and the first observation position is fixed. Therefore, the area outside the ladle outlet changes little, allowing for pre-collection of background data and extraction of background features. Then, through feature matching, the correlation between the two-dimensional point clouds included in the first, second, third, and fourth planar point cloud data and the corresponding background features of the background area is determined, thereby filtering out the first, second, third, and fourth background point cloud data. Specifically, the background features can be constructed using a geometric model in the form of a KD-tree, allowing for similarity calculation using geometric distance to match the background features with the two-dimensional point clouds.
[0034] Step S3: Perform weighted fusion of the first background point cloud data, the second background point cloud data, the third background point cloud data, and the fourth background point cloud data to determine the fifth background point cloud data.
[0035] In practice, although planar projection compresses 3D point clouds into 2D point clouds, improving data processing speed, the 3D-to-2D conversion loses information in one dimension. Therefore, this disclosure uses a weighted fusion method to process the first, second, third, and fourth background point cloud data as a whole. Specifically, the 2D point clouds included in the first, second, third, and fourth background point cloud data all correspond to background similarity. Taking the first background point cloud data as an example, the background similarity of its included 2D point clouds represents the similarity between the 2D point cloud matching the background features from the perspective of the first plane and the background features. Therefore, the background similarities of multiple 2D point clouds corresponding to the same 3D point clouds included in the first, second, third, and fourth background point cloud data can be weighted and summed, filtered using a background similarity threshold (the weighted background similarity being greater than the threshold), and a 2D-to-3D remapping can be performed to obtain the fifth background point cloud data. In the fifth background point cloud data, the background similarity of the three-dimensional point clouds included is greater than the background similarity threshold. For example, taking the first background point cloud data as an example, its two-dimensional point clouds match the background features. However, in the second, third, and fourth planar point cloud data, the two-dimensional point clouds corresponding to the same three-dimensional point clouds may or may not match the background features. Therefore, a weighted fusion method is used to determine whether the corresponding three-dimensional point cloud belongs to the background region by weighting the background similarity of the four two-dimensional point clouds corresponding to the same three-dimensional point cloud data. The weighted fusion can be performed using an average weighting method (the weights can be 0.25, 0.25, 0.25, 0.25 respectively). In addition, the weights can be adjusted according to actual needs, which is not limited here.
[0036] Step S4: Perform planar radius filtering on the first planar point cloud data, the second planar point cloud data, the third planar point cloud data, and the fourth planar point cloud data respectively to obtain the first filtered point cloud data, the second filtered point cloud data, the third filtered point cloud data, and the fourth filtered point cloud data.
[0037] Specifically, the first filtered point cloud data consists of outliers (two-dimensional point cloud) included in the first planar point cloud data. The second filtered point cloud data consists of outliers (two-dimensional point cloud) included in the second planar point cloud data. The third filtered point cloud data consists of outliers (two-dimensional point cloud) included in the third planar point cloud data. The fourth filtered point cloud data consists of outliers (two-dimensional point cloud) included in the fourth planar point cloud data.
[0038] In practice, conventional planar radius filtering algorithms for 3D point clouds construct spheres to identify outliers (noise). This disclosure, however, converts the 3D point cloud data into 2D point clouds. Therefore, when performing planar radius filtering on the aforementioned first, second, third, and fourth planar point cloud data, a circle is constructed on the plane. When the number of 2D point clouds within the circle is less than the filtering threshold, the 2D point cloud corresponding to the circle's center is identified as an outlier. This achieves planar radius filtering for 2D point clouds, resulting in the first, second, third, and fourth filtered point cloud data.
[0039] Step S5: Perform weighted fusion of the first filtered point cloud data, the second filtered point cloud data, the third filtered point cloud data, and the fourth filtered point cloud data to obtain the fifth filtered point cloud data.
[0040] In practice, similar to step S3, although the planar projection method compresses the 3D point cloud into a 2D point cloud, improving data processing speed, the 3D-to-2D conversion loses information in one dimension. Specifically, the 2D point clouds (outliers) included in the first, second, third, and fourth filtered point cloud data all have outlier values. Outlier value represents the ratio of the number of 2D point clouds contained within a 2D point cloud to the filtering threshold. Since the number of 2D point clouds (outliers) within the corresponding circles in the first, second, third, and fourth filtered point cloud data is less than the filtering threshold, the outlier values corresponding to the 2D point clouds (outliers) included in the first, second, third, and fourth filtered point cloud data are all less than 1. Therefore, this disclosure can use a weighted fusion method to perform weighted fusion on multiple two-dimensional point clouds (outliers) corresponding to the same point cloud included in the first, second, third, and fourth filtered point cloud data, thereby further filtering out the corresponding three-dimensional point clouds with a weighted outlier of less than 1, which are then used as the fifth filtered point cloud data. For example, taking the first filtered point cloud data as an example, the outlier of the two-dimensional point clouds it includes is less than 1, but the two-dimensional point clouds corresponding to the three-dimensional point clouds it includes may not be considered outliers or may be considered outliers in the second, third, and fourth plane point cloud data. Therefore, by performing a weighted sum of the outliers of the four two-dimensional point clouds corresponding to the same three-dimensional point cloud, it is determined whether the corresponding three-dimensional point cloud is an outlier, thus obtaining the fifth filtered point cloud data. One approach is to use a weighted average (with weights of 0.25, 0.25, 0.25, and 0.25 respectively) to sum the outliers of four corresponding 2D point clouds of the same 3D point cloud. The weights can also be adjusted as needed, and this is not limited here.
[0041] Step S6: Based on the fifth background point cloud data and the fifth filtered point cloud data, remove point cloud data from the three-dimensional point cloud data to obtain the removed three-dimensional point cloud data.
[0042] In practice, the aforementioned execution master can remove the 3D point clouds located in the fifth background point cloud data and the 3D point clouds located in the fifth filtered point cloud data from the 3D point cloud data to obtain the removed 3D point cloud data.
[0043] Step S7: Perform geometric feature search based on the removed 3D point cloud data to obtain the first location information.
[0044] In practice, since the drain outlet is similar to a cylinder, the first location information can be obtained through geometric feature search methods such as cylinder model fitting and plane segmentation.
[0045] Step 103: Perform coordinate transformation on the first position information to obtain the second observation position.
[0046] In some embodiments, the execution entity may perform coordinate transformation on the first position information to obtain the second observation position.
[0047] The second observation position represents the updated image acquisition position of the quadcopter.
[0048] In practice, assuming that the shooting direction of the second camera and the distance between it and the ladle outlet are fixed, the second observation position along the shooting direction, corresponding to the first position information, can be obtained through coordinate transformation.
[0049] Step 104: At the second observation position, acquire a high-resolution image of the ladle outlet using the second camera.
[0050] In some embodiments, the aforementioned execution entity may acquire high-resolution images of the ladle outlet using a second camera at a second observation position.
[0051] The second camera is a 2D high-resolution industrial camera located at the end of the quadcopter.
[0052] In practice, after the quadcopter moves to the second observation position, the aforementioned actuator can control the second camera located at the end of the quadcopter to acquire high-resolution images of the three-dimensional area corresponding to the first position information.
[0053] In some optional implementations of certain embodiments, the aforementioned executing entity, at the aforementioned second observation position, acquires a high-resolution image of the ladle outlet using a second camera, including: Step S1: Control the quadcopter to move to the second observation position.
[0054] In practice, the aforementioned executing entity can send displacement commands to the quadcopter to move the quadcopter to the aforementioned second observation position.
[0055] Step S2: In response to the quadcopter moving to the second observation position, control the second camera to face the three-dimensional region corresponding to the first position information.
[0056] In practice, when the second camera is connected to the four-axis controller via a gimbal, the gimbal can be controlled to orient the second camera toward the three-dimensional area corresponding to the first position information. When the second camera is not connected to the four-axis controller via a gimbal, the axis movement of the four-axis controller can be controlled to orient the second camera toward the three-dimensional area corresponding to the first position information.
[0057] Step S3: Adjust the focal length of the second camera.
[0058] In practice, autofocus can be performed based on the current focus mode of the second camera. Focus modes can include spot focus mode, area focus mode, and tracking focus mode, among others.
[0059] Step S4: In response to the completion of focus adjustment, control the second camera to acquire a high-resolution image of the ladle drain outlet.
[0060] In practice, the aforementioned executing entity can send image acquisition commands to the second camera to control the second camera to acquire high-resolution images of the ladle outlet.
[0061] Step 105: Determine the second location information based on the pre-trained image segmentation model and the high-resolution image.
[0062] In some embodiments, the aforementioned execution entity may determine the second location information based on a pre-trained image segmentation model and a high-resolution image.
[0063] The second location information represents the sub-pixel level location of the ladle's drain outlet. Image segmentation models such as U-Net and DeepLab can be used.
[0064] Optionally, the image segmentation model consists of a backbone feature extractor, a multi-scale feature extractor, a feature decoder, and a location regressor. The backbone feature extractor uses a ResNet101 residual network model. The multi-scale feature extractor consists of a 1×1 convolutional layer, a 3×3 convolutional layer with a dilation rate of 6, a 3×3 convolutional layer with a dilation rate of 6, and a 3×3 convolutional layer with a dilation rate of 18. These layers are configured in parallel to perform multi-scale feature extraction on the feature maps extracted by the backbone feature extractor. The feature decoder consists of a concatenation layer, a 1×1 convolutional layer, and an upsampling network. The concatenation layer stitches together the multiple feature maps output by the multi-scale feature extractor. The 1×1 convolutional layer performs original-size feature convolution on the features output by the concatenation layer. The upsampling network upsamples the features output from the 1×1 convolutional layer to restore a feature map of the same size as the high-resolution image. The location regressor uses a regressor based on preset anchor boxes. The location regressor performs location regression using the output of the upsampling network as input to obtain the second location information.
[0065] In some optional implementations of certain embodiments, the execution entity determines the second location information based on a pre-trained image segmentation model and the high-resolution image, including: Step S1: Use the above-mentioned backbone feature extractor to extract features from the high-resolution image to obtain an initial feature map.
[0066] Step S2: Using the multi-scale feature extractor described above, multi-scale feature extraction is performed on the initial feature map to obtain a set of multi-scale feature maps.
[0067] Step S3: Using the stitching layer included in the above feature decoder, the multi-scale feature maps in the above multi-scale feature map set are stitched together to obtain a stitched feature map; Step S4: The concatenated feature map is convolved using the lightweight convolutional network included in the feature decoder to obtain the convolved feature map. Step S5: The convolutional feature map is upsampled using the upsampling network included in the feature decoder to obtain the upsampled feature map.
[0068] The feature map size of the upsampled feature map is consistent with the image size of the high-resolution image.
[0069] The second location information is generated using the aforementioned location regressor and the aforementioned upsampled feature map.
[0070] In practice, a backbone feature extractor is introduced to accelerate the training rate of the image segmentation model during the training phase. Secondly, a multi-scale feature extractor is designed to encode multi-scale information, thereby enhancing the image segmentation model's ability to extract features for image edge details. Next, a feature decoder is designed to achieve fine-grained semantic segmentation at the original image size. Finally, a location regressor is combined to regress the location of the ladle outlet. The image segmentation model disclosed herein can overcome complex background interference such as ambient lighting and high-temperature radiation in complex casting scenarios, achieving accurate boundary segmentation.
[0071] Step 106: Determine the third location information based on the first location information and the second location information.
[0072] In some embodiments, the aforementioned executing entity may determine the third location information based on the first location information and the second location information.
[0073] Among them, the third location information represents the location of the ladle outlet in the three-dimensional coordinate system.
[0074] In practice, the first position information represents the three-dimensional coordinates of the three-dimensional region where the ladle outlet is located, while the second position information represents the two-dimensional coordinates of the ladle outlet. Therefore, given that the first and second cameras are pre-calibrated, the aforementioned execution entity can convert the second position information into three-dimensional coordinates using the rotation matrix and translation vector corresponding to the cameras. Based on this, the aforementioned execution entity can obtain new three-dimensional center coordinates by weighted summing of the three-dimensional coordinates corresponding to the first position information (e.g., the center coordinates of the three-dimensional region and the three-dimensional coordinates corresponding to the second position information). Then, based on the positional relationship between the center coordinates and corner coordinates included in the first position information, the corner coordinates are updated on the new three-dimensional center coordinates to obtain the third position information.
[0075] In some optional implementations of certain embodiments, the execution entity determines the third location information based on the first location information and the second location information, including: Step S1: Perform a three-dimensional coordinate transformation on the above second position information to obtain candidate position information.
[0076] In practice, the second position information represents the two-dimensional image coordinates in the high-resolution image. Under the premise that the second camera is pre-calibrated, the second position information can be converted into candidate position information in the three-dimensional coordinate system according to the rotation matrix and translation vector corresponding to the second camera, thereby transforming the second position information into the same coordinate system as the first position information.
[0077] Step S2: Perform position weighting on the above first position information and the above candidate position information to obtain weighted position information.
[0078] In practice, the center coordinates included in the first position information and the corresponding three-dimensional coordinates of the candidate position information can be weighted and summed to obtain the weighted position information.
[0079] Step S3: Using the position corresponding to the weighted position information above as the center, perform edge fitting of the drain outlet based on the first position information above to determine the direction vector.
[0080] In practice, since the first location information corresponds to a three-dimensional region, the edge of the ladle drain can be determined by using the weighted location as the center and the three-dimensional point cloud within the three-dimensional region corresponding to the first location information as the basis, through circle fitting. The direction vector can then be determined based on the plane where the drain edge is located.
[0081] Step S4: Generate the third position information based on the weighted position information and the direction vector.
[0082] The third location information can include: weighted location information and a direction vector. The weighted location information can be used to characterize the docking position of the long nozzle and the ladle outlet. The direction vector can characterize the docking direction of the long nozzle.
[0083] Step 107: Based on the third position information, control the four-axis gripper to automatically align the lifted long water inlet with the ladle water outlet.
[0084] In some embodiments, the aforementioned execution entity can control the four-axis holder to automatically align the lifted long water inlet with the ladle water outlet based on the third position information.
[0085] In practice, the aforementioned execution entity can take the current position of the long water inlet supported by the four-axis holders as the starting position, the weighted position information included in the third position information as the ending position, and the direction corresponding to the direction vector as the docking direction. Through path planning, the movement path of the four-axis holders can be generated to achieve automatic alignment of the supported long water inlet with the ladle water outlet.
[0086] In some optional implementations of certain embodiments, the aforementioned execution entity controls the four-axis gripper to automatically align the lifted long nozzle with the ladle drain outlet based on the third position information, including: Step S1: Based on the third position information and the current position of the four-axis gripper, a movement path is planned to generate the gripper movement path.
[0087] In practice, the starting position of the long nozzle held by the quadcopter can be used as the starting position, and the weighted position information included in the third position information can be used as the ending position. By using polynomial interpolation, such as cubic or quintic polynomials, the path of the quadcopter can be planned to obtain the movement path of the quadcopter.
[0088] Step S2: Send the above-mentioned gripper movement path to the gripper control system corresponding to the above-mentioned four-axis gripper.
[0089] The gripper control system can be a control system used to control the displacement movement of a four-axis gripper.
[0090] Step S3: Drive the four-axis holder along the holder's movement path to the position corresponding to the third position information through the holder control system, so as to automatically align the lifted long water inlet with the ladle water outlet.
[0091] In practice, the holder control system can convert the holder movement path into control commands to drive the surrounding holders along the holder movement path, move the long water inlet to the position corresponding to the third position information (weighted position information), and connect the long water inlet and the ladle water outlet according to the direction corresponding to the direction vector.
[0092] In practice, this disclosure achieves automatic identification of ladle drain and long drain by combining a 3D industrial camera and a 2D high-resolution industrial camera. Specifically: (1) Compared with the identification method using line-scanning lidar, it greatly reduces structural complexity, improves identification efficiency and accuracy, and avoids the safety risks associated with line-scanning lidar needing to extend into the narrow area under the ladle for scanning. (2) Compared with the method of using a temperature sampling sensor to assist the four-axis holder in sensing the ladle drain, it also avoids the problem of temperature sampling sensors being easily interfered with by environmental factors such as high temperature and vibration on site, resulting in error accumulation and poor system stability.
[0093] The above embodiments of this disclosure have the following beneficial effects: Through the automatic alignment method for the long nozzle and ladle outlet based on multi-source visual fusion in some embodiments of this disclosure, automatic alignment of the long nozzle and ladle outlet is achieved, greatly reducing the work safety risks for smelting workers. Specifically, firstly, in response to the four-axis gripper moving to a preset first observation position, three-dimensional point cloud data of the bottom of the ladle is acquired by a first camera, wherein the first camera is a 3D industrial camera installed at the end of the four-axis gripper. In practice, existing four-axis grippers are mainly manually controlled and lack visual perception capabilities; therefore, this disclosure achieves the acquisition of three-dimensional point cloud data through a 3D industrial camera at the end of the four-axis gripper. Secondly, based on the three-dimensional point cloud data, first position information is determined, wherein the first position information represents the three-dimensional region where the ladle outlet is located for coarse positioning. This, combined with the three-dimensional point cloud data, achieves coarse positioning of the ladle outlet. Next, coordinate transformation is performed on the first position information to obtain a second observation position. This allows for the adjustment of the observation position of the 2D high-resolution industrial camera, ensuring that the second camera can acquire high-resolution images of the ladle outlet from a full-view perspective. Further, based on a pre-trained image segmentation model and the aforementioned high-resolution images, second position information is determined, representing the sub-pixel level position of the ladle outlet. Additionally, based on the first and second position information, third position information is determined, representing the position of the ladle outlet in a three-dimensional coordinate system. This achieves further precise positioning of the ladle outlet. Finally, based on the third position information, the four-axis gripper is controlled to automatically align the lifted long nozzle with the ladle outlet. In summary, this method achieves automatic alignment of the long nozzle and the ladle outlet, significantly reducing the safety risks for smelting workers.
[0094] Further reference Figure 3 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an automatic alignment device for long nozzles and ladle outlets based on multi-source visual fusion. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, this automatic alignment device for long water inlets and ladle drain outlets based on multi-source visual fusion can be specifically applied to various electronic devices.
[0095] like Figure 3As shown, an automatic alignment device 300 for long water inlet and ladle drain outlet based on multi-source visual fusion in some embodiments includes: a first acquisition unit 301, a first determination unit 302, a coordinate transformation unit 303, a second acquisition unit 304, a second determination unit 305, a third determination unit 306, and a control unit 307. The first acquisition unit 301 is configured to acquire three-dimensional point cloud data of the bottom of the ladle via a first camera in response to the four-axis gripper moving to a preset first observation position. The first camera is a 3D industrial camera located at the end of the four-axis gripper. The first determination unit 302 is configured to determine first position information based on the three-dimensional point cloud data. The first position information represents the three-dimensional region where the ladle drain outlet is located (coarsely positioned). The coordinate transformation unit 303 is configured to perform coordinate transformation on the first position information. The transformation yields a second observation position; a second acquisition unit 304 is configured to acquire a high-resolution image of the ladle outlet at the second observation position using a second camera, wherein the second camera is a 2D high-resolution industrial camera located at the end of the quadcopter; a second determination unit 305 is configured to determine second position information based on a pre-trained image segmentation model and the high-resolution image, wherein the second position information represents the sub-pixel level position of the ladle outlet; a third determination unit 306 is configured to determine third position information based on the first and second position information, wherein the third position information represents the position of the ladle outlet in a three-dimensional coordinate system; and a control unit 307 is configured to control the quadcopter to automatically align the lifted long nozzle with the ladle outlet based on the third position information.
[0096] It is understandable that the units described in the multi-source visual fusion-based automatic alignment device 300 for long water inlets and ladle outlets are related to the reference... Figure 1 The steps described in the method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the automatic alignment device 300 for long water inlets and ladle drain outlets based on multi-source visual fusion and the units contained therein, and will not be repeated here.
[0097] The following is for reference. Figure 4 It shows a schematic diagram of the structure of an electronic device (e.g., a computing device) 400 suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0098] like Figure 4As shown, the electronic device 400 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory 402 or a program loaded from a storage device 408 into a random access memory 403. The random access memory 403 also stores various programs and data required for the operation of the electronic device 400. The processing unit 401, the read-only memory 402, and the random access memory 403 are interconnected via a bus 404. An input / output interface 405 is also connected to the bus 404.
[0099] Typically, the following devices can be connected to the input / output interface 405: input devices 406 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 408 including, for example, magnetic tape, hard disk, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 4 Each box shown can represent a device or multiple devices as needed.
[0100] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a storage device 408, or installed from a read-only memory 402. When the computer program is executed by the processing device 401, it performs the functions defined in the methods of some embodiments of this disclosure.
[0101] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0102] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0103] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: respond to the quadcopter moving to a preset first observation position; acquire three-dimensional point cloud data of the bottom of the ladle using a first camera, wherein the first camera is a 3D industrial camera located at the end of the quadcopter; determine first position information based on the three-dimensional point cloud data, wherein the first position information represents the three-dimensional region where the ladle outlet is located (coarsely located); perform coordinate transformation on the first position information to obtain a second observation position; and at the second observation position, through... The second camera acquires a high-resolution image of the ladle drain outlet, wherein the second camera is a 2D high-resolution industrial camera installed at the end of the quadcopter. Based on a pre-trained image segmentation model and the high-resolution image, second position information is determined, wherein the second position information represents the sub-pixel position of the ladle drain outlet. Based on the first and second position information, third position information is determined, wherein the third position information represents the position of the ladle drain outlet in a three-dimensional coordinate system. Based on the third position information, the quadcopter is controlled to automatically align the lifted long nozzle with the ladle drain outlet.
[0104] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0106] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0107] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. An automatic alignment method for long water inlets and ladle drain outlets based on multi-source visual fusion, characterized in that, include: In response to the quadcopter moving to a preset first observation position, three-dimensional point cloud data of the bottom of the ladle is acquired by a first camera, wherein the first camera is a 3D industrial camera set at the end of the quadcopter. Based on the three-dimensional point cloud data, a first location information is determined, wherein the first location information represents the three-dimensional region where the ladle drain is located in coarse positioning; The first position information is transformed by coordinates to obtain the second observation position; At the second observation position, a high-resolution image of the ladle outlet is acquired by a second camera, wherein the second camera is a 2D high-resolution industrial camera set at the end of the quadcopter. Based on a pre-trained image segmentation model and the high-resolution image, a second location information is determined, wherein the second location information characterizes the subpixel level location of the ladle drain outlet. Based on the first location information and the second location information, a third location information is determined, wherein the third location information represents the location of the ladle outlet in a three-dimensional coordinate system; Based on the third position information, the four-axis gripper is controlled to automatically align the lifted long water inlet with the ladle water outlet.
2. The automatic alignment method for long water inlet and ladle drain outlet based on multi-source visual fusion according to claim 1, characterized in that, Determining the first location information based on the three-dimensional point cloud data includes: The three-dimensional point cloud data is projected onto a plane to obtain first plane point cloud data, second plane point cloud data, third plane point cloud data, and fourth plane point cloud data. The first plane point cloud data is obtained by projecting along a first plane, the second plane point cloud data is obtained by projecting along a second plane, the third plane point cloud data is obtained by projecting along a third plane, and the fourth plane point cloud data is obtained by projecting along a fourth plane. The first plane and the second plane are perpendicular to each other, the first plane and the third plane are perpendicular to each other, the second plane and the third plane are perpendicular to each other, and the fourth plane is perpendicular to the first plane and forms a preset plane angle with the second plane. Background feature matching is performed on the first planar point cloud data, the second planar point cloud data, the third planar point cloud data, and the fourth planar point cloud data respectively to determine the first background point cloud data, the second background point cloud data, the third background point cloud data, and the fourth background point cloud data; The fifth background point cloud data is determined by weighted fusion of the first background point cloud data, the second background point cloud data, the third background point cloud data, and the fourth background point cloud data.
3. The automatic alignment method for long water inlet and ladle drain outlet based on multi-source visual fusion according to claim 2, characterized in that, The step of determining the first location information based on the three-dimensional point cloud data further includes: Planar radius filtering is performed on the first planar point cloud data, the second planar point cloud data, the third planar point cloud data, and the fourth planar point cloud data respectively to obtain the first filtered point cloud data, the second filtered point cloud data, the third filtered point cloud data, and the fourth filtered point cloud data. The fifth filtered point cloud data is obtained by weighted fusion of the first filtered point cloud data, the second filtered point cloud data, the third filtered point cloud data, and the fourth filtered point cloud data. Based on the fifth background point cloud data and the fifth filtered point cloud data, the three-dimensional point cloud data is removed to obtain the removed three-dimensional point cloud data. Geometric feature search is performed on the removed 3D point cloud data to obtain the first location information.
4. The automatic alignment method for long water inlet and ladle drain outlet based on multi-source visual fusion according to claim 3, characterized in that, At the second observation position, a high-resolution image of the ladle outlet is acquired using a second camera, including: Control the quadcopter to move to the second observation position; In response to the quadcopter moving to the second observation position, the second camera is controlled to face the three-dimensional region corresponding to the first position information; Adjust the focus of the second camera; In response to the completion of focus adjustment, the second camera is controlled to acquire a high-resolution image of the ladle drain outlet.
5. The automatic alignment method for long water inlet and ladle drain outlet based on multi-source visual fusion according to claim 4, characterized in that, The image segmentation model consists of a backbone feature extractor, a multi-scale feature extractor, a feature decoder, and a location regressor. The feature decoder consists of a concatenation layer, a lightweight convolutional network, and an upsampling network. The step of determining the second location information based on the pre-trained image segmentation model and the high-resolution image includes: The high-resolution image is used to extract features through the backbone feature extractor to obtain an initial feature map; The multi-scale feature extractor is used to extract multi-scale features from the initial feature map to obtain a set of multi-scale feature maps. The feature decoder uses a stitching layer to stitch together the multi-scale feature maps in the multi-scale feature map set to obtain a stitched feature map. The concatenated feature map is processed by a lightweight convolutional network included in the feature decoder to obtain a convolutional feature map. The convolutional feature map is upsampled using the upsampling network included in the feature decoder to obtain an upsampled feature map, wherein the feature map size of the upsampled feature map is consistent with the image size of the high-resolution image; The second location information is generated using the location regressor and the upsampled feature map.
6. The automatic alignment method for long water inlet and ladle drain outlet based on multi-source visual fusion according to claim 5, characterized in that, The step of determining the third location information based on the first location information and the second location information includes: Perform a three-dimensional coordinate transformation on the second location information to obtain candidate location information; The first location information and the candidate location information are weighted to obtain weighted location information; Using the position corresponding to the weighted position information as the center, the direction vector is determined by fitting the drain outlet edge based on the first position information; The third position information is generated based on the weighted position information and the direction vector.
7. The automatic alignment method for long water inlet and ladle drain outlet based on multi-source visual fusion according to claim 6, characterized in that, The step of controlling the four-axis gripper to automatically align the lifted long nozzle with the ladle drain outlet based on the third position information includes: Based on the third position information and the current position of the four-axis gripper, a movement path is planned to generate the gripper's movement path; The movement path of the gripper is sent to the gripper control system corresponding to the four-axis gripper; The four-axis gripper is driven by the gripper control system to move along the gripper movement path to the position corresponding to the third position information, so as to automatically align the lifted long water inlet with the ladle water outlet.
8. An automatic alignment device for long water inlets and ladle drain outlets based on multi-source visual fusion, characterized in that, include: The first acquisition unit is configured to acquire three-dimensional point cloud data of the bottom of the ladle through a first camera in response to the movement of the quadcopter to a preset first observation position. The first camera is a 3D industrial camera set at the end of the quadcopter. The first determining unit is configured to determine first location information based on the three-dimensional point cloud data, wherein the first location information represents the three-dimensional region where the ladle drain is located by coarse positioning. A coordinate transformation unit is configured to perform coordinate transformation on the first position information to obtain a second observation position; The second acquisition unit is configured to acquire high-resolution images of the ladle drain outlet at the second observation position via a second camera, wherein the second camera is a 2D high-resolution industrial camera located at the end of the quadcopter. The second determining unit is configured to determine second location information based on a pre-trained image segmentation model and the high-resolution image, wherein the second location information characterizes the subpixel level ladle outlet location. The third determining unit is configured to determine third position information based on the first position information and the second position information, wherein the third position information represents the position of the ladle drain in a three-dimensional coordinate system; The control unit is configured to control the four-axis gripper to automatically align the lifted long water inlet with the ladle water outlet based on the third position information.
9. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.
10. A computer-readable medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1 to 7.