An intelligent matching method and system for virtual and real coordinates of an additive repair robot
By extracting the feature of the real image and fusion of multi-scale feature, and calculating the center coordinates of the object with the minimum external rectangle algorithm, the problems of low efficiency and poor accuracy of virtual and real coordinate matching of additive repair robots are solved, and efficient virtual and real coordinate matching is achieved, which improves repair accuracy and reduces the labor force for manual adjustment.
Patent Information
- Application Number
- CN202210390357.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-04-14
AI Technical Summary
In the prior art, the virtual and real coordinate matching of additive repair robots has problems such as low efficiency, poor accuracy, and manual adjustments require time and effort.
An intelligent matching method for virtual and real coordinates of additive repair robots is adopted. By extracting features and fusion of real images and enhancing feature descriptions using the Attention module, the center coordinates of the object are calculated in combination with the minimum external rectangle algorithm to achieve coordinate matching between the real objects and the simulated three-dimensional model.
It improves the image segmentation quality, reduces the matching error of virtual and real coordinates, improves the accuracy and efficiency of additive repair, reduces the labor force for manual adjustment, and meets the needs of industrial applications.
Smart Images

Figure CN114662612B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of virtual technology and relates to a method and system for intelligent matching of virtual and real coordinates of an additive repair robot. Background Art
[0002] With the continuous improvement of the demand of the national manufacturing industry, industrial robots are more and more widely used in the fields of automobile manufacturing, machining, grinding and polishing, welding, additive manufacturing, etc., and are also the main development direction for realizing automatic and precise additive repair. The complex and diverse repair requirements pose a huge challenge to the trajectory planning of robots. Traditional robot teaching programming has problems such as low efficiency and poor accuracy.
[0003] Digital twin first plans the repair trajectory of the object to be repaired in the simulation environment and generates numerical control codes, and then controls the robot to complete the additive repair task in the real world according to the numerical control codes. The error-free matching of the simulation coordinates and the real coordinates of the repaired object is the premise and guarantee for realizing the precise repair of the object. When the repaired object, the end effector, the robot working space, etc. are repositioned or changed, the results of the previous offline programming need to be readjusted according to the changed scenario, and currently, manual methods are often used for adjustment, which is time-consuming and laborious. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the prior art and provide a method and system for intelligent matching of virtual and real coordinates of an additive repair robot.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An intelligent matching method for virtual and real coordinates of an additive repair robot includes the following steps:
[0007] S1: Extract features from the physical image to obtain multi-scale features of different stages of the physical image;
[0008] S2: Fuse the multi-scale features of each stage with the upsampled features of the same dimension to obtain the fused features of the image; enhance the fused features to obtain the enhanced integrated features, and use the integrated features for prediction to obtain the segmented image of the physical object;
[0009] S3: Describe the position of the obtained segmented image to obtain the real coordinates of the physical object, and at the same time establish a simulation three-dimensional model of the physical object, and adjust the position of the simulation three-dimensional model based on the real coordinates of the physical object to achieve coordinate matching between the physical object and the simulation three-dimensional model.
[0010] A further improvement of the present invention is that:
[0011] The feature extraction of the physical image in S1 includes the following steps of extracting features from the physical image through a feature extraction network:
[0012] S1.1: Input size is C j ×H j ×W j The feature map P j After one convolution block, the module performs channel splitting in the first residual stage to obtain the split feature P Aj :
[0013]
[0014] S1.2: Calculate direct connection features:
[0015] P Bj = Cov(P j ) (2)
[0016] S1.3: Split feature P Aj After one convolution block, it enters the secondary residual stage. The output feature of the secondary residual stage is P mj :
[0017]
[0018] S1.4: Fusion of equation (2) and equation (3) to obtain multi-scale feature F j , where j = 2, 3, 4,
[0019]
[0020] In the formula, Cov(.) represents the convolution operation, Indicates the features of the index part after the channel is split. This is the channel fusion operation.
[0021] The S2 comprises the following steps:
[0022] S2.1: Establish a segmentation prediction network, define the upsampled features of each stage under the segmentation prediction network as G1~G4, and define the multi-scale features F of each stage j Fuse with the upsampled features at the same latitude to obtain the fused feature S i ; and fusion feature S i To enhance:
[0023] S2.2: Fusion feature S i Compress and obtain the reduced feature T i ;
[0024] S2.3: Complete Feature T i Encoding in the vertical coordinate (H, 1) and horizontal coordinate (1, W) directions to obtain encoding features of different dimensions and Among them, the encoded representation with the vertical coordinate h in the c channel is:
[0025]
[0026] The encoded representation with the horizontal coordinate w is:
[0027]
[0028] In the formula, x c (h, n) represents the eigenvalue at the position (h, n) in the c channel;
[0029] S2.4: Introduce the attenuation coefficient r and output the integrated feature through Equation (7)
[0030]
[0031] In the formula, σ(.) represents the activation function, as shown in Equation (8):
[0032]
[0033] In the formula, ω is an input feature, and min(.) and max(.) represent the functions for obtaining the minimum value and the maximum value respectively;
[0034] S2.5: Use the dimension splitting function to split Predict the fusion coefficients of two dimensions through Equations (9) and (10) and and respectively represent the importance degrees of the position information and the spatial information;
[0035]
[0036]
[0037] Estimate the output feature A through Equation (11) i :
[0038]
[0039] In the formula, φ(·, site) represents the eigenvalue in the site direction after dimension splitting, and λ(·) is the Sigmoid activation function, is the product operation;
[0040] S2.6: Use a prediction convolution and a Sigmoid activation function to predict the segmented image and obtain the segmented image of the physical object.
[0041] In S2.1, channel compression is performed through a single-layer 3×3 convolutional network.
[0042] S3 includes the following steps:
[0043] S3.1: Describe the position of the obtained segmented image, using the top view T view of the physical image as the input image, and output the position information of the physical object through Equation (12):
[0044] x, y, h p , w p , γ = ψ(f(T view )) (12)
[0045] In the formula, f(.) represents the object segmentation network, ψ(.) is the minimum bounding rectangle algorithm, (x, y), h p , w p respectively represent the center coordinates of the object to be repaired, the length and width of the bounding rectangle, and γ is the deflection angle;
[0046] S3.2: Use the Graham algorithm to calculate the minimum convex hull U min of the object image, and obtain the total number N min of the edges in U sum ;
[0047] S3.3: Define an array V rect and initialize it;
[0048] S3.4: Take any edge S min in U num as the starting edge, and define the left endpoint O num of S left as the rotation center;
[0049] S3.5: Rotate S num , and judge whether S num is parallel to the horizontal axis of the image coordinates. If it holds, insert the number num, rotation angle R num of edge S angle , and the minimum bounding rectangle R area into the array V rect , let N sum = N sum - 1, and execute S3.6; otherwise, continue to execute S3.5;
[0050] S3.6: Judge whether N sum = 0 holds. If it holds, select the next edge clockwise, update the information of S num , and execute S3.5; otherwise, execute S3.7;
[0051] S3.7: According to the area R of the minimum bounding rectanglearea Sort the V rect array to obtain the minimum bounding rectangle, and obtain the object position information x, y, h according to Equation (12) p , w p , γ;
[0052] S3.8: Adjust the rotation position of the physical object at the simulation end according to γ;
[0053] S3.9: Obtain the position representation x', y', h of the object at the simulation end p ', w p ', γ', and obtain the deviation d of the X and Y axes x = x - x', d y = y - y', and according to d x , d y Translate the repaired object at the simulation end. After the translation is completed, update d x , d y , and judge d x ≤ 1 pixel and d y ≤ 1 pixel. If it holds, the matching is completed; otherwise, continue to execute S3.9
[0054] The established physical three-dimensional model is consistent with the actual size of the physical object
[0055] When performing coordinate matching, place the simulation three-dimensional model of the physical object and the physical object on the virtual operation platform and the actual operation platform respectively. Taking the operation platform as a reference, with the Z axis regarded as coaxial, adjust the horizontal coordinates for coordinate matching
[0056] An intelligent virtual-real coordinate matching system for an additive repair robot includes a feature extraction module, a segmentation prediction module, and a coordinate matching module;
[0057] The feature extraction module is used to extract features from the physical object image to obtain multi-scale features of the physical object at different stages;
[0058] The segmentation prediction module is used to fuse the multi-scale features of each stage with the upsampled features of the same dimension to obtain the fused features of the image; enhance the fused features to obtain the enhanced integrated features, and use the integrated features for prediction to obtain the segmentation image of the physical object;
[0059] The coordinate matching module is used to describe the position of the obtained segmentation image to obtain the real coordinates of the physical object, and at the same time establish a simulation three-dimensional model of the physical object, and adjust the position of the simulation three-dimensional model based on the real coordinates of the physical object to achieve coordinate matching between the physical object and the simulation three-dimensional model
[0060] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of claims 1-7 are implemented.
[0061] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] The present invention discloses an intelligent matching method for virtual and real coordinates of an additive repair robot. By performing segmentation processing on a physical image, not only can multi-scale features of the physical object be obtained, but also the number of parameters can be reduced, thereby enhancing the global and position information of the object, improving the feature space description ability, and improving the image segmentation quality. Based on this, the virtual and real coordinates of the physical object are matched and corrected, reducing the position error between the simulation object model and the physical object, improving the additive repair accuracy and efficiency, meeting the requirements of industrial applications, replacing manual adjustment, and saving labor. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0065] Figure 1 It is a network diagram of object segmentation of the present invention;
[0066] Figure 2 It is a CSPBlock multi-scale residual module diagram of the present invention;
[0067] Figure 3 It is an Attention module diagram of the present invention;
[0068] Figure 4 It is a position information description diagram of the present invention;
[0069] Figure 5 It is a comparative experiment test result diagram of the present invention;
[0070] Figure 6 It is a digital twin operation result diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0072] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0073] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0074] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is habitually placed during use, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0075] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "installed", "connected", "coupled" are to be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0076] The following further describes the present invention in detail with reference to the accompanying drawings:
[0077] An intelligent matching method for virtual and real coordinates of an additive repair robot is disclosed in an embodiment of the present invention. Specifically, an image segmentation model is used to obtain a segmented image of an object in an actual scene, and the true coordinates of the object are estimated by the minimum bounding rectangle. Accordingly, the simulation end performs translation and rotation operations to adjust the position of the 3D model of the simulated object, realizing intelligent coordinate matching. To minimize the virtual and real coordinate error, a multi-scale residual UNet image segmentation model based on Attention is proposed. The multi-scale features of the object are extracted by a multi-level residual strategy, and the Attention strategy is used to enhance the network's description of position information, improve the feature space expression ability, and enhance the image segmentation quality.
[0078] See Figure 1 , the multi-scale residual UNet object image segmentation model based on Attention disclosed in the embodiment of the present invention is composed of a left-right symmetric feature extraction network and a segmentation prediction network.
[0079] Step 1: The feature extraction network first obtains the encoded feature F1 by using a traditional convolution kernel, and then uses three CSPBlock multi-scale residual modules as shown in Figure 2 to obtain the information description of the object, and the encoded features of the three modules are defined as F2 to F4 respectively. The CSPBlock multi-scale residual module, through a multi-level residual connection method, can significantly reduce the computational cost while better retaining multi-scale features and enhancing the description ability of the object. The specific extraction process is as follows:
[0080] Step 1.1: Assume that the input is a feature map P j ×H j ×W j of size C j . In the first-level residual stage, channel splitting is performed through formula (1) to obtain the split feature P Aj , effectively reducing the network parameters.
[0081]
[0082] Step 1.2: Calculate the direct connection feature P Bj using formula (2) to completely retain the current feature.
[0083] P Bj = Cov(P j ) (2)
[0084] Step 1.3: The split feature P Aj calculates the output feature P mj of the second-level residual through formula (3).
[0085]
[0086] Step 1.4: Fuse PBj With P mj Obtain multi-scale feature F j , j = 2, 3, 4.
[0087]
[0088] In the formula, Cov(.) represents the convolution operation, is the feature of the index part after channel splitting, is the channel fusion operation.
[0089] Step 2.1: Define the upsampling features of each stage in the segmentation prediction network as G1 to G4, and directly fuse the upsampling feature G i of the segmentation prediction network at each stage with the encoded feature F i of the same dimension, i = 1, 2, 3, 4 to obtain the fused feature S i , see Figure 3 , introduce the Attention module to enhance the position description of the fused feature S i , effectively suppress interference features, and improve the spatial expressiveness of features.
[0090] Step 2.2: Compress the channels of the fused feature S 1i of size C 1i ×H 1i ×W i through a single-layer 3×3 convolutional network to obtain the reduced feature T i , reducing the computational loss.
[0091] Step 2.3: Re-encode the features in the horizontal and vertical dimensions using the feature map decomposition method. Specifically, complete the encoding in the vertical coordinate (H, 1) and horizontal coordinate (1, W) directions of the features to obtain the encoded features and The encoded representation of the vertical coordinate h in the c channel is expressed by Equation (5), and the encoded representation of the horizontal coordinate w is expressed by Equation (6).
[0092]
[0093]
[0094] In the formula, x c (h, n) represents the feature value at the position (h, n) in the c channel.
[0095] After that, use the channel fusion operation, convolution operation and H_swish activation function to coordinate and integrate the encoded features of different dimensions, while highlighting the description of the object location information and improving the correlation of spatial information.
[0096] Step 2.4: To limit the size of the module, a decay coefficient r is introduced in the convolution operation, and the integrated feature T of size D i / r×1×(W 1i +H 1i ) is output. In Equation (7), the specific representation of the activation function represented by σ(.) is shown in Equation (8); S , in the equation, ω is a feature of the input, and min(.) and max(.) respectively represent the functions for obtaining the minimum value and the maximum value;
[0097]
[0098]
[0099] where ω is an input feature, and min(.) and max(.) represent the functions for obtaining the minimum value and the maximum value respectively;
[0100] Step 2.5: The dimension splitting function is used to split the feature and the fusion coefficients of two dimensions are predicted through Equations (9) and (10) and to represent the importance degrees of the position information and the spatial information,
[0101]
[0102]
[0103] and the final output feature A i is estimated through Equation (11).
[0104]
[0105] Step 2.6: After the segmentation prediction network undergoes four upsampling operations and four attention modules, an accurate description A4 of the object segmentation information is obtained, and a segmentation image is predicted through a 1×1 convolution. The size of the segmentation image is the same as that of the input image, improving the accuracy of object coordinate prediction.
[0106] Step 3: Coordinate correction
[0107] Step 3.1: When matching coordinates, a simulated three-dimensional model of an object is constructed. The size of the simulated model is the same as that of the object, and the virtual and real objects are placed on a virtual-real operation platform with a fixed position. It can be regarded as having no deviation in the Z-axis and mainly focuses on horizontal adjustment. In the embodiment of the present invention, the minimum circumscribed rectangle algorithm is introduced, and the object center coordinates are estimated according to the segmentation image. An overhead view T of a terracotta warrior head to be repaired is input view , and the output object position information is Equation (12), and the specific description can be seen in Figure 4 .
[0108] x,y,h p ,wp , γ = ψ(f(T view )) (12)
[0109] where f(.) represents the object segmentation network, ψ(.) is the minimum bounding rectangle algorithm, (x, y), h p , w p represent the center coordinates of the object to be repaired, the length and width of the bounding rectangle respectively, and γ is the deflection angle;
[0110] Step 3.2: Use the Graham algorithm with a lower time complexity to calculate the minimum convex hull U of the object image min , and obtain the total number N of edges in U min ; sum
[0111] Step 3.3: Define an array V rect and initialize it;
[0112] Step 3.4: Take any edge S min in U as the starting edge, and define the left endpoint O num of S num as the rotation center; left
[0113] Step 3.5: Rotate S num , and determine whether S num is parallel to the horizontal axis of the image coordinates. If it holds, insert the number num of edge S num , the rotation angle R angle , and the minimum bounding rectangle R area into the array V rect . Let N sum = N sum - 1, and execute Step 3.6; otherwise, continue to execute Step 3.5;
[0114] Step 3.6: Determine whether N sum = 0 holds. If it holds, select the next edge clockwise, update the information of S num , and execute Step 3.5; otherwise, execute Step 3.7;
[0115] Step 3.7: Sort the array V area according to the area R of the minimum bounding rectangle to obtain the minimum bounding rectangle, and obtain the object position information x, y, h rect , w p , γ according to Equation (12); p
[0116] Step 3.8: Adjust the rotation position of the physical object at the simulation end according to γ;
[0117] Step 3.9: Obtain the position representation x', y', h of the object at the simulation endp ', w p ', γ', obtain the deviation d of the X and Y axes x = x - x', d y = y - y', according to d x , d y Translate the repair object at the simulation end. After the translation ends, update d x , d y , judge d x ≤ 1 pixel and d y ≤ 1 pixel holds. If it holds, the intelligent matching is completed; otherwise, continue to execute step 3.9.
[0118] Obtaining an accurate image segmentation result is an important guarantee for the coordinate matching of the additive repair object. Since there is no publicly available dataset in the field of additive repair, in the embodiments of the present invention, the handbag publicly available dataset and the self-built additive repair dataset are selected to conduct experiments to verify the effectiveness of the proposed image segmentation method. In addition, the commonly used evaluation index in the field of image segmentation, the average intersection over union (P mIoU ) is used for the analysis of the experimental results.
[0119]
[0120] In the formula, N is the total number of all images used for testing, and the variable descriptions are shown in Table 1.
[0121] Table 1 Variable descriptions of the average intersection over union (PmIoU)
[0122]
[0123] The handbag publicly available dataset has a complex shooting background and a large variety of types, including 550 RGB color images and their labeled ground truth images. 500 images are used for network training, and 50 images are used for result testing; the self-built additive repair dataset consists of five additive repair object bodies, namely ceramic cups, screws, intersection pipes, terracotta warrior heads, and impellers, which will be used in subsequent experiments. A total of 300 image samples are collected, 270 images are used for model training, and 30 images are used for result testing.
[0124] To fully train the model, data augmentation methods are adopted to randomly crop, rotate, scale, etc. the sample images, expand the number of dataset samples, and enhance the robustness of the model. Taking the self-built dataset as an example, after data augmentation, the number of training images is increased to 800, and the number of testing images is increased to 200, meeting the learning requirements of the model.
[0125] Some experimental results of the self-built dataset are as Figure 5 shown, and the comparison results of multiple methods for the handbag dataset and the self-built dataset are shown in Table 2.
[0126] See Figure 5 , the input images are from five different objects in the self-built dataset, and there are obvious differences in the shape, color, size, etc. of the objects. The experimental results in the figure show that the method of the embodiment of the present invention has comparable segmentation performance to UNet on the self-built dataset, indicating that the method of the embodiment of the present invention has strong adaptability.
[0127] Table 2 Comparison results of various segmentation methods
[0128]
[0129] It can be seen from Table 2 that compared with UNet, when the number of parameters of the model proposed in the embodiment of the present invention decreases by 20.02M, the prediction accuracy of the self-built dataset increases by 0.01; the segmentation accuracy of the handbag dataset is comparable, proving that the method in this paper has good generalization ability. Compared with the advanced algorithms Res-Net, SegNet and L-UNet, the segmentation effect is slightly inferior. The reason is that in order to maintain a relatively optimal number of parameters to meet the requirements of industrial real-time performance, this paper greatly compresses the channels of the fused feature S i , resulting in the loss of some features and causing a slight decrease in the segmentation result. The model with fewer parameters in the embodiment of the present invention is still a better choice. The experimental results of the ablation study of the Attention module show that the model achieves a more accurate prediction effect with only an increase of 0.02M in the number of parameters, proving the effectiveness of the Attention module proposed in the embodiment of the present invention.
[0130] After obtaining the accurately segmented image, using the minimum rectangle to obtain the center coordinates of the object is another important link to achieve the matching of virtual and real coordinates. The embodiment of the present invention selects the TURIN robot 1, the end effector 2 of the robot, the camera 3, the object to be repaired 4, and the repair platform 5 of 600mm×600mm as shown in Figure 6 to complete the construction of the real scene. In the experiment, the size of the input image is set to 640×640, and multiple matching experiments are carried out using the terracotta warrior heads in different placement positions. The matching relationship between the estimated physical center point coordinates and the center point coordinates of the object on the simulation side is shown in Table 3.
[0131] Table 3 Matching of center point coordinate positions
[0132]
[0133] It can be seen from Table 3 that the difference between the estimated object center coordinates and the real object center coordinates is less than or equal to 1 in both the horizontal and vertical directions, indicating that the matching error is less than 1mm, which can meet the requirements of most industrial digital twin tasks and verifies the effectiveness of the coordinate intelligent matching algorithm in the embodiment of the present invention.
[0134] Through Figure 6The digital twin experiment verifies the method described in the embodiments of the present invention. Through actual measurement, the virtual-real coordinate error of the first repair point of the robot is less than 1 mm, further verifying the effectiveness of the method in this article.
[0135] An embodiment of the present invention discloses a virtual-real coordinate intelligent matching system, including a feature extraction module, a segmentation prediction module, and a coordinate matching module;
[0136] The feature extraction module is used to extract features from the physical image to obtain multi-scale features of different stages of the physical object;
[0137] The segmentation prediction module is used to fuse the multi-scale features of each stage with the upsampled features of the same dimension to obtain the fused features of the image; enhance the fused features to obtain the enhanced integrated features, and predict the integrated features to obtain the segmentation image of the physical object;
[0138] The coordinate matching module is used to describe the position of the obtained segmentation image to obtain the real coordinates of the physical object, and at the same time establish a simulation three-dimensional model of the physical object, and adjust the position of the simulation three-dimensional model based on the real coordinates of the physical object to achieve coordinate matching between the physical object and the simulation three-dimensional model.
[0139] A schematic diagram of a terminal device provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned various method embodiments are implemented. Or, when the processor executes the computer program, the functions of each module / unit in the above-mentioned various device embodiments are implemented.
[0140] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.
[0141] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.
[0142] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0143] The memory can be used to store the computer program and / or module. By running or executing the computer program and / or module stored in the memory, and by invoking the data stored in the memory, the processor implements various functions of the terminal device.
[0144] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0145] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent matching method for virtual and real coordinates of an additive repair robot, characterized in that, It includes the following steps: S1: Extract features from the physical object image to obtain multi-scale features at different stages of the physical object image; S2: Fuse the multi-scale features at each stage with the upsampled features of the same dimension to obtain the fused features of the image; enhance the fused features to obtain the enhanced integrated features, and use the integrated features for prediction to obtain the segmented image of the physical object image; S3: Obtain the true coordinates of the physical object by describing the position of the segmented image; establish a simulated 3D model of the physical object, and adjust the position of the simulated 3D model based on the true coordinates of the physical object to achieve coordinate matching between the physical object and the simulated 3D model; In the above S1, the steps for extracting features from the physical object image include the following steps: extract features from the physical object image through a feature extraction network: S1.1: Input size is C j × H j × W j Feature map P j After one convolution block, the module performs channel splitting in the first residual stage to obtain split features P Aj : S1.2: Calculate the direct connection features: S1.3: Split features After passing through one convolutional block, it enters the secondary residual stage, and the output feature of the secondary residual stage is P mj : S1.4: Fuse Equation (2) and Equation (3) to obtain multi-scale features F j , where j = 2, 3, 4, In the formula, Cov (.) represents a convolution operation, represents taking the index part of the features after channel splitting, is a channel fusion operation.
2. The intelligent matching method for virtual and real coordinates of an additive repair robot according to claim 1, characterized in that, The above S2 includes the following steps: S2.1: Establish a segmentation prediction network, and define the upsampling features at each stage of the segmentation prediction network as G 1 to G 4, fuse the multi-scale features at each stage F j with the upsampling features of the same dimension to obtain fused features S i ; and enhance the fused features S i as follows: S2.2: Compress the fused feature S i to obtain the reduced feature T i ; S2.3: Complete features T i Encode in the vertical coordinate (H, 1) and horizontal coordinate (1, W) directions to obtain encoded features of different dimensions and , where c the encoding expression of the vertical coordinate in the channel is h : The horizontal coordinate is w The encoded representation of which is: In the formula, x c ( h , n ) represents c the eigenvalue at the position in the channel ( h , n ); S2.4: Introduce the attenuation coefficient r and output the integrated features through Equation (7). : In the formula, represents the activation function, see Equation (8): In the formula, is a feature of the input, min (.) and max (.) represent functions for obtaining the minimum value and the maximum value respectively; S2.5: Split using the dimensionality splitting function , and predict the fusion coefficients of the two dimensions through Equations (9) and (10) and , and respectively represent the importance levels of the position information and the spatial information; Features of the predicted output through formula (11) A i : In the formula, represents the eigenvalue taken in the site direction after dimensionality splitting, is the Sigmoid activation function, is the multiplication operation; S2.6: Use a prediction convolution and a Sigmoid activation function to predict the segmented image to obtain the segmented image of the physical object.
3. An intelligent matching method for virtual and real coordinates of an additive repair robot according to claim 2, characterized in that In the above S2.1, channel compression is performed through a single-layer 3×3 convolution network.
4. An intelligent matching method for virtual and real coordinates of an additive repair robot according to claim 1, characterized in that The above S3 includes the following steps: S3.
1. Describe the position of the obtained segmented image, using the top view of the physical image T view as the input image, and output the position information of the physical object through Equation (12): In the formula, f (.) represents the object segmentation network, is the minimum bounding rectangle algorithm, ( x , y ), h p , w p respectively represent the center coordinates of the object to be repaired, the length and width of the bounding rectangle, γ is the deflection angle; S3.2: Calculate the minimum convex hull of the object image using Graham's algorithm U min , and obtain U min the total number of edges in N sum ; S3.3: Define an array V rect and initialize it; S3.4: Starting from U min any one side S num as the starting side, define S num the left endpoint O left as the rotation center; S3.5: Rotation S num , determine S num whether it is parallel to the horizontal axis of the image coordinates. If it holds, insert the number S num of the side num , the rotation angle R angle , the minimum area bounding rectangle R area into the array V rect , and let N sum = N sum -1, and execute S3.6; otherwise, continue to execute S3.5; S3.6: Determine N sum if = 0 holds, then clockwise select the next edge, update S num information, and execute S3.5; otherwise, execute S3.7; S3.7: According to the area R of the minimum bounding rectangle area sort V rect the array to obtain the minimum circumscribed rectangle, and obtain the object position information according to Equation (12) x , y , h p , w p , γ ; S3.8: According to γ Adjust the rotation position of the physical object on the simulation side; S3.9: Obtain the position representation of the object on the simulation side x' , y' , h p ' , w p ' , γ' , and obtain the deviation of the X and Y axes d x = x - x' , d y = y - y' , according to d x , d y translate the repaired object on the simulation side. After the translation is completed, update d x , d y , and judge d x ≤1 pixel and d y ≤1 pixel holds. If it holds, the matching is completed; otherwise, continue to execute S3.9 5. An intelligent matching method for virtual and real coordinates of an additive repair robot according to claim 4, characterized in that The established 3D model of the physical object is the same size as the actual physical object.
6. The intelligent matching method for virtual and real coordinates of an additive repair robot according to claim 1, wherein, During the above coordinate matching, place the simulated 3D model of the physical object and the physical object on the virtual operation platform and the actual operation platform respectively. Taking the operation platform as a reference, with the Z-axis regarded as coaxial, adjust the horizontal coordinates for coordinate matching.
7. An intelligent matching system for virtual and real coordinates of an additive repair robot, characterized in that, It includes a feature extraction module, a segmentation prediction module, and a coordinate matching module; The feature extraction module is used to extract features from the physical object image to obtain multi-scale features at different stages of the physical object; In the above feature extraction module, the steps for extracting features from the physical object image include the following steps: extract features from the physical object image through a feature extraction network: S1.1: The input is a feature map of size C j × H j × W j . After passing through 1 convolutional block, the module performs channel splitting at the first-level residual stage to obtain split features P j : P Aj S1.2: Calculate the direct connection features: S1.3: Feature splitting After passing through one convolutional block, it enters the secondary residual stage, and the output feature of the secondary residual stage is P mj : S1.4: Fuse Equation (2) and Equation (3) to obtain multi-scale features F j , where j = 2, 3, 4, In the formula, Cov (.) represents a convolution operation, represents taking the index part of the features after channel splitting, is the channel fusion operation; The segmentation prediction module is used to fuse the multi-scale features at each stage with the upsampled features of the same dimension to obtain the fused features of the image; enhance the fused features to obtain the enhanced integrated features, and use the integrated features for prediction to obtain the segmented image of the physical object image; The coordinate matching module is used to obtain the true coordinates of the physical object by describing the position of the obtained segmented image; establish a simulated 3D model of the physical object, and adjust the position of the simulated 3D model based on the true coordinates of the physical object to achieve coordinate matching between the physical object and the simulated 3D model.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1-6.
Citation Information
Patent Citations
Method and device for realizing virtual-real fusion, electronic equipment and storage medium
CN113409473A
Remote sensing image semantic segmentation method and system based on multi-scale information fusion
CN113780296A