Special yarn three-dimensional reconstruction method and system based on deep learning

Through deep learning methods, a multi-view special yarn image is acquired and an improved MVSNet network is constructed. Combined with the ECA attention mechanism and the AFPN feature fusion module, the problem of low robustness and automation in complex structure processing is solved, and high-precision three-dimensional reconstruction of special yarn is realized, suitable for industrial detection and simulation analysis.

CN120388145APending Publication Date: 2025-07-29SHAOXING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432926.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When dealing with complex structures, traditional three-dimensional reconstruction methods often cause incomplete structural restoration, blurred or distorted surface texture, and fail to effectively align, crop and filter the image, resulting in serious interference from background impurities and light noise during subsequent reconstruction, affecting the quality of the three-dimensional model, and making it difficult to apply to industrial detection or simulation modeling tasks with high accuracy requirements, resulting in low robustness and automation.

Method used

采用基于深度学习的特种纱线三维重建方法,通过获取多张不同视角的特种纱线图像及其元数据,构建改进的MVSNet网络,结合ECA注意力机制和AFPN特征融合模块,进行三维重建并进行误差校正,生成高精度点云。

Benefits of technology

It improves the integrity and robustness of yarn spatial structure reconstruction, enhances the model's ability to pay attention to key channel information, realizes high-precision point cloud generation, meets the needs of industrial detection and simulation analysis, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388145A_ABST
    Figure CN120388145A_ABST
Patent Text Reader

Abstract

The invention provides a special yarn three-dimensional reconstruction method and system based on deep learning, and relates to the technical field of image processing, and the method comprises the steps: obtaining a plurality of special yarn images of different visual angles and corresponding metadata; based on each special yarn image and the metadata, establishing a three-dimensional reconstruction data set of the special yarn; a deep learning three-dimensional reconstruction algorithm based on the improved MVSNet network is constructed, and the deep learning three-dimensional reconstruction algorithm comprises an ECA attention mechanism and an AFPN feature fusion module; based on the three-dimensional reconstruction data set, performing three-dimensional reconstruction on the special yarn through a deep learning three-dimensional reconstruction algorithm, and determining a reconstruction point cloud; and performing error correction on the reconstructed point cloud to determine a final reconstruction result. According to the method, high-precision point cloud generation can be realized, and the robustness of three-dimensional reconstruction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a three-dimensional reconstruction method and system for special yarns based on deep learning. Background Art

[0002] The three-dimensional reconstruction method for special yarns based on deep learning is a technology that automatically restores the spatial geometric structure of special yarns (functional textile yarns with complex three-dimensional morphology and fine structures) from multi-angle images by using technologies such as convolutional neural networks (CNNs), multi-view stereo (MVS), or voxel prediction, which is of great significance for intelligent textile design, quality inspection, and digital simulation.

[0003] As the core of intelligent textile materials, the microscopic structure of special yarns directly affects their performance. Deep learning has powerful feature extraction and pattern recognition capabilities, can more effectively process complex image information, realize automatic and accurate reconstruction of the three-dimensional structure of yarns, provide key data support for subsequent performance analysis, defect detection, and intelligent design, and promote the digital and intelligent transformation of the textile industry.

[0004] However, when traditional three-dimensional reconstruction methods process complex structures, due to insufficient perspective coverage, the structure restoration is often incomplete, the surface texture is blurred or distorted, the images are not effectively aligned, cropped, and filtered, resulting in serious interference from background impurities and light noise in the subsequent reconstruction process, affecting the quality of the three-dimensional model, lacking a systematic error correction scheme, and being difficult to apply to industrial inspection or simulation modeling tasks with high precision requirements, resulting in too low robustness and automation in the three-dimensional reconstruction of special yarns. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, the purpose of the embodiments of the present invention is to provide a three-dimensional reconstruction method for special yarns based on deep learning, which can solve the technical problems that when traditional three-dimensional reconstruction methods process complex structures, due to insufficient perspective coverage, the structure restoration is often incomplete, the surface texture is blurred or distorted, the images are not effectively aligned, cropped, and filtered, resulting in serious interference from background impurities and light noise in the subsequent reconstruction process, affecting the quality of the three-dimensional model, lacking a systematic error correction scheme, and being difficult to apply to industrial inspection or simulation modeling tasks with high precision requirements, resulting in too low robustness and automation in the three-dimensional reconstruction of special yarns.

[0006] In the first aspect of the embodiments of the present invention, a three-dimensional reconstruction method for special yarns based on deep learning is proposed, including:

[0007] S1: Obtain multiple special yarn images from different perspectives and the metadata corresponding to the special yarn images;

[0008] S2: Based on each special yarn image and metadata, establish a 3D reconstruction dataset of the special yarn;

[0009] S3: Construct a deep learning 3D reconstruction algorithm based on the improved MVSNet network, where the deep learning 3D reconstruction algorithm includes an ECA attention mechanism and an AFPN feature fusion module;

[0010] S4: Based on the 3D reconstruction dataset, perform 3D reconstruction on the special yarn through the deep learning 3D reconstruction algorithm to determine the reconstructed point cloud;

[0011] S5: Perform error correction on the reconstructed point cloud to determine the final reconstruction result.

[0012] In the second aspect of the embodiments of the present invention, a 3D reconstruction system of special yarn based on deep learning is proposed, including: a processor and a memory.

[0013] The memory stores a program or instructions that can run on the processor. When the program or instructions are executed by the processor, the steps of the 3D reconstruction method of special yarn based on deep learning as in the first aspect are implemented.

[0014] In the third aspect of the embodiments of the present invention, a readable storage medium is proposed. A program or instructions are stored on the readable storage medium. When the program or instructions are executed by the processor, the steps of the 3D reconstruction method of special yarn based on deep learning as in the first aspect are implemented.

[0015] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:

[0016] In the embodiments of the present invention, through multiple special yarn images from different perspectives, multi-angle coverage is ensured, and the training semantics are enriched with metadata, improving the integrity and robustness of the yarn spatial structure reconstruction. By fusing the attention mechanism and multi-scale feature extraction, the model performance is enhanced. The ECA attention mechanism can improve the network's attention to key channel information, and the AFPN feature fusion module can effectively fuse features at different levels and multi-resolutions. By constructing a deep learning 3D reconstruction algorithm based on the improved MVSNet network, high-precision point cloud generation can be achieved, significantly improving the spatial reduction degree of the yarn model. By performing error correction on the reconstructed point cloud, it is ensured that the results meet the requirements of industrial inspection or simulation analysis, and the evaluation criteria can be further optimized according to the application scenario, with good scalability. Description of the Drawings

[0017] The accompanying drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic flowchart of a three-dimensional reconstruction method for special yarns based on deep learning provided by an embodiment of the present invention;

[0019] Figure 2 is a schematic diagram of the special yarn evenness shooting of a multi-camera system provided by an embodiment of the present invention;

[0020] Figure 3 is a schematic diagram of the network structure of an improved MVSNet deep learning three-dimensional reconstruction algorithm provided by the present invention;

[0021] Figure 4 is a schematic diagram of the structure of a three-dimensional reconstruction system for special yarns based on deep learning provided by an embodiment of the present invention. Detailed Embodiments

[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are exemplary and not intended to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0023] The three-dimensional reconstruction method for special yarns based on deep learning provided by the embodiments of the present invention will be described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0024] Refer to the attached drawings of the specification Figure 1 , which shows a schematic flowchart of a three-dimensional reconstruction method for special yarns based on deep learning provided by an embodiment of the present invention.

[0025] The embodiments of the present invention provide a three-dimensional reconstruction method for special yarns based on deep learning, which may include the following steps:

[0026] Refer to the attached drawings of the specification Figure 2 , which shows a schematic diagram of the special yarn evenness shooting of a multi-camera system provided by the present invention.

[0027] As Figure 2As shown, four cameras are evenly distributed along the moving direction of the yarn, and the lens of each camera is perpendicular to the moving direction of the yarn to ensure that clear images of the yarn can be captured. The angular interval between the cameras is 90 degrees and they are on the same horizontal plane to ensure that the images collected cover all-round features of the yarn while being of the same section of special yarn.

[0028] S1: Obtain multiple images of the special yarn from different perspectives and the metadata corresponding to the special yarn images.

[0029] It should be noted that by obtaining the special yarn images and their corresponding image metadata (such as camera parameters, depth maps, and masks) from multiple perspectives, the morphological features and texture details of the yarn in space can be comprehensively captured, providing rich and accurate input information for subsequent 3D reconstruction and significantly improving the reconstruction quality and stability.

[0030] In a possible implementation, the metadata includes camera parameters, a true depth map, and a mask.

[0031] Among them, the camera parameters are used to describe the geometric and optical characteristics during the camera imaging process, mainly including internal parameters and external parameters. The internal parameters describe the internal structure of the camera, such as focal length, principal point position, pixel scaling factor, etc., and are used to project 3D points onto the image plane. The external parameters describe the position and orientation of the camera in the world coordinate system (i.e., rotation matrix and translation vector) and are used to transform points in the world coordinate system to the camera coordinate system. The true depth map is a grayscale image, where each pixel value represents the actual depth at that position in the image. The mask is a binary image that marks the valid area and the background area in the image. Generally, white (pixel value of 1) represents the valid area, such as the yarn body, and black (pixel value of 0) represents the invalid area or the background.

[0032] In the present invention, according to the characteristics of the target area, the GrabCut method is selected to generate a mask image for the image mask, and post-processing is performed on the generated mask, where the pixel value of the target area is 255 (white) and the pixel value of the background area is 0 (black).

[0033] S1 is specifically as follows:

[0034] Through a multi-camera system, obtain multiple images of the special yarn from different perspectives and the corresponding metadata.

[0035] The multi-camera system adopts an annular configuration and includes multiple imaging units.

[0036] Each imaging unit is evenly distributed on the circumference of the same plane.

[0037] Specifically, to achieve dynamic continuous acquisition of multi-angle characteristic yarn images, cameras are arranged at multiple perspectives along the yarn movement direction. The lens of each camera is perpendicular to the yarn movement direction to obtain clear images. Through the software synchronization mechanism PTP protocol, it is ensured that multiple cameras capture images at the same time point, avoiding misalignment of multi-perspective images caused by time deviation. The image stream is transmitted in real time to a computer or server through the GigE interface.

[0038] S2: Based on each special yarn image and metadata, establish a three-dimensional reconstruction dataset of the special yarn.

[0039] It should be noted that by fusing multiple special yarn images from different perspectives and their corresponding metadata to construct a three-dimensional reconstruction dataset with a unified structure and complete annotation, it not only provides multi-source input information for the deep learning model, but also improves the data diversity and accuracy during training, thereby enhancing the generalization ability and reconstruction effect of the model.

[0040] In a possible implementation manner, S2 specifically includes:

[0041] S201: Perform optimization operations on the special yarn images to determine the optimized images, where the optimization operations include alignment operations and cropping operations.

[0042] S202: Perform median filtering on the optimized images:

[0043] f output x,y=median f input x+i,y+j|i,j∈W

[0044] where, f output (x,y) represents the pixel at position (x,y) in the output image, median represents the median operation, f input (x+i,y+j) represents the pixel at position (x+i,y+j) in the input image, both i and j represent relative offsets, and W represents the neighborhood window.

[0045] S203: Construct a three-dimensional reconstruction dataset based on the optimized images after median filtering.

[0046] In the present invention, the three-dimensional reconstruction dataset of the special yarn contains 500 scenes each containing a section of the special yarn sliver. Each scene contains 24 yarn pictures in different directions, with a total of 12,000 yarn images. The image resolution is 1600×1200, including 400 scenes as the training set, 50 scenes with 50 training samples as the validation set, and 50 scenes as the test set for subsequent model training and performance evaluation.

[0047] Refer to the attached drawings of the specification Figure 3, showing a schematic diagram of the network structure of an improved MVSNet deep learning three-dimensional reconstruction algorithm provided by the present invention.

[0048] As Figure 3 shown, on the basis of the original MVSNet, the improved MVSNet network structure introduces an AFPN feature fusion module and an ECA attention mechanism module to significantly improve the multi-scale feature extraction ability and the attention to key features. In this algorithm, the input is multiple images taken from different perspectives of the same scene, including a reference view and multiple source views. First, through the AFPN progressive feature fusion mechanism, the low-level features and high-level features are gradually fused to ensure that the features at different levels are fused in an appropriate proportion, thereby enhancing the model's detection ability for multi-scale targets. Subsequently, the ECA attention mechanism is introduced to further enhance the dependency relationship between feature channels. The ECA module first performs global average pooling on the input feature map to compress the spatial information into the channel dimension, then captures the local dependency relationship between channels through one-dimensional convolution operations, and finally generates channel weights through the Sigmoid activation function. This lightweight architecture significantly enhances the model's ability to focus on key features without adding excessive computational costs, especially performing well in dealing with fine structure features.

[0049] S3: Construct a deep learning three-dimensional reconstruction algorithm based on the improved MVSNet network, where the deep learning three-dimensional reconstruction algorithm includes an ECA attention mechanism and an AFPN feature fusion module.

[0050] Among them, the MVSNet network is a multi-view stereo three-dimensional reconstruction network based on deep learning, which generates a dense point cloud by constructing a cost volume and predicting a depth map, and is widely used in three-dimensional modeling tasks. The ECA attention mechanism is a lightweight channel attention mechanism that captures the correlation between channels through local one-dimensional convolution, improving the neural network's ability to focus on key information without significantly increasing the computational overhead. The AFPN feature fusion module is a structure for multi-scale feature fusion, which can achieve efficient information flow between feature maps of different resolutions, thereby enhancing the model's recognition ability for different scale structures, especially suitable for targets with complex textures.

[0051] It should be noted that by constructing a deep learning three-dimensional reconstruction algorithm based on the improved MVSNet, fusing the ECA attention mechanism and the AFPN feature fusion module, the model's ability to extract multi-scale structure features and its response ability to key channel information are significantly enhanced, effectively improving the three-dimensional modeling accuracy and detail restoration quality of the complex morphology of special yarns.

[0052] In a possible implementation manner, S3 is specifically:

[0053] Based on the MVSNet network, the ECA attention mechanism and the AFPN feature fusion module are introduced to construct a deep learning three-dimensional reconstruction algorithm.

[0054] The AFPN feature fusion module is used to extract multi-scale features at different image resolutions:

[0055]

[0056] Among them, represents the fused feature, represents the feature from the first layer to the l-th layer, represents the feature from the second layer to the l-th layer, represents the feature from the third layer to the l-th layer, and represent the weights of features at different levels.

[0057] The principle of the ECA attention mechanism is specifically as follows:

[0058] w = σ(C1D k (y))

[0059] Among them, w represents the output of the ECA attention mechanism, σ represents the Sigmoid activation function, and C1D k represents the one-dimensional convolution operation of the k-th convolution kernel.

[0060] In the present invention, adding ECA to the MVSNet network can help the model more accurately focus on the fine structural features in the special yarn. The specific steps are as follows: (1) Perform global average pooling on the input image. (2) Perform a one-dimensional convolution operation with a convolution kernel size of k, and obtain the weights w of each channel through the Sigmoid activation function.

[0061] In a possible implementation manner, the processing process of the MVSNet network for the three-dimensional reconstruction dataset specifically includes:

[0062] Adopt a joint feature extraction strategy based on a convolutional neural network to extract features from the three-dimensional reconstruction dataset and determine the feature map.

[0063] Among them, the joint feature extraction strategy is a way to extract shared semantic features from multiple images using a convolutional neural network (CNN), which can capture the structural similarity between different perspectives and generate comparable feature maps for analysis. The feature map is an intermediate representation after extracting image information in the neural network, retaining the key spatial structure and semantic information for subsequent depth estimation.

[0064] According to the camera parameters, map the feature map into the camera coordinates to construct the cost volume:

[0065]

[0066] Among them, C represents the cost volume, M represents the variance metric, and V i represents the i-th feature map, represents the average value of the feature map, where i = 0, 1, …, N, and N represents the total number of feature maps.

[0067] Among them, the cost volume is a multi-dimensional tensor used to store the pixel matching costs between multiple views, reflecting the matching credibility of each pixel under different depth hypotheses, and is constructed in the camera space coordinates.

[0068] The 3D U-Net is used to regularize the cost volume.

[0069] Among them, the 3D U-Net is a three-dimensional convolutional network structure used to perform feature fusion and regularization on the cost volume, improving the continuity and accuracy of depth prediction.

[0070] Calculate the depth values of all pixels in the regularized cost volume to determine the depth map:

[0071]

[0072] Among them, D represents the depth value, d represents the depth sampling value, P() represents the probability value, and d min represents the minimum value of depth sampling, and d max represents the maximum value of depth sampling.

[0073] Specifically, the depth sampling value refers to the distance information from a certain point in the three-dimensional scene to the sensor or camera collected by a laser scanner. The depth sampling value is usually represented in numerical form and can be an absolute distance (such as meters, centimeters, etc.) or a relative distance (such as a value normalized between 0 and 1). These values reflect the geometric shape and spatial distribution of the object surface in the scene.

[0074] The Gipuma tool is used to optimize the depth map to generate a three-dimensional point cloud model.

[0075] Among them, the Gipuma tool is a dense point cloud generation tool used to extract three-dimensional coordinate points from the depth map, with point cloud filtering and refinement functions, and is suitable for the post-processing stage of multi-view reconstruction tasks.

[0076] It should be noted that by extracting multi-view image features through a convolutional neural network, constructing a cost volume based on camera parameters, and then using a 3D U-Net for regularization processing, high-precision depth map estimation is achieved. Combining with the Gipuma tool to generate a high-quality point cloud model not only improves the density and continuity of 3D reconstruction, but also enhances the ability to restore the detailed structure of complex yarns, making the final model more realistic and stable.

[0077] In the present invention, the input feature map adopts a cascaded cost volume structure, and the resolution of the feature map gradually increases as the stage progresses. Specifically, the resolution of the feature map in the first stage is 1 / 16 of the original input image, the resolution of the feature map in the second stage is 1 / 4 of the original input image, and the resolution of the feature map in the third stage is the same as that of the original input image, which is 1. This progressive design from low resolution to high resolution can effectively reduce the consumption of computing resources while gradually improving the accuracy of depth estimation.

[0078] In a possible implementation manner, the size of the regularized cost volume is specifically:

[0079]

[0080] where W in 、H in and D in respectively represent the width, height and depth of the cost volume, W out 、H out and D out respectively represent the width, height and depth of the regularized cost volume, s represents the sliding step size, p represents the padding value, and w, h and d respectively represent the width, height and depth of the convolutional kernel.

[0081] In the present invention, this parameter relationship formula assumes that the scale of the input layer is W in ×H in ×D in ×C in , and the scale of the output layer is W out ×H out ×D out ×C out . The sliding step size defines the interval at which the convolutional kernel moves on the input data. The size of the step directly affects the size of the output feature map: the larger the step size, the smaller the size of the output feature map. The padding value refers to the set of values added around the edge of the input data during the convolution operation to control the size of the output feature map. These values are usually zero, and the main purpose of padding is to maintain the spatial dimension of the input data during the convolution process, especially when using a larger convolutional kernel or when it is desired that the size of the output feature map is the same as the input.

[0082] In a possible implementation, after S3, it further includes:

[0083] Input the three-dimensional reconstruction dataset into the deep learning three-dimensional reconstruction algorithm for training until the loss function value is less than the preset loss function value.

[0084] It should be noted that by continuously training the constructed three-dimensional reconstruction dataset in the deep learning three-dimensional reconstruction algorithm until the loss function converges to the preset threshold, it ensures that the model has stronger fitting ability and generalization performance in terms of the structure, depth, and texture restoration of special yarns, thereby significantly improving the accuracy and stability of three-dimensional reconstruction.

[0085] In the present invention, the Adaptive Moment Estimation method Adam is used as the optimizer for training. The number of input views is 4, the picture input size is 640×512, the depth plane is selected as 48, the initial learning rate is set to 0.001, the Batch_size is set to 64, and the Epoch is set to 16.

[0086] It should be noted that those skilled in the art can set the size of the preset loss function value according to the actual situation, and the present invention does not make any limitations.

[0087] The specific calculation formula of the loss function value is:

[0088]

[0089] Among them, Loss represents the loss function value, D l (x) represents the predicted depth of pixel x after being processed by the l-th layer, x v represents the set of pixels, λ l represents the loss weight of the l-th layer, |||| represents the norm symbol, represents the true depth of pixel x after being processed by the l-th layer, l = 1, 2, 3.

[0090] S4: Based on the three-dimensional reconstruction dataset, perform three-dimensional reconstruction on the special yarn through the deep learning three-dimensional reconstruction algorithm to determine the reconstructed point cloud.

[0091] It should be noted that by using the deep learning three-dimensional reconstruction algorithm and the constructed three-dimensional reconstruction dataset to perform three-dimensional reconstruction on the special yarn, an accurate reconstructed point cloud is generated. Through the optimization of the deep learning model, it can better capture the details and complex structures of the yarn, thereby improving the reconstruction accuracy and making the final three-dimensional point cloud more realistic and rich in details, which is suitable for high-precision three-dimensional modeling tasks.

[0092] S5: Perform error correction on the reconstructed point cloud to determine the final reconstruction result.

[0093] It should be noted that by performing error correction on the point cloud data obtained through reconstruction, and combining evaluation indicators such as accuracy, integrity, and overall score to identify and correct the deviations and missing parts in the point cloud, the overall quality and structural restoration degree of the 3D model can be effectively improved, ensuring that the final reconstruction results meet the application requirements in terms of accuracy, continuity, and integrity.

[0094] In a possible implementation manner, S5 is specifically:

[0095] By calculating the accuracy, integrity, and overall score of the reconstructed point cloud, error correction is performed on the reconstructed point cloud to determine the final reconstruction result.

[0096] In a possible implementation manner, the accuracy is specifically:

[0097]

[0098] Among them, Accuracy represents the accuracy of the reconstructed point cloud, R represents the point set of the reconstructed point cloud, min represents minimization, r represents a point in the reconstructed point cloud, g represents a point in the true value point cloud, and G represents the point set of the true value point cloud.

[0099] The integrity is specifically:

[0100]

[0101] Among them, Completeness represents the integrity of the reconstructed point cloud.

[0102] The overall score is specifically:

[0103]

[0104] Among them, Ovaerall represents the overall evaluation of the reconstructed point cloud.

[0105] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:

[0106] The embodiments of the present invention ensure multi-angle coverage through special yarn images from multiple different perspectives, and enrich the training semantics with metadata, improving the integrity and robustness of the yarn spatial structure reconstruction. By fusing the attention mechanism and multi-scale feature extraction, the model performance is enhanced. The ECA attention mechanism can improve the network's attention to key channel information, and the AFPN feature fusion module can effectively fuse features at different levels and multi-resolutions. By constructing a deep learning 3D reconstruction algorithm based on the improved MVSNet network, high-precision point cloud generation can be achieved, significantly improving the spatial restoration degree of the yarn model. By performing error correction on the reconstructed point cloud, it is ensured that the results meet the requirements of industrial inspection or simulation analysis, and the evaluation criteria can be further optimized according to the application scenario, having good scalability.

[0107] Refer to the attached instruction manual Figure 4 , which shows a schematic structural diagram of a three-dimensional reconstruction system for special yarns based on deep learning provided by an embodiment of the present invention.

[0108] An embodiment of the present invention provides a three-dimensional reconstruction system 20 for special yarns based on deep learning, including: a processor 201 and a memory 202;

[0109] The memory 202 stores programs or instructions that can run on the processor 201. When the programs or instructions are executed by the processor 201, the steps of the above-mentioned three-dimensional reconstruction method for special yarns based on deep learning are implemented, and the same technical effects can be achieved. To avoid repetition, the present invention will not be elaborated herein.

[0110] It should be understood that the processor 201 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0111] It should also be understood that the memory 202 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0112] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0113] It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or posterior. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0114] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0115] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0116] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0117] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0119] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0120] An embodiment of the present invention provides a readable storage medium, including: programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, the steps of the above-mentioned three-dimensional reconstruction method of special yarn based on deep learning are implemented, and the same technical effects can be achieved. To avoid repetition, the present invention will not be described in detail again.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A three-dimensional reconstruction method for special yarns based on deep learning, characterized in that, Including: S1: Obtain special yarn images from multiple different perspectives and the metadata corresponding to the special yarn images; S2: Based on each of the special yarn images and the metadata, establish a 3D reconstruction dataset for the special yarn; S3: Construct a deep learning 3D reconstruction algorithm based on the improved MVSNet network, where the deep learning 3D reconstruction algorithm includes an ECA attention mechanism and an AFPN feature fusion module; S4: Based on the 3D reconstruction dataset, perform 3D reconstruction on the special yarn through the deep learning 3D reconstruction algorithm to determine the reconstructed point cloud; S5: Perform error correction on the reconstructed point cloud to determine the final reconstruction result.

2. The three-dimensional reconstruction method of special yarn based on deep learning according to claim 1, characterized in that The metadata includes camera parameters, a ground truth depth map, and a mask; The specific content of S1 is: Through a multi-camera system, obtain special yarn images from multiple different perspectives and the metadata corresponding to the special yarn images; The multi-camera system adopts a circular configuration and includes multiple imaging units; Each of the imaging units is evenly distributed on the circumference of the same plane.

3. The three-dimensional reconstruction method of special yarn based on deep learning according to claim 1, characterized in that The specific content of S2 includes: S201: Perform optimization operations on the special yarn images to determine optimized images, where the optimization operations include alignment operations and cropping operations; S202: Perform median filtering on the optimized images; f output x,y = median f input x + i,y + j | i,j ∈ W; Among them, f output (x, y) represents the pixel of the output image at the position (x, y), median represents the median operation, and f input (x + i, y + j) represents the pixel of the input image at the position (x + i, y + j), where both i and j represent relative offsets, and W represents the neighborhood window; S203: Based on the optimized images after median filtering, construct the 3D reconstruction dataset.

4. The three-dimensional reconstruction method of special yarns based on deep learning according to claim 1, characterized in that, The specific content of S3 is: Based on the MVSNet network, introduce an ECA attention mechanism and an AFPN feature fusion module to construct the deep learning 3D reconstruction algorithm; The AFPN feature fusion module is used to extract multi-scale features at different image resolutions; Among them, represents the fused feature, represents the features from the first layer to the l-th layer, represents the features from the second layer to the l-th layer, represents the features from the third layer to the l-th layer, and represent the weights of features at different levels; The principle of the ECA attention mechanism is specifically: w = σ(C1D k (y)); Among them, w represents the output of the ECA attention mechanism, σ represents the Sigmoid activation function, and C1D k represents the one-dimensional convolution operation of the k-th convolution kernel.

5. The three-dimensional reconstruction method of special yarns based on deep learning according to claim 4, characterized in that, The processing process of the MVSNet network for the 3D reconstruction dataset specifically includes: Adopt a joint feature extraction strategy based on a convolutional neural network to perform feature extraction on the 3D reconstruction dataset to determine a feature map; According to the camera parameters, map the feature map into the camera coordinates to construct a cost volume; Among them, C represents the cost volume, M represents the variance metric, and V i represents the i-th feature map, represents the average value of the feature map, where i = 0, 1, …, N, and N represents the total number of feature maps; Adopt a 3D U-Net to regularize the cost volume to determine a regularized cost volume; Calculate the depth values of all pixels in the regularized cost volume to determine a depth map; where D represents the depth value, d represents the depth sampling value, P() represents the probability value, d min represents the minimum value of depth sampling, d max represents the maximum value of depth sampling; Through the Gipuma tool, optimize the depth map to generate a 3D point cloud model.

6. The three-dimensional reconstruction method of special yarn based on deep learning according to claim 5, characterized in that The size of the regularized cost volume is specifically: Among them, W in , H in and D in respectively represent the width, height and depth of the cost volume, W out , H out and D out respectively represent the width, height and depth of the regularized cost volume, s represents the sliding step size, p represents the padding value, and w, h and d respectively represent the width, height and depth of the convolution kernel.

7. The three-dimensional reconstruction method of special yarn based on deep learning according to claim 1, characterized in that After S3, it further includes: Input the 3D reconstruction dataset into the deep learning 3D reconstruction algorithm for training until the loss function value is less than a preset loss function value; The calculation formula of the loss function value is specifically: Among them, Loss represents the loss function value, D l (x) represents the predicted depth of pixel x after being processed by the l-th layer, and x v represents the set of pixels, and λ l represents the loss weight of the l-th layer, |||| represents the norm symbol, represents the true depth of pixel x after being processed by the l-th layer, where l = 1, 2, 3.

8. The three-dimensional reconstruction method of special yarns based on deep learning according to claim 1, characterized in that The specific content of S5 is: By calculating the accuracy, completeness, and overall score of the reconstructed point cloud, perform error correction on the reconstructed point cloud to determine the final reconstruction result.

9. The three-dimensional reconstruction method of special yarn based on deep learning according to claim 8, characterized in that, The accuracy is specifically: Where Accuracy represents the accuracy of the reconstructed point cloud, R represents the point set of the reconstructed point cloud, min represents minimization, r represents a point in the reconstructed point cloud, g represents a point in the ground truth point cloud, and G represents the point set of the ground truth point cloud; The completeness is specifically: Among them, Completeness represents the completeness of the reconstructed point cloud; The specific overall score is as follows: Among them, Ovaerall represents the overall evaluation of the reconstructed point cloud.

10. A three-dimensional reconstruction system for special yarns based on deep learning, characterized in that, It includes: A processor and a memory; The memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the 3D reconstruction method for special yarns based on deep learning according to any one of claims 1 to 9 are implemented.