Processing method for capturing object image in real time based on computer vision technology
By adopting computer vision technology in image processing, combining Gaussian filters, histogram equalization, hybrid architecture feature extraction, full convolutional network segmentation, multi-view geometry three-dimensional recovery and Poisson surface reconstruction, the problem of insufficient noise removal and feature extraction accuracy in image processing is solved, high-precision object segmentation and three-dimensional model generation are achieved, and the rendering quality of digital twins is improved.
Patent Information
- Application Number
- CN202510425797.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art cannot effectively remove all types of noise in image processing, especially when enhancing contrast, new distortion will be introduced. In complex scenarios, feature extraction and segmentation accuracy is not accurate enough, affecting subsequent reconstruction operations.
Using a computer vision technology-based processing method, the object image is captured in real time through a USB camera or a webcam, the noise is removed using a Gaussian filter and the contrast is enhanced by histogram equalization. Then, image features are extracted using the hybrid architecture of Inception V3 and ResNet-50, pixel-level segmentation is performed by combining a full convolutional network, and then the three-dimensional information is restored using the multi-view geometry method, and a smooth three-dimensional surface model is generated through the Poisson surface reconstruction algorithm. Finally, the realism and rendering quality of digital twins are improved through style transfer and global lighting technologies.
It realizes high-precision and high-efficiency object segmentation, and the generated three-dimensional model has high realistic and rendering quality, meeting the needs of real-time rendering and interaction.
Smart Images

Figure CN120047458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and specifically to a processing method for real-time capturing of object images based on computer vision technology. Background Art
[0002] The processing of captured object images refers to the process of using computer vision technology to capture images of objects and analyzing, processing, and optimizing these images through a series of algorithms to achieve purposes such as recognition, classification, measurement, or three-dimensional reconstruction.
[0003] When performing image processing currently, it is impossible to effectively remove all types of noise, especially new distortions will be introduced while enhancing the contrast. Especially in complex scenes, the extraction and segmentation methods often cannot achieve high precision, resulting in inaccurate segmentation results and affecting subsequent reconstruction operations. Therefore, we propose a processing method for real-time capturing of object images based on computer vision technology to solve the problems raised above. Summary of the Invention
[0004] The purpose of the present invention is to provide a processing method for real-time capturing of object images based on computer vision technology to solve the problems of limited image preprocessing effect and inaccurate feature extraction and segmentation accuracy in the current market proposed in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A processing method for real-time capturing of object images based on computer vision technology includes the following steps:
[0007] Step 1: Use a USB camera and a network camera to capture object images in real time, and transmit the image data to the server side in real time through TCP / IP for processing;
[0008] Step 2: Use a Gaussian filter to remove noise, histogram equalization to enhance the contrast, extract image features through a hybrid architecture based on Inception V3 and ResNet-50, the hybrid architecture optimizes feature extraction through a multi-scale attention mechanism, use a fully convolutional network for pixel-level segmentation, divide the image into different object regions, and then perform post-processing on the segmentation results;
[0009] Step 3: Use multi-view geometry to recover the three-dimensional information of the object from two-dimensional images from different perspectives, generate three-dimensional point cloud data, use a three-dimensional surface reconstruction algorithm based on the Poisson equation to convert the three-dimensional point cloud data into a smooth three-dimensional surface model, and at the same time combine the segmentation results output by the CutX model to process noise and missing data;
[0010] Step 4: Ensure the visual consistency of the digital twin from different perspectives through style transfer, simulate the lighting and shadow effects in the real world using global illumination, use GLSL language to edit the material of the digital twin, write a shader program, integrate the digital twin model into Unity for real-time rendering and interaction;
[0011] Step 5: Transmit the processed digital twin model to the client device through the network interface, convert the position, posture, and actions of the digital twin model into controllable signals, and then drive the external hardware to perform corresponding actions.
[0012] Preferably, in Step 2, the specific steps of using Inception to extract image features are as follows:
[0013] Introduce the numpy and tensorflow.keras modules. Numpy is used for numerical calculations, and tensorflow.keras is a tool for building and training neural networks. Use the trained Inception V3 model and prepare the image to be analyzed. After loading the image, resize its size to 299x299 pixels, convert the image into an array and add dimensions, preprocess the image to match the data format during model training, run the image data through the model using the predict function and output the features, flatten the features, then use a graph to display these features, and finally, save the features as a.npy file.
[0014] Preferably, the specific steps of removing small noise through morphological operations and extracting the object boundary using an edge detection algorithm are as follows:
[0015] Step 1: Read the segmentation map output by the FCN, which is a two-dimensional array where each element represents the class label of a pixel. Create a 3x3 square structuring element, and the value of each pixel represents its class label;
[0016] Step 2: Apply the erosion operation to the segmentation result to remove small noise, apply the dilation operation to the eroded result to restore the main shape of the object, and iterate the erosion and dilation operations multiple times until the small noise is effectively removed;
[0017] Step 3: When the seg_map is in color, convert it to a grayscale image for edge detection. Use the Canny algorithm to extract the edges of the object. When finer edges are required, use the Zhang-Suen algorithm. Superimpose the result of edge detection on the original segmentation map to clearly display the boundary of the object.
[0018] Preferably, the specific steps of using multi-view geometry to recover the three-dimensional information of an object from two-dimensional images from different perspectives and generate three-dimensional point cloud data are as follows:
[0019] Use the segmentation results output by the CutX model to assist in finding the pixel points corresponding to the same object in multi-view images. Utilize the previously calibrated camera parameters to convert the corresponding points in the two-dimensional image into points in three-dimensional space. For each pair of matched points, use the principle of triangulation to calculate their positions in three-dimensional space. Pool all the three-dimensional points obtained through triangulation to form the three-dimensional point cloud of the object. Finally, perform optimization processing on the generated point cloud.
[0020] Preferably, the specific steps of using the Poisson surface reconstruction algorithm to convert the three-dimensional point cloud data into a smooth three-dimensional surface model are as follows:
[0021] Input the cleaned point cloud data into the Poisson surface reconstruction algorithm, read the point cloud data. Before performing Poisson reconstruction, estimate a normal for each point in the point cloud to create a normal estimator, merge the point cloud and the normals, perform Poisson surface reconstruction, save or display the reconstructed mesh. After Poisson surface reconstruction is completed, optimize the mesh quality through mesh smoothing and simplification. When the original image contains texture information, map these textures onto the reconstructed surface.
[0022] Preferably, in step two, the training of the fully convolutional network adopts a composite loss function, which is composed of 60% Dice loss and 40% cross-entropy loss weighted. The specific formula is:
[0023] L total = 0.6·L Dice + 0.4·L CE
[0024] where, L Dice optimizes the boundary accuracy by calculating the overlap degree between the predicted segmentation and the ground truth label, and L CE optimizes the class consistency through the pixel-level classification error;
[0025] The weight assignment mechanism of the multi-scale attention module is: assign a weight of 0.3 to the low-level features (res2c) of Inception V3, a weight of 0.5 to the middle-level features (res3c), and a weight of 0.2 to the high-level features (res5c) to balance local details and global semantics.
[0026] Preferably, in step four, the total loss function of style transfer is defined as:
[0027] L total = α||F res5c (I g ) - F res5c (I c )|| 2 + β∑ l ||G(F l (Ig )) - G(F l (I s )) || 2
[0028] Among them, I g is the generated image, I c is the content image, I s is the style image;
[0029] F res5c represents the content features extracted from the res5c layer in ResNet101, F l are the style features extracted from the res2c / res3c / res4c / res5c layers;
[0030] G(·) is the Gram matrix, which is used to quantify the statistical distribution of style features;
[0031] β is the weight coefficient.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] The present invention adopts Gaussian filter and histogram equalization technology to effectively remove image noise and enhance contrast. It uses the Inception network to extract image features and combines with the fully convolutional network for pixel-level segmentation, achieving high-precision and high-efficiency object segmentation. Then, through the multi-view geometry method, the three-dimensional information of the object is recovered from two-dimensional images with different perspectives to generate three-dimensional point cloud data. Finally, the global illumination technology is applied in the digital twin model to simulate the lighting and shadow effects in the real world, improving the realism and rendering quality of the digital twin.
[0034] The present invention uses GLSL language to edit materials and write shader programs, providing highly customizable visual effects for the digital twin. Integrating the digital twin model into Unity realizes real-time rendering and interaction.
[0035] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will be readily apparent by referring to the drawings and the following detailed description. Brief Description of the Drawings
[0036] Figure 1 is the flowchart of the method for processing images captured in real time by the present invention;
[0037] Figure 2 is the flowchart of testing and evaluation in the present invention;
[0038] Figure 3It is a flowchart for pixel-level segmentation of pictures using FCN in the present invention;
[0039] Figure 4 It is a flowchart for Poisson surface reconstruction in the present invention. Specific embodiments
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0041] Embodiment 1
[0042] Please refer to Figures 1 - 4 , a method for processing real-time capture of object images based on computer vision technology, including the following steps:
[0043] Step 1: Use a USB camera and a network camera to perform real-time capture of object images, and transmit the image data to the server side for processing in real time through TCP / IP;
[0044] Step 2: Use a Gaussian filter to remove noise, perform histogram equalization to enhance contrast, extract image features through a hybrid architecture based on Inception V3 and ResNet-50. The hybrid architecture optimizes feature extraction through a multi-scale attention mechanism, uses a fully convolutional network for pixel-level segmentation, divides the image into different object regions, and then performs post-processing on the segmentation results;
[0045] Step 3: Use multi-view geometry to recover the three-dimensional information of the object from two-dimensional images from different perspectives, generate three-dimensional point cloud data, use a three-dimensional surface reconstruction algorithm based on the Poisson equation to convert the three-dimensional point cloud data into a smooth three-dimensional surface model, and at the same time combine the segmentation results output by the CutX model to process noise and missing data;
[0046] Step 4: Ensure the visual consistency of the digital twin from different perspectives through style transfer, use global illumination to simulate the lighting and shadow effects in the real world, use GLSL language to edit the material of the digital twin, write a shader program, integrate the digital twin model into Unity for real-time rendering and interaction;
[0047] Step 5: Transmit the processed digital twin model to the client device through a network interface, convert it into a controllable signal according to the position, posture, and actions of the digital twin model, and then drive the external hardware to perform corresponding actions.
[0048] In step 2, the specific steps of using Inception to extract image features are as follows:
[0049] Import the numpy and tensorflow.keras modules. Numpy is used for numerical calculations, and tensorflow.keras is a tool for building and training neural networks. Use the trained Inception V3 model and prepare the image to be analyzed. After loading the image, resize its size to 299x299 pixels, convert the image into an array and add dimensions, preprocess the image to match the data format during model training, run the image data through the model using the predict function, and output the features. Flatten the features, then use a graph to display these features. Finally, save the features as a.npy file.
[0050] Specifically, the histogram equalization formula used is:
[0051] s(k) = T(r(k))
[0052] where r(k) is the gray level of the original image, s(k) is the gray level after equalization, and T is the cumulative distribution function (CDF).
[0053] In a convolutional neural network, the formula for the convolution operation is:
[0054]
[0055] where f is the input image, g is the convolution kernel, and * represents the convolution operation.
[0056] ReLU function: f(x) = max(0, x)
[0057] The formula for the pooling operation is: p(s, y) = max f(x + i, y + j)
[0058] where p is the feature after pooling and f is the feature before pooling.
[0059] The formula for the fully connected layer is:
[0060] where y is the output, w i is the weight, x i is the input, b is the bias, and σ is the activation function.
[0061] In step 2, the specific steps of using a fully convolutional network for pixel-level segmentation and dividing the image into different object regions are as follows:
[0062] First, construct the FCN model. Increase the spatial resolution of the feature map output by the Inception network through a transposed convolutional layer to make it close to the size of the original image. Merge the feature map of the middle layer in the Inception network with the output of the upsampling layer through skip connections to retain detailed information. Add a 1x1 convolutional layer at the end of the network to output the class prediction for each pixel.
[0063] Step 2: Take the feature map extracted by the Inception network as input and feed it into the constructed FCN model.
[0064] Step 3: The feature map is gradually restored to the resolution of the original image through the upsampling layer. Through skip connections, perform element-wise addition of the output of the upsampling layer and the feature map of the middle layer in the Inception network to fuse features at different scales. Finally, output the class prediction for each pixel through a 1x1 convolutional layer.
[0065] Step 4: Apply the Softmax function to the class prediction output by the FCN to convert the prediction values into a probability distribution, and each pixel will obtain a probability belonging to each class.
[0066] Step 5: For each pixel, select the class with the highest probability as the final class label of the pixel. According to the class label of each pixel, construct a segmentation map with the same size as the original image, where the value of each pixel represents its corresponding class. Use morphological operations to smooth the segmentation boundary and output the final segmentation map.
[0067] The specific steps for removing small noise through morphological operations and extracting the object boundary using an edge detection algorithm are as follows:
[0068] Step 1: Read the segmentation map output by the FCN, which is a two-dimensional array, where each element represents the class label of a pixel. Create a 3x3 square structuring element, and the value of each pixel represents its class label.
[0069] Step 2: Apply the erosion operation to the segmentation result to remove small noise, and then apply the dilation operation to the eroded result to restore the main shape of the object. Iterate the erosion and dilation operations multiple times until the small noise is effectively removed.
[0070] Step 3: When the seg_map is colored, it needs to be converted to a grayscale image for edge detection. Use the Canny algorithm to extract the edges of the object. When finer edges are required, use the Zhang-Suen algorithm. Superimpose the result of edge detection on the original segmentation map to clearly display the object boundary.
[0071] Specifically, the relationship between the point X in the three-dimensional world projected onto the point x on the two-dimensional image plane is as follows:
[0072]
[0073] Among them, f is the camera focal length, Z is the distance from the object to the camera, X is the point in the world coordinate system, and x is the point on the image plane.
[0074] To transform the three-dimensional point from the world coordinate system to the camera coordinate system, use the rotation matrix R and the translation vector t: X camera = R·X world + t
[0075] To project the point in the camera coordinate system onto the image plane, use the camera intrinsic matrix K: x = K·[X comera ; 1] T
[0076] Among them, K contains the focal length f and the principal point coordinates (cx, cy).
[0077] The specific steps to recover the three-dimensional information of the object from the two-dimensional images from different perspectives using the multi-view geometry method and generate the three-dimensional point cloud data are as follows:
[0078] Use the segmentation result output by the CutX model to assist in finding the pixel points corresponding to the same object in the multi-view images. Utilize the previously calibrated camera parameters to convert the corresponding points in the two-dimensional images into points in the three-dimensional space. For each pair of matching points, use the triangulation principle to calculate their positions in the three-dimensional space. Pool all the three-dimensional points obtained through triangulation to form the three-dimensional point cloud of the object. Finally, perform optimization processing on the generated point cloud.
[0079] The specific steps to convert the three-dimensional point cloud data into a smooth three-dimensional surface model using the Poisson surface reconstruction algorithm are as follows:
[0080] Input the cleaned point cloud data into the Poisson surface reconstruction algorithm, read the point cloud data. Before performing Poisson reconstruction, estimate a normal for each point in the point cloud to create a normal estimator, merge the point cloud and the normal, perform Poisson surface reconstruction, save or display the reconstructed mesh. After the Poisson surface reconstruction is completed, optimize the mesh quality through mesh smoothing and simplification. When the original image contains texture information, map these textures onto the reconstructed surface.
[0081] In step four, the specific steps to perform style transfer using the ResNet101 model are as follows:
[0082] Step 1: Load the pre-trained ResNet101 model from TensorFlow, and load the original image to be kept and the image of the target style.
[0083] Step 2: Measure the differences in content features between the generated image and the content image, measure the differences in style features between the generated image and the style image, maintain the smoothness of the generated image, and reduce noise;
[0084] Step 3: Use the ResNet101 model to extract the features of the content image and the style image, use res5c to extract the content loss, use res2c, res3c, res4c, res5c to extract the style loss, initialize the generated image as the content image or a random noise image. For the generated image, calculate the content features and style features through ResNet101, and calculate the content loss, style loss, and total variation loss. Use L-BFGS and Adam to update the pixel values of the generated image according to the gradients of the loss function. Propagate the generated image forward through ResNet101 to obtain the features until the preset number of iterations is reached, and the optimization process is completed. Save the final generated image.
[0085] Specifically, in Step 4, the reflection equation is:
[0086] L O (P, W) = L e + ∫ Ω f r (P, W′, W) L i (P, W′) (W′·n) dw
[0087] where: L O is the outgoing radiance from point P in the direction of W, L e is the self-emission radiance of point P, f r is the BRDF (Bidirectional Reflectance Distribution Function), L i is the incident radiance from point P in the direction of W′, W′ is a direction on the hemisphere Ω , n is the surface normal of point P, and dw is the infinitesimal solid angle in the direction of W′.
[0088] The formula used for the bidirectional reflectance distribution function is:
[0089] where ρ is the reflectivity of the surface, w i is the incident direction, w o is the reflection direction, and π is the normalization factor of the hemisphere solid angle.
[0090] The formula used for radiance is: L(P) = E(P) + ρ(P) ∫ Ω f r (P, W) L(P′, W) (W·n) dw
[0091] Where, L(P) is the radiance of point P, E(P) is the self-luminance of point P, ρ(P) is the reflectivity of point P, L(P′,W) is the incident radiance from point P′ along direction W, and f r (P,W) is the BRDF function that defines the reflection characteristics, and (W·n) is the cosine term of the incident direction W and the surface normal n;
[0092] The specific steps for using global illumination to simulate lighting and shadow effects in the real world are as follows:
[0093] Step 1: Create a virtual scene in 3D modeling software and add light sources to the scene;
[0094] Step 2: Use ray tracing and radiosity algorithms to simulate how light reflects, refracts, scatters, and shadows; calculate the indirect lighting effect, i.e., the effect after light is reflected multiple times in the scene;
[0095] Step 3: Create light maps for the objects in the scene. Among them, static lighting information is used to simulate complex shadow and lighting effects. In baked lighting, the lighting information is "baked" into the texture of the object, so that the lighting effect can be reproduced during rendering, reducing the real-time calculation amount;
[0096] Step 4: Use Unreal Engine 4 for real-time global illumination. Simulate the human eye's perception of light by adjusting tone mapping. Enhance the contact shadows and occlusion areas between objects by adding ambient occlusion effects. Simulate the glow and halo effects generated by a real camera lens in strong light. Finally, use global illumination technology to render the scene, output the final image or video, and adjust the light sources, materials, and scene settings according to the rendering results.
[0097] The specific steps for integrating the digital twin model into Unity for real-time rendering and interaction are as follows:
[0098] Step 1: Reduce the number of polygons, merge meshes, and optimize textures of the model, and ensure that the digital twin model is in a format supported by Unity. In Unity, select Import New Asset through the Assets menu and select the model file. In the import settings, adjust the scaling, coordinate system, and materials of the model as needed;
[0099] Step 2: Create new materials in Unity and assign them to different parts of the model. Specify the built-in shaders in Unity or custom GLSL / HLSL shaders for the materials, and adjust the parameters of the shaders in the material editor, such as color, texture, and reflectivity;
[0100] Step 3: Add a light source to the scene and configure global illumination in the Lighting window of Unity, including baking lightmaps and real-time light probes, and set shadow casting and receiving properties for the light source and objects.
[0101] Step 4: Place the model in a suitable position in the scene, add a background, ground, and skybox, add colliders to the digital twin for physical interaction, write scripts in C# to control the movement, rotation, and scaling of objects, click the Play button in Unity to run the scene, view the real-time rendering effect of the digital twin, reduce DrawCalls for the scene according to the test results, select Build Settings in the File menu of Unity, set the target platform, and build the project.
[0102] Example 2
[0103] Use an industrial camera of Basler ace acA2000-50gc to collect multi-view images of metal parts, with a fixed resolution of 1920×1080 and uniform LED cold light source for lighting conditions.
[0104] Transmit the images to the server via the TCP / IP protocol, which is configured with an NVIDIA Tesla V100 GPU and 32GB of memory).
[0105] Poisson equation
[0106] where is the Laplace operator, representing the second derivative of the surface φ; the divergence of the vector field V, reflecting the convergence / divergence degree of the point cloud normal direction;
[0107] For the 3D model reconstructed by the Poisson equation, combined with the CutX model segmentation results, the surface defect detection accuracy reaches 98.7%. Test dataset: 5000 metal part images with cracks and scratches.
[0108] The single-frame processing time is 50ms, including the entire process of image denoising, feature extraction, segmentation, and 3D reconstruction, meeting the real-time detection requirements of the production line.
[0109] Compared with traditional manual inspection, the efficiency is increased by 20 times, and the missed detection rate is reduced to less than 0.5%.
[0110] Example 3
[0111] In this example, the method is used for real-time rendering and interaction in a virtual reality scene. The specific implementation steps are as follows:
[0112] Hardware environment configuration: Oculus Quest 2 headset (six-degree-of-freedom tracking), and the PC is configured with an RTX 3080 graphics card.
[0113] Software environment configuration: Unity 2021.3 LTS engine, integrated with custom shaders written in GLSL.
[0114] Performance metrics:
[0115] For the rendering frame rate, the digital twin model with global illumination and dynamic shadows achieves stable rendering at 90 FPS in Unity.
[0116] For the interaction latency, the latency from head movement input to digital twin response is less than 10 ms, supporting gesture recognition and physical collision detection.
[0117] The visual consistency error of the digital twin after style transfer is less than 3% at different perspectives, significantly reducing the user's sense of dizziness.
[0118] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices.
[0119] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0120] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0121] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for capturing an object image in real time based on computer vision technology, characterized in that: The steps include: Step 1: Use a USB camera or a webcam to capture the object image in real time, and transmit the image data to the server for processing in real time via TCP / IP; Step 2: Use Gaussian filter to remove noise, histogram equalization to enhance contrast, and extract image features through a hybrid architecture based on Inception V3 and ResNet-50. The hybrid architecture optimizes feature extraction through a multi-scale attention mechanism, uses a fully convolutional network for pixel-level segmentation, divides the image into different object regions, and then post-processes the segmentation results; Step 3: Use multi-view geometry to recover the 3D information of the object from 2D images with different view angles, generate 3D point cloud data, and use the Poisson equation The 3D surface reconstruction algorithm converts the 3D point cloud data into a smooth 3D surface model, and combines the segmentation results output by the CutX model to process noise and missing data; Step 4: Use style transfer to ensure the visual consistency of the digital twin under different viewing angles, use global illumination to simulate the lighting and shadow effects in the real world, use GLSL language to edit the material of the digital twin, and write a shader program to integrate the digital twin model into Unity for real-time rendering and interaction; Step 5: Transmit the processed digital twin model to the client device through the network interface, and convert it into a controllable signal according to the position, posture, and movement of the digital twin model, thereby driving the external hardware to perform corresponding actions.
2. The method for capturing an object image in real time based on computer vision technology according to claim 1, characterized in that: In step 2, the specific steps of using Inception to extract image features are: We introduce the numpy and tensorflow.keras modules. Numpy is used to handle numerical calculations, and tensorflow.keras is a tool for building and training neural networks. We use the trained Inception V3 model and prepare the image to be analyzed. After loading the image, resize it to 299x299 pixels, convert the image into an array and increase the dimension. We preprocess the image to match the data format used when training the model. The predict function runs the image data through the model and outputs the features. We flatten the features and then use a chart to display these features. Finally, we save the features as a .npy file.
3. The method for capturing an object image in real time based on computer vision technology according to claim 1, characterized in that: The specific steps of removing small noise points through morphological operations and extracting the object boundary through edge detection algorithm are as follows: Step 1: Read the segmentation map output by FCN, which is a two-dimensional array, in which each element represents the category label of a pixel, and create a 3x3 square structure element, in which the value of each pixel represents its category label; Step 2: Apply corrosion operation to the segmentation result to remove small noise points, apply dilation operation to the corrosion result to restore the main shape of the object, and iterate corrosion and dilation operations multiple times until the small noise points are effectively removed; Step 3: When the seg_map is in color, it needs to be converted to a grayscale image for edge detection. The Canny algorithm is used to extract the edges of the object. When finer edges are needed, the Zhang-Suen algorithm is used to superimpose the edge detection results on the original segmentation map to clearly display the boundaries of the object.
4. The method for capturing an object image in real time based on computer vision technology according to claim 1, characterized in that: The specific steps of using multi-view geometry to recover the three-dimensional information of an object from two-dimensional images of different perspectives and generate three-dimensional point cloud data are as follows: The segmentation results output by the CutX model are used to assist in finding the pixel points corresponding to the same object in the multi-view images. The corresponding points in the two-dimensional image are converted into points in the three-dimensional space using the camera parameters obtained by the previous calibration. For each matching point pair, the triangulation principle is used to calculate their positions in the three-dimensional space. All the three-dimensional points obtained by triangulation are brought together to form a three-dimensional point cloud of the object. Finally, the generated point cloud is optimized.
5. The method for capturing an object image in real time based on computer vision technology according to claim 1, characterized in that: The specific steps of converting 3D point cloud data into a smooth 3D surface model using the Poisson surface reconstruction algorithm are as follows: Input the cleaned point cloud data into the Poisson surface reconstruction algorithm, read the point cloud data, estimate a normal for each point in the point cloud to create a normal estimator before Poisson reconstruction, merge the point cloud and normal, perform Poisson surface reconstruction, save or display the reconstructed mesh, and after Poisson surface reconstruction is completed, optimize the mesh quality through mesh smoothing and simplification. When the original image contains texture information, map these textures onto the reconstructed surface.
6. The method for capturing an object image in real time based on computer vision technology according to claim 1, characterized in that: In step 2, the training of the fully convolutional network adopts a composite loss function, which is composed of 60% Dice loss and 40% cross entropy loss weighted, and the specific formula is: L total 0.6 L Dice +0.4 L CE Among them, L Dice The boundary accuracy is optimized by calculating the overlap between the predicted segmentation and the true label, L CE Optimizing category consistency through pixel-level classification error; The weight allocation mechanism of the multi-scale attention module is as follows: the low-level features of Inception V3 are given a weight of 0.3, the middle-level features are given a weight of 0.5, and the high-level features are given a weight of 0.
2.
7. The method for capturing an object image in real time based on computer vision technology according to claim 1, characterized in that: In step 4, the total loss function of style transfer is defined as: L total =α||F res5c (I g )-F res5c (I c )|| 2 +β∑ l ||G(F l (I g ))-G(F l (I s ))|| 2 Among them, I g To generate an image, I c is the content image, I s is the style image; F res5c represents the content features extracted by the res5c layer in ResNet101, F l Style features extracted for res2c / res3c / res4c / res5c layers; G(·) is the Gram matrix, which is used to quantify the statistical distribution of style features; β is the weight coefficient.
Citation Information
Cited By
Method and system for detecting appearance defects of automobile parts
CN120374600A
Livestock head data analysis and management system in breeding process
CN120615770A