Visual processing method, apparatus, device, and storage medium

CN122820697APending Publication Date: 2026-09-25SHENZHEN ANGSTROM EXCELLENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611253025.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,上述3D重建过程设涉及空间平移、仿射变换、各个视角下各个像素点位置对齐等步骤,数据计算量较大;而且,3D重建过程是直接融合不同视角下在探测器的二维平面上的像素特征,对于同一像素点,容易混合半导体器件不同区域的信息,像素特征精度低,导致后续视觉处理结果的准确度不高

Benefits of technology

[0009]本申请实施例提供了一种视觉处理方法、装置、设备及存储介质,根据本申请提供的方案,获取多个视角下的像素数据;像素数据是由探测器对射线透过待测物体后的投影进行采集得到的二维图像;将多个视角下的像素数据以及各个视角对应的探测器的相机标定参数,输入深度学习模型,通过深度学习模型对待测物体的几何位置特征进行学习,输出目标视角下的修正几何特征;深度学习模型是训练完成的端到端模型,专注于每个视角下的像素数据与待测物体的实际几何位置之间的关系,不需要对每个视角中的各个像素点均进行位置对齐,放弃3D重建过程,降低数据计算量,提高了数据处理效率。目标视角为多个视角中的一个或多个视角;相机标定参数指示射线方向与视角之间的物理空间联系,深度学习模型具有围绕物理空间联系聚合射线方向上的跨视角特征的功能;沿射线方向对多个视角下位于同一多维空间区域的像素特征进行聚合,保证了像素特征传输的几何一致性,提高了修正几何特征的精度。根据目标视角下的修正几何特征进行视觉处理,得到待测物体的视觉处理结果,提高了视觉处理的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820697A_ABST
    Figure CN122820697A_ABST
Patent Text Reader

Abstract

The application discloses a kind of visual processing method, device, equipment and storage medium, belong to computer technical field.The method comprises: obtaining pixel data under multiple perspectives;Pixel data is the two-dimensional image obtained by the detector to the projection after the projection of the ray through the object to be measured;The camera calibration parameters of the detector corresponding to each perspective are input into the deep learning model, and the geometric position characteristics of the object to be measured are learned by the deep learning model, and the corrected geometric characteristics under the target perspective are output;Target perspective is one or more perspectives in multiple perspectives;Along the ray direction, the pixel characteristics in the same multidimensional space region under multiple perspectives are aggregated, the geometric consistency of pixel feature transmission is ensured, and the accuracy of the corrected geometric characteristics is improved.According to the corrected geometric characteristics under the target perspective, visual processing is carried out, and the visual processing result of the object to be measured is obtained, and the accuracy of visual processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a visual processing method, apparatus, device, and storage medium. Background Technology

[0002] In semiconductor non-destructive testing scenarios, two-dimensional (2D) X-ray projection images of semiconductor devices are acquired from multiple perspectives using detectors. Three-dimensional (3D) reconstruction is then performed on the 2D images from multiple perspectives, such as CT reconstruction or feature alignment in the UV plane. In the UV plane, the U-axis represents the horizontal direction and the V-axis represents the vertical direction, in order to reconstruct a 3D stereo model. Then, according to the reverse process of 3D reconstruction, the 3D stereo model is mapped onto the 2D image, and the 2D image is then subjected to visual processing to achieve non-destructive testing of semiconductor devices.

[0003] However, the aforementioned 3D reconstruction process involves steps such as spatial translation, affine transformation, and alignment of pixel positions under various viewpoints, resulting in a large amount of data computation. Moreover, the 3D reconstruction process directly fuses pixel features on the detector's two-dimensional plane from different viewpoints. For the same pixel, it is easy to mix information from different regions of the semiconductor device, leading to low pixel feature accuracy and consequently, low accuracy of subsequent visual processing results. Summary of the Invention

[0004] This application provides a visual processing method, apparatus, device, and storage medium, which can improve the accuracy of visual processing and reduce the amount of data computation. The technical solution is as follows: In a first aspect, a visual processing method is provided, the method comprising: acquiring pixel data from multiple viewpoints; the pixel data being a two-dimensional image obtained by a detector capturing the projection of rays through an object under test; inputting the pixel data from the multiple viewpoints and the camera calibration parameters of the detector corresponding to each viewpoint into a deep learning model; learning the geometric position features of the object under test through the deep learning model; and outputting corrected geometric features from a target viewpoint; the target viewpoint being at least one of the multiple viewpoints; the camera calibration parameters indicating the physical spatial relationship between the ray direction and the viewpoint; the deep learning model having the function of aggregating cross-viewpoint features along the ray direction around the physical spatial relationship; and performing visual processing based on the corrected geometric features from the target viewpoint to obtain the visual processing result of the object under test.

[0005] Secondly, a visual processing device is provided, comprising: an acquisition module for acquiring pixel data from multiple viewpoints; the pixel data being a two-dimensional image obtained by a detector capturing the projection of rays through an object under test; a deep learning module for inputting the pixel data from the multiple viewpoints and the camera calibration parameters of the detector corresponding to each viewpoint into a deep learning model, learning the geometric position features of the object under test through the deep learning model, and outputting corrected geometric features from a target viewpoint; the target viewpoint being at least one of the multiple viewpoints; the camera calibration parameters indicating the physical spatial relationship between the ray direction and the viewpoint, and the deep learning model having the function of aggregating cross-viewpoint features along the ray direction around the physical spatial relationship; and a visual processing module for performing visual processing based on the corrected geometric features from the target viewpoint to obtain the visual processing result of the object under test.

[0006] Thirdly, a computer device is provided, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program implementing the method described in the first aspect when executed by the processor.

[0007] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0008] Fifthly, a computer program product containing instructions is provided that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0009] This application provides a visual processing method, apparatus, device, and storage medium. According to the solution provided, pixel data from multiple viewpoints is acquired. The pixel data is a two-dimensional image obtained by a detector capturing the projection of rays through an object under test. The pixel data from multiple viewpoints, along with the camera calibration parameters of the detectors corresponding to each viewpoint, are input into a deep learning model. The deep learning model learns the geometric position features of the object under test and outputs corrected geometric features from the target viewpoint. The deep learning model is a trained end-to-end model that focuses on the relationship between pixel data from each viewpoint and the actual geometric position of the object under test. It does not require position alignment of each pixel in each viewpoint, abandoning the 3D reconstruction process, reducing data computation, and improving data processing efficiency. The target viewpoint is one or more of the multiple viewpoints. The camera calibration parameters indicate the physical spatial relationship between the ray direction and the viewpoint. The deep learning model has the function of aggregating cross-viewpoint features along the ray direction based on the physical spatial relationship. Aggregating pixel features located in the same multi-dimensional spatial region from multiple viewpoints along the ray direction ensures the geometric consistency of pixel feature transmission and improves the accuracy of corrected geometric features. Visual processing is performed based on the corrected geometric features from the target perspective to obtain the visual processing results of the object under test, thereby improving the accuracy of visual processing. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of a multi-view fusion process provided in an embodiment of this application; Figure 2 This is a flowchart of a visual processing method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a pixel data acquisition process provided in an embodiment of this application; Figure 4 This is a schematic diagram of a pixel data conversion process provided in an embodiment of this application; Figure 5 This is a schematic diagram of a pixel data mapping process provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a vision processing device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0013] It should be understood that "multiple" as mentioned in this application refers to two or more. In the description of this application, unless otherwise stated, " / " indicates "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, to facilitate a clear description of the technical solutions of this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and that "first," "second," etc., do not necessarily imply differences.

[0014] Before providing a detailed explanation of the embodiments of this application, the application scenarios and related technologies of the embodiments of this application will be described first.

[0015] The visual processing method provided in this application embodiment can be applied to X-ray inspection scenarios such as semiconductor packaging, through-silicon vias, solder joints, voids, and cracks, for example, noise reduction, image smoothing, and defect detection.

[0016] Multi-view X-ray projection differs from ordinary video frames. The position of the same object structure changes in different viewpoints. Within the same viewpoint, pixel data differs depending on the ray direction, and vice versa. In other words, the same detector coordinates in different viewpoints typically do not correspond to the same object position. If adjacent viewpoints are directly fused based on two-dimensional pixel positions—for example, in CT reconstruction or feature alignment in the UV plane—it may result in the mixing of object regions that do not actually correspond, leading to low pixel feature accuracy.

[0017] like Figure 1 As shown, Figure 1 This is a schematic diagram of a multi-view fusion process provided in an embodiment of this application. Figure 1 Figure A illustrates a multi-view fusion scheme for CT reconstruction or feature alignment in the UV plane, where fusion is performed directly based on the same detector coordinates. However, the same pixel coordinates do not necessarily represent the same object region or the same defect. For example, the same object region may be in the upper left corner of one image but in the upper right corner of another. This multi-view fusion method can lead to cross-view information mismatch, causing weak defect structures to be confused. The geometric correspondence cannot be explained by simply using "images from adjacent views". Figure 1Figure B in the figure illustrates the physical spatial relationship between the ray direction and the viewing angle provided in the embodiments of this application, and the scheme of aggregating cross-view features in the ray direction. With the ray direction as the organization center, pixel data from multiple viewing angles are projected onto a multi-dimensional reference object space. The calibration geometry of the imaging system (i.e., camera calibration parameters) is used to associate the cross-view object space, and then mapped to the target viewing angle to obtain the corrected geometric features under the target viewing angle. By aggregating cross-view features around the same ray direction, the fusion of cross-view features is based on imaging geometry, rather than the consistency of two-dimensional coordinates.

[0018] Based on the above Figure 1 This application provides a visual processing method, such as... Figure 2 As shown, Figure 2 This is a flowchart of a visual processing method provided in an embodiment of this application. The visual processing method includes: S101. Acquire pixel data from multiple perspectives; pixel data is a two-dimensional image obtained by the detector from the projection of rays through the object under test.

[0019] The object under test can be a finished semiconductor product, a semi-finished semiconductor product, or other devices related to a single semiconductor, including but not limited to integrated circuits, optoelectronic devices (such as light-emitting diodes (LEDs), laser diodes, image sensors, optocouplers, etc.), discrete devices (diodes, transistors, etc.), packaged chips, packaged wafers, packaged batteries, etc.

[0020] For example, such as Figure 3 As shown, Figure 3 This is a schematic diagram illustrating a pixel data acquisition process provided in an embodiment of this application. Figure 3 The diagram illustrates how rays from a certain perspective are projected onto the detector's projection domain through the object under test. The ray source and detector are positioned on opposite sides of the 3D object under test (voxel center xyz), with the ray emitted from the source pointing towards the object. The detector receives the rays penetrating the object and generates pixel data (detector feature pixel) of the object based on the rays penetrating the object and the rays from the source. The pixel data is a two-dimensional image.

[0021] For example, the above Figure 3In this process, a homogeneous projection transformation (PTM, u=p0 / p2, v=p1 / p2) is performed using the geometric relationship between the X-ray source and the detector. The detector performs valid depth sampling, such as bilinear grid sampling, and performs valid-view weighted mean on each viewpoint to obtain the pixel data (canonical lifted feature volume) of each viewpoint, that is, the pixel data of the object under test (detector feature pixel).

[0022] In related technologies, after the detector feature pixel is acquired, 3D reconstruction (3D grid sample along depth) is performed on the UV plane (ray-box intersecton) of the detector. Finally, the pixel data at different stages (stage features) is equivalent to the reverse process of the above process, realizing the conversion between detector pixel and object-space ray.

[0023] S102. Input the pixel data from multiple viewpoints and the camera calibration parameters of the detectors corresponding to each viewpoint into the deep learning model. The deep learning model learns the geometric position features of the object under test and outputs the corrected geometric features under the target viewpoint.

[0024] The target viewpoint is at least one of multiple viewpoints. The camera calibration parameters indicate the physical spatial relationship between the ray direction and the viewpoint. The deep learning model has the function of aggregating cross-viewpoint features in the ray direction around the physical spatial relationship.

[0025] Camera calibration parameters (also known as camera calibration parameters) reflect the correspondence between three-dimensional space and two-dimensional image coordinates. They can also be understood as characterizing the geometric relationships between rays, detectors, target points, and objects. Camera calibration parameters are fixed and do not change with the viewpoint. They include, but are not limited to, camera intrinsic parameters, camera extrinsic parameters, and distortion coefficients.

[0026] Because the position of the same object region changes in different projection viewpoints, multi-view information cannot be fused solely based on two-dimensional coordinates. Therefore, the multi-view fusion scheme provided in this application differs from related technologies that "view images from adjacent viewpoints together," instead adopting a "geometrically based view around the ray direction." The deep learning model does not fuse multi-view features based on the surface similarity of two-dimensional coordinates, but rather utilizes the calibration geometry of the imaging system (i.e., camera calibration parameters) to aggregate cross-view features around the ray direction in the projection viewpoint.

[0027] The calibration geometry of the imaging system (i.e., camera calibration parameters) establishes the physical spatial relationship between ray directions and multiple viewpoints; that is, it describes the relationship between the projection and ray directions of the object space and the viewpoints of each detector. For example... Figure 4 As shown, Figure 4 This is a schematic diagram illustrating a pixel data conversion process provided in an embodiment of this application. The calibration geometry of the imaging system (i.e., camera calibration parameters) and pixel data from multiple viewpoints are input into a deep learning model. The deep learning model can process pixel data from multiple viewpoints simultaneously and share processing parameters across these viewpoints, but it does not average the features from different viewpoints. The deep learning model aggregates cross-viewpoint features along ray directions based on physical spatial relationships. For example, it aggregates cross-viewpoint features along ray directions A, B, and C, and outputs a set of corrected geometric features along the ray directions while maintaining the order and correspondence. Figure 4 The corrected geometric features for ray direction A, ray direction B, and ray direction C are shown in the figure.

[0028] S103. Perform visual processing based on the corrected geometric features from the target viewpoint to obtain the visual processing result of the object under test.

[0029] As mentioned above Figure 4 As shown, the corrected geometric features along the ray direction output by the deep learning model (i.e., the corrected geometric features from the target's perspective) can be used for subsequent visual processing. Visual processing includes, but is not limited to, image denoising, defect detection, size measurement, image enhancement, target recognition, target classification, 3D reconstructed stereo defect detection, and volume measurement.

[0030] Correspondingly, visual processing results include, but are not limited to, noise reduction results, defect detection results, object size and volume, image enhancement results, and object category.

[0031] According to the scheme provided in this application, pixel data from multiple perspectives is acquired. The pixel data is a two-dimensional image obtained by a detector capturing the projection of rays through the object under test. The pixel data from multiple perspectives, along with the camera calibration parameters of the detectors corresponding to each perspective, are input into a deep learning model. The deep learning model learns the geometric position features of the object under test and outputs corrected geometric features from the target perspective. The deep learning model is a trained end-to-end model that focuses on the relationship between pixel data from each perspective and the actual geometric position of the object under test. It does not require position alignment of each pixel in each perspective, abandoning the 3D reconstruction process, reducing data computation, and improving data processing efficiency. The target perspective is one or more perspectives from multiple perspectives. The camera calibration parameters indicate the physical spatial relationship between the ray direction and the perspective. The deep learning model has the function of aggregating cross-perspective features along the ray direction based on the physical spatial relationship. Aggregating pixel features located in the same multi-dimensional spatial region from multiple perspectives along the ray direction ensures the geometric consistency of pixel feature transmission and improves the accuracy of the corrected geometric features. Visual processing is performed based on the corrected geometric features from the target perspective to obtain the visual processing result of the object under test, improving the accuracy of visual processing.

[0032] In some embodiments, the above Figure 2 S102 can also be implemented in the following way. The deep learning model transforms the pixel data from multiple viewpoints into a multi-dimensional reference object space according to the camera calibration parameters of the detectors corresponding to each viewpoint, and obtains latent spatial features; the latent spatial features include the corrected features after aligning and fusing the pixel data from multiple viewpoints according to the geometric position of the object to be tested; the latent spatial features are mapped according to the target viewpoint to obtain the corrected geometric features under the target viewpoint.

[0033] For example, such as Figure 5 As shown, Figure 5 This is a schematic diagram illustrating a pixel data mapping process provided in an embodiment of this application. The correlation between multi-view X-ray projections (i.e., the correlation between pixel data and corrected geometric features) is determined by the three-dimensional geometric position of the object. Based on this, multi-view two-dimensional features (i.e., pixel data from multiple viewpoints) are input, i.e. Figure 3The pixel data of the object under test (detector feature pixel) is sampled. A deep learning model learns the geometric position features of the object, transforming the multi-view 2D features into a multi-dimensional reference object space (canonical object-space grid) to obtain latent space features. The multi-dimensional reference object space includes multiple latent space sub-blocks, which can be used to uniformly transform the multi-view 2D features into individual latent space sub-blocks. These are then mapped to the target viewpoint to obtain the corrected geometric features under the target viewpoint, completing the calibration geometric transfer path of "multi-view 2D features (object xyz) → multi-dimensional reference object space → corrected geometric features under the target viewpoint (detector UV 2D pixel grid)".

[0034] This application provides a multi-view X-ray projection method based on latent spatial feature transfer. Utilizing a multi-dimensional reference object space defined by scanning calibration geometry, cross-view feature transfer is established between pixel data from multiple viewpoints acquired by a two-dimensional detector in multi-view X-ray projection, the multi-dimensional reference object space, and each viewpoint. Based on the obtained pixel data from the target viewpoint, conditional deep learning prediction is performed on its projection to generate a projection corresponding to the target viewpoint (i.e., corrected geometric features). Cross-view feature alignment is transferred from the two-dimensional detector plane to the multi-dimensional reference object space defined by scanning calibration geometry, enabling pixel data from different viewpoints to be aligned and fused according to the geometric position coordinates of the same object under test, improving the accuracy of latent spatial features. Then, the latent spatial features of the object under test are remapped according to the target viewpoint geometry to form corrected geometric features under the target viewpoint, completing the detector's projection domain prediction process. By aggregating pixel features located in the same multi-dimensional spatial region from multiple viewpoints along the ray direction, geometric consistency of pixel feature transfer is ensured, improving the accuracy of corrected geometric features.

[0035] It should be noted that the embodiments of this application do not limit the network structure and output parameterization method of the deep learning model. As long as it can realize the unified transformation of multi-view two-dimensional features using multi-dimensional reference object space and remap them to the target viewpoint to form the corrected geometric features under the target viewpoint, the projection domain prediction process of the conditional detector can be realized.

[0036] In some embodiments, the visual processing method further includes a deep learning model training process. This involves acquiring pixel data samples of an object sample from multiple viewpoints, along with geometric position annotations corresponding to the pixel data samples from any viewpoint; inputting the pixel data samples from multiple viewpoints, the geometric position annotations, and the camera calibration parameters of the detectors corresponding to each viewpoint into an initial deep learning model; using the initial deep learning model to learn the geometric position features of the object sample; and outputting predicted data from multiple viewpoints; and updating and training the initial deep learning model based on the pixel data samples from multiple viewpoints and the predicted data from multiple viewpoints to obtain the deep learning model.

[0037] Related technologies rely on 3D reconstruction, which requires sample-by-sample orthographic projection. That is, each pixel data sample needs to be labeled with geometric information, or additional external teacher data needs to be added to supplement geometric information. This increases the amount of computation and data preparation costs, and may introduce teacher bias, resulting in low feature accuracy.

[0038] In this embodiment, training constraints are constructed from the input projection (corresponding to pixel data samples of an object sample under multiple viewpoints) and the calibration geometry itself (corresponding to the camera calibration parameters of the detector for each viewpoint). This maintains geometric consistency in pixel feature transmission without relying on per-sample orthographic projection or additional external teacher data (corresponding to the geometric position annotations of pixel data samples under any viewpoint). In other words, using geometric constraints and output constraints to train the deep learning model improves training accuracy.

[0039] The target latent feature reference (corresponding to predicted data from multiple views) is generated from the input projection (corresponding to pixel data samples of the object sample from multiple viewpoints) and the calibration geometry itself (corresponding to the camera calibration parameters of the detector for each viewpoint). Training is then updated based on the geometric position annotations and predicted data. By constraining the geometric consistency of the multi-dimensional baseline object space, an intermediate feature space is organized and passed on multi-view features under geometric and output constraints. The deep learning model can form a set of reference features based on the same input information and the same calibration geometry, used to constrain whether the learned features maintain reasonable geometric relationships. Layered convergence is achieved by employing boundary updates and gradient stopping, continuing training after fixing some parameters until the model converges, thus improving training efficiency.

[0040] In some embodiments, the above Figure 2 S103 can also be implemented in the following way: Extract features from the corrected geometric features under the target viewpoint to obtain pixel-coded features; perform visual processing on the pixel-coded features to obtain the visual processing result of the object under test.

[0041] The corrected geometric features from the target's perspective are information representations in the spatial dimension and cannot be directly used for subsequent visual processing. Therefore, as mentioned above... Figure 5 As shown, after mapping the corrected geometric features (detector UV 2Dpixel grid) under the target viewpoint, feature extraction can be performed on it. The extracted pixel-coded feature grid is a regular machine language, which facilitates subsequent visual processing and improves visual processing efficiency.

[0042] In some embodiments, visual processing includes noise reduction and / or defect detection, and the visual processing result includes a noise reduction result or a defect detection result. The steps described above for visually processing pixel-coded features to obtain the visual processing result of the object under test can also be implemented in the following ways: Perform noise reduction processing on the pixel-coded features to obtain a noise reduction result of the object under test; or, perform defect detection on the pixel-coded features or the noise reduction result to obtain a defect detection result of the object under test.

[0043] During the production process, the object under test may develop various internal defects, such as cracks, dark marks, and excessive gaps. These internal defects can affect the quality and safety of the object under test.

[0044] The relevant technology employs a 3D reconstruction process, whose original projection may contain noise. Small defects often have low contrast, and directly denoising the features after 3D reconstruction can easily smooth out weak defects along with the original. In scenarios with strong noise and sparse defects, directly denoising the features after 3D reconstruction makes it difficult to distinguish between random noise and real weak structures, and can easily damage key defect signals such as voids and cracks in through silicon vias (TSVs).

[0045] In this embodiment, a deep learning model is used to aggregate cross-view features along the physical space connection ray direction. The accuracy of the corrected geometric features under the target viewpoint is high. Noise reduction is performed based on the pixel-coded features corresponding to these corrected geometric features. The defect detection result can be shown in the form of a mask or annotation, which can improve the accuracy of the noise reduction result. Defect detection based on the pixel-coded features corresponding to these corrected geometric features can improve the accuracy of the defect detection result.

[0046] Furthermore, noise reduction can be performed on the pixel-coded features before defect detection, thereby improving the accuracy of the defect detection results.

[0047] In some embodiments, the steps of visually processing the pixel-coded features to obtain the visual processing result of the object under test can also be implemented in the following ways: denoising the pixel-coded features using a denoising network to obtain a denoising result; and performing defect detection on the pixel-coded features or the denoising result using a defect detection network to obtain a defect detection result.

[0048] In this embodiment, a network is used to perform noise reduction or defect detection. Compared with noise reduction or defect detection algorithms, the network can be used as the task head (including noise reduction head or defect detection head) of a deep learning model, which can achieve lightweight deployment, reduce deployment complexity, and as an end-to-end model, the path is shorter and the data processing efficiency is high.

[0049] It should be noted that the embodiments of this application do not limit the network structure and output parameterization method of the denoising network and the defect detection network, as long as it can realize the technical relationship between the target view denoising projection and defect detection based on the multi-view two-dimensional features.

[0050] In this embodiment, the projection image quality optimization (i.e., noise reduction), defect detection, and latent spatial features are processed in a hierarchical manner. This ensures that the noise reduction data (i.e., pixel-coded features), the defect detection data (i.e., pixel-coded features or noise reduction results), and the cross-view geometric constraints each have clear data sources and boundaries of action. This allows for independent training of the deep learning model, the noise reduction network, and the defect detection network. The noise reduction network can serve as the noise reduction head of the deep learning model, and the defect detection network can serve as the defect detection head of the deep learning model. Both are trained independently, reducing the number of training iterations, enabling the noise reduction network to converge quickly, reducing data processing volume, and improving training efficiency.

[0051] This application provides an overall architecture in which the projected image quality layer (i.e., denoising network), the defect detection layer (i.e., defect detection network), and the geometric consistency layer of latent spatial features (i.e., deep learning model) are separated and trained collaboratively. This overall architecture has the function of outputting denoising processing results or defect detection results based on pixel data from multiple viewpoints. Furthermore, because this overall architecture abandons the 3D reconstruction process and improves the accuracy of correcting geometric features under the target viewpoint by aggregating cross-viewpoint features along the ray direction around the physical space, it improves the accuracy of denoising processing results and defect detection results.

[0052] In some embodiments, the visual processing method further includes a training process for a denoising network. This involves acquiring a first pixel data sample and a noisy data sample after adding noise to the first pixel data sample; using the noisy data sample as input to an initial denoising network, and outputting predicted denoising data through the initial denoising network; obtaining a first loss value by applying a first preset loss function based on the first pixel data sample and the predicted denoising data; and updating and training the initial denoising network based on the first loss value to obtain the denoising network.

[0053] For example, multiple noisy data samples are input into an initial denoising network, which performs denoising processing and outputs predicted denoised data corresponding to each noisy data sample. A first loss value is calculated based on the multiple predicted denoised data, multiple first pixel data samples, and a first preset loss function. The training process is backpropagated based on the output first loss value, and the initial denoising network is iteratively trained using the first loss value. Each iteration updates the parameters in the initial denoising network until a training termination condition is met, such as reaching a preset number of training iterations or the first loss value meeting a predetermined threshold, thus obtaining the denoising network. This embodiment does not limit the preset number of iterations or the predetermined threshold.

[0054] The first preset loss function can be a cross-entropy loss function, a mean-squared error loss function, an L1 norm loss function, a smooth L1 loss, etc., and this application embodiment does not limit it.

[0055] In this embodiment, the training process of the denoising network is independent of the training processes of the deep learning model and the defect detection network. It can be used as the denoising head of the deep learning model and trained independently, reducing the number of training iterations, enabling the denoising network to converge quickly, reducing the amount of data processing, and improving training efficiency.

[0056] In some embodiments, the visual processing method further includes a training process for a defect detection network. This involves acquiring second pixel data samples and corresponding defect location annotations; using the second pixel data samples as input to an initial defect detection network, and outputting predicted defect locations through the initial defect detection network; obtaining a second loss value by applying a second preset loss function based on the predicted defect locations and defect location annotations; and updating and training the initial defect detection network based on the second loss value to obtain the defect detection network.

[0057] For example, multiple second-pixel data samples are input into an initial defect detection network. The initial defect detection network performs defect detection and outputs the predicted defect location corresponding to each second-pixel data sample. A second loss value is calculated based on multiple predicted defect locations, multiple defect location labels, and a second preset loss function. The training process is backpropagated based on the output second loss value. The initial defect detection network is iteratively trained using the second loss value. Each iteration updates the parameters in the initial defect detection network until a training termination condition is met, such as reaching a preset number of training iterations or the second loss value meeting a predetermined threshold. The defect detection network is then obtained. This embodiment does not limit the preset number of iterations or the predetermined threshold.

[0058] The second preset loss function can be a cross-entropy loss function, a mean-squared error loss function, an L1 norm loss function, a smooth L1 loss, etc., and this application embodiment does not limit it.

[0059] In this embodiment, the training process of the defect detection network is independent of the training processes of the deep learning model and the denoising network. It can be used as the defect detection head of the deep learning model and trained independently, reducing the number of training iterations, enabling the denoising network to converge quickly, reducing the amount of data processing, and improving training efficiency.

[0060] Based on the visual processing method provided in the above embodiments Figure 6 This is a schematic diagram of the structure of a visual processing device provided in an embodiment of this application. This device can be implemented as part or all of a computer device by software, hardware, or a combination of both. See also... Figure 6 The vision processing device 60 includes: an acquisition module 601 for acquiring pixel data from multiple viewpoints; the pixel data is a two-dimensional image obtained by a detector collecting the projection of rays through the object under test; a deep learning module 602 for inputting the pixel data from multiple viewpoints and the camera calibration parameters of the detector corresponding to each viewpoint into a deep learning model, learning the geometric position features of the object under test through the deep learning model, and outputting corrected geometric features under the target viewpoint; the target viewpoint is at least one of the multiple viewpoints; the camera calibration parameters indicate the physical spatial relationship between the ray direction and the viewpoint, and the deep learning model has the function of aggregating cross-viewpoint features in the ray direction around the physical spatial relationship; and a vision processing module 603 for performing vision processing based on the corrected geometric features under the target viewpoint to obtain the vision processing result of the object under test.

[0061] Optionally, the deep learning module 602 is further configured to use a deep learning model to transform pixel data from multiple viewpoints into a multi-dimensional reference object space based on the camera calibration parameters of the detector corresponding to each viewpoint, thereby obtaining latent spatial features; the latent spatial features include corrected features after aligning and fusing pixel data from multiple viewpoints according to the geometric position of the object to be tested; and to map the latent spatial features according to the target viewpoint to obtain corrected geometric features under the target viewpoint.

[0062] Optionally, the visual processing device 60 further includes a training module 604, which is used to acquire pixel data samples of an object sample from multiple viewpoints and geometric position annotations corresponding to pixel data samples from any viewpoint among the multiple viewpoints; input the pixel data samples from multiple viewpoints, geometric position annotations, and camera calibration parameters of the detectors corresponding to each viewpoint into an initial deep learning model; use the initial deep learning model to learn the geometric position features of the object sample and output prediction data from multiple viewpoints; update and train the initial deep learning model based on the pixel data samples from multiple viewpoints and the prediction data from multiple viewpoints to obtain a deep learning model.

[0063] Optionally, the vision processing module 603 is also used to extract features from the modified geometric features under the target viewpoint to obtain pixel-coded features; and to perform visual processing on the pixel-coded features to obtain the visual processing result of the object under test.

[0064] Optionally, the visual processing includes noise reduction processing and / or defect detection, and the visual processing result includes noise reduction processing result or defect detection result; the visual processing module 603 is also used to perform noise reduction processing on the pixel-coded features to obtain the noise reduction processing result of the object under test; or, to perform defect detection on the pixel-coded features or noise reduction processing result to obtain the defect detection result of the object under test.

[0065] Optionally, the visual processing module 603 is further configured to perform noise reduction processing on the pixel coding features according to the noise reduction network to obtain the noise reduction result; and to perform defect detection on the pixel coding features or the noise reduction result according to the defect detection network to obtain the defect detection result; wherein the training processes of the deep learning model, the noise reduction network and the defect detection network are independent of each other.

[0066] Optionally, the training module 604 is further configured to acquire a first pixel data sample and a noise data sample after adding noise to the first pixel data sample; use the noise data sample as the input of the initial denoising network, and output predicted denoising data through the initial denoising network; according to the first pixel data sample and the predicted denoising data, adopt a first preset loss function to obtain a first loss value; and update and train the initial denoising network according to the first loss value to obtain a denoising network.

[0067] Optionally, the training module 604 is further configured to acquire the second pixel data sample and the defect location label corresponding to the second pixel data sample; use the second pixel data sample as the input of the initial defect detection network, and output the predicted defect location through the initial defect detection network; obtain the second loss value by using the second preset loss function based on the predicted defect location and the defect location label; and update and train the initial defect detection network based on the second loss value to obtain the defect detection network.

[0068] It should be noted that the vision processing device provided in the above embodiments is only illustrated by the division of the above functional modules when performing vision processing. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0069] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.

[0070] The visual processing device and visual processing method embodiments provided in the above embodiments belong to the same concept. The specific working process and technical effects of the units and modules in the above embodiments can be found in the method embodiment section, and will not be repeated here.

[0071] Based on the visual processing method provided in the above embodiments Figure 7 This application provides a schematic diagram of the structure of a computer device, as shown in the embodiment of the present application. Figure 7 As shown, the computer device 70 includes a processor 701, a memory 702, and a computer program 703 stored in the memory 702 and executable on the processor 701. When the processor 701 executes the computer program 703, it implements the steps in the visual processing method in the above embodiments.

[0072] The computer device 70 can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device 70 can be a desktop computer, a portable computer, a network server, a handheld computer, a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. This application embodiment does not limit the type of computer device 70. Those skilled in the art will understand that... Figure 7The computer device 70 is merely an example and does not constitute a limitation on the computer device 70. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0073] Processor 701 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0074] In some embodiments, memory 702 may be an internal storage unit of computer device 70, such as a hard disk or memory of computer device 70. In other embodiments, memory 702 may be an external storage device of computer device 70, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on computer device 70. Furthermore, memory 702 may include both internal and external storage units of computer device 70. Memory 702 is used to store operating system, application programs, boot loader, data, and other programs. Memory 702 may also be used to temporarily store data that has been output or will be output.

[0075] This application also provides a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.

[0076] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the various method embodiments described above.

[0077] This application provides a computer program product that, when run on a computer, causes the computer to perform the steps described in the various method embodiments above.

[0078] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above method embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices. The computer-readable storage medium mentioned in this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.

[0079] It should be understood that all or part of the steps of the above embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-described computer-readable storage medium.

[0080] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0081] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0082] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A visual processing method, characterized in that, The method includes: Pixel data is acquired from multiple perspectives; the pixel data is a two-dimensional image obtained by the detector from the projection of rays passing through the object under test. The pixel data from the multiple viewpoints and the camera calibration parameters of the detectors corresponding to each viewpoint are input into the deep learning model. The deep learning model learns the geometric position features of the object under test and outputs the corrected geometric features under the target viewpoint. The target viewpoint is at least one of the multiple viewpoints. The camera calibration parameters indicate the physical spatial relationship between the ray direction and the viewpoint. The deep learning model has the function of aggregating cross-viewpoint features in the ray direction around the physical spatial relationship. Visual processing is performed based on the corrected geometric features under the target viewpoint to obtain the visual processing result of the object under test; The process of inputting pixel data from multiple viewpoints and camera calibration parameters of the detectors corresponding to each viewpoint into a deep learning model, learning the geometric position features of the object under test through the deep learning model, and outputting corrected geometric features from the target viewpoint includes: The deep learning model transforms the pixel data from the multiple viewpoints into a multi-dimensional reference object space based on the camera calibration parameters of the detectors corresponding to each viewpoint, thereby obtaining latent spatial features. The latent spatial features include corrected features after aligning and fusing the pixel data from the multiple viewpoints according to the geometric position of the object under test. The latent spatial features are mapped according to the target perspective to obtain the corrected geometric features under the target perspective.

2. The method as described in claim 1, characterized in that, The method further includes: Obtain pixel data samples of an object sample from multiple viewpoints and the geometric position annotations corresponding to pixel data samples from any of the multiple viewpoints; The pixel data samples from the multiple viewpoints, the geometric position annotations, and the camera calibration parameters of the detectors corresponding to each viewpoint are input into the initial deep learning model. The initial deep learning model is used to learn the geometric position features of the object samples and output the prediction data from the multiple viewpoints. The initial deep learning model is updated and trained based on pixel data samples and prediction data from the multiple perspectives to obtain the deep learning model.

3. The method as described in claim 1 or 2, characterized in that, The step of performing visual processing based on the corrected geometric features under the target viewpoint to obtain the visual processing result of the object under test includes: Feature extraction is performed on the modified geometric features under the target viewpoint to obtain pixel-coded features; The pixel-coded features are subjected to visual processing to obtain the visual processing result of the object under test.

4. The method as described in claim 3, characterized in that, The visual processing includes noise reduction processing and / or defect detection, and the visual processing result includes the noise reduction processing result or the defect detection result. The step of visually processing the pixel-encoded features to obtain the visual processing result of the object under test includes: The pixel-coded features are subjected to noise reduction processing to obtain the noise reduction result of the object under test; Alternatively, defect detection can be performed on the pixel encoding features or the noise reduction processing results to obtain the defect detection results of the object under test.

5. The method as described in claim 4, characterized in that, The step of performing noise reduction processing on the pixel-encoded features to obtain the noise reduction result of the object under test includes: The pixel-encoded features are denoised using a denoising network to obtain the denoising result; the training processes of the deep learning model and the denoising network are independent of each other. The method further includes: Acquire a first pixel data sample and a noise data sample after adding noise to the first pixel data sample; The noise data samples are used as input to the initial noise reduction network, and the initial noise reduction network outputs predicted noise reduction data. Based on the first pixel data sample and the predicted noise reduction data, a first preset loss function is used to obtain a first loss value; The initial denoising network is updated and trained based on the first loss value to obtain the denoising network.

6. The method as described in claim 4, characterized in that, The step of performing defect detection on the pixel encoding features or the noise reduction processing result to obtain the defect detection result of the object under test includes: Defect detection is performed on the pixel-encoded features or the noise reduction result by the defect detection network to obtain the defect detection result; the training processes of the deep learning model and the defect detection network are independent of each other; The method further includes: Obtain the second pixel data sample and the corresponding defect location annotation; The second pixel data sample is used as the input to the initial defect detection network, and the predicted defect location is output by the initial defect detection network. Based on the predicted defect location and the defect location label, a second preset loss function is used to obtain a second loss value; The initial defect detection network is updated and trained based on the second loss value to obtain the defect detection network.

7. A vision processing device, characterized in that, The device includes: The acquisition module is used to acquire pixel data from multiple perspectives; the pixel data is a two-dimensional image obtained by the detector collecting the projection of rays through the object under test. The deep learning module is used to input pixel data from the multiple viewpoints and camera calibration parameters of the detectors corresponding to each viewpoint into the deep learning model. The deep learning model learns the geometric position features of the object under test and outputs corrected geometric features from the target viewpoint. The target viewpoint is at least one of the multiple viewpoints. The camera calibration parameters indicate the physical spatial relationship between the ray direction and the viewpoint. The deep learning model has the function of aggregating cross-viewpoint features in the ray direction around the physical spatial relationship. The visual processing module is used to perform visual processing based on the corrected geometric features under the target viewpoint to obtain the visual processing result of the object under test. The deep learning module is further configured to use the deep learning model to convert pixel data from multiple viewpoints into a multi-dimensional reference object space based on the camera calibration parameters of the detectors corresponding to each viewpoint, thereby obtaining latent spatial features; the latent spatial features include corrected features after aligning and fusing pixel data from multiple viewpoints according to the geometric position of the object under test; and to map the latent spatial features according to the target viewpoint to obtain corrected geometric features under the target viewpoint.

8. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-6.