Multi-modal fusion product quality detection method, device and equipment and storage medium
Through the multimodal data fusion technology of event cameras and lidar, combined with depth information for adaptive partitioning and boundary segmentation, and using fuzzy theory and low-rank matrix reduction algorithm to optimize image processing, the problems of low precision and low efficiency in traditional detection methods are solved, and more efficient and reliable product quality detection is achieved.
Patent Information
- Application Number
- CN202510626674.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional product quality inspection methods suffer from low detection accuracy and efficiency due to incomplete image segmentation and insufficient pixel relationship mining, making it difficult to meet the needs of real-time feedback and rapid adjustment in industrial production.
The multimodal data fusion technology of event camera and lidar is adopted, and the depth information is combined for adaptive partitioning. The boundary is segmented using the geodesic frame model, and the image is optimized through fuzzy theory and low-rank matrix reduction algorithm.
It achieves more accurate, efficient and reliable product quality detection results, overcomes the limitations of single modality data, enhances anti-interference ability and computational efficiency, and improves the accuracy and completeness of image segmentation.
Smart Images

Figure CN120688910A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of quality inspection technology, and in particular to a multimodal fusion product quality inspection method, apparatus, device, and storage medium. Background Art
[0002] With the continuous development of artificial intelligence technology, product inspection methods based on image segmentation, feature extraction, and comparative analysis are becoming mainstream. These methods use image recognition and processing technology to determine product specifications, classify and organize products, and detect surface features, thereby improving the accuracy and efficiency of product quality inspection.
[0003] Traditional product quality inspection methods primarily rely on single-modality image acquisition devices to capture product images and extract features of the target area through image segmentation techniques. However, traditional image segmentation methods fail to fully consider the edge features of the image, resulting in incomplete segmentation and the tendency to miss or misjudge subtle surface defects on the product. The inherent relationships between image data pixels are not fully explored, resulting in low segmentation accuracy. Furthermore, factors such as image noise, center point selection, and distance measurement can affect segmentation accuracy and, in turn, the effectiveness of product quality inspection. Processing complex product surface features often requires multiple iterations and complex calculations, resulting in low inspection efficiency and making it difficult to meet the needs of real-time feedback and rapid adjustment in industrial production.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a multimodal fusion product quality inspection method, device, equipment and storage medium, aiming to solve the technical problems of low accuracy and low efficiency of industrial product quality inspection caused by incomplete image segmentation and insufficient pixel relationship mining.
[0006] To achieve the above objectives, this application proposes a multimodal fusion product quality detection method, which includes:
[0007] Obtain image data and depth information of the area to be detected collected by the event camera and lidar;
[0008] Partitioning the image data according to the depth information to obtain a partition set;
[0009] Segmenting the boundary of the area to be detected based on a geodesic frame model to obtain multiple loop areas;
[0010] Optimizing the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain a target region image;
[0011] The target area image is compared with the standard quality rule paradigm, and product quality detection is completed based on the comparison result.
[0012] In one embodiment, the step of partitioning the image data according to the depth information to obtain a set of partitions includes:
[0013] determining an initial number of partitions, an initial number of super pixels, and an image gradient according to the depth information and the image data;
[0014] Adjusting the initial partition number and the initial super pixel number by using a morphological closing reconstruction algorithm and the image gradient to obtain a target initial partition number and a target super pixel number;
[0015] Performing fuzzy processing on the target initial partition number and the target super pixel number to obtain a fuzzy member set and a group center point set;
[0016] Partitioning is performed according to the fuzzy member set and the group center point set to obtain a partition set.
[0017] In one embodiment, the step of segmenting the boundary of the area to be detected based on the geodesic frame model to obtain multiple loop areas includes:
[0018] Obtain boundary thresholds based on the geodesic frame model;
[0019] Dividing the area to be detected to obtain boundary segments, and determining a curve neighborhood of the boundary segments according to the boundary threshold;
[0020] determining a target path adjacent to the boundary based on the boundary segment and the curve neighborhood;
[0021] A gradient path is determined according to the target path, and neighborhood segments are determined according to the gradient path to obtain multiple loop regions.
[0022] In one embodiment, the step of determining a target path adjacent to a boundary based on the boundary segment and the curve neighborhood includes:
[0023] Calculating a weighted distance based on the boundary segment and the curve neighborhood;
[0024] Calculating geodesic distance according to the weighted distance;
[0025] A connection domain is determined according to the geodesic distance, and a target path of adjacent boundaries is determined according to the connection domain.
[0026] In one embodiment, the step of optimizing the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain the target region image includes:
[0027] Performing fuzzy equivalence partitioning on the partition set according to fuzzy theory, and constructing a fuzzy relationship matrix;
[0028] According to a low-rank matrix reduction algorithm, the fuzzy relationship matrix is reduced to generate a low-rank pixel matrix and a spatial relationship matrix;
[0029] The partition images of the partition set are optimized according to the low-rank pixel matrix, the spatial relationship matrix, and the multiple loop regions to obtain a target region image.
[0030] In one embodiment, the step of comparing the target area image with the standard quality rule paradigm and completing product quality inspection based on the comparison result includes:
[0031] Comparing the target area image with the standard quality rule paradigm, and calculating a matching score based on the comparison result;
[0032] When the matching score is less than a preset matching threshold, determining that the product is unqualified;
[0033] When the matching score is greater than or equal to the preset matching threshold, determining whether the target area image matches the standard quality rule paradigm;
[0034] When the target area image does not match the standard quality rule paradigm, generating a repair suggestion report according to the matching result;
[0035] When the target area image matches the standard quality rule pattern, the product is determined to be qualified and the product quality inspection is completed.
[0036] In one embodiment, the step of obtaining image data and depth information of the area to be detected collected by the event camera and the lidar includes:
[0037] Acquire the depth image signal of the area to be detected collected by the event camera, and acquire the high-frequency pixel signal of the area to be detected collected by the lidar;
[0038] generating a multimodal image by fusing the depth image signal and the high-frequency pixel signal;
[0039] Distortion correction is performed on the multimodal image to obtain image data and depth information.
[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a multimodal fusion product quality detection device, which includes: an acquisition module for acquiring image data and depth information of the area to be detected collected by an event camera and a lidar;
[0041] A partitioning module, configured to partition the image data according to the depth information to obtain a partition set;
[0042] A segmentation module, configured to segment the boundary of the area to be detected based on a geodesic frame model to obtain multiple loop areas;
[0043] An optimization module is used to optimize the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain a target region image;
[0044] The detection module is used to compare the target area image with the standard quality rule paradigm and complete product quality detection based on the comparison result.
[0045] In addition, to achieve the above-mentioned purpose, the present application also proposes a multimodal fusion product quality detection device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multimodal fusion product quality detection method as described above.
[0046] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the multimodal fusion product quality detection method described above are implemented.
[0047] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the multimodal fusion product quality detection method as described above.
[0048] One or more technical solutions proposed in this application have at least the following technical effects:
[0049] By utilizing multimodal data fusion technology from event cameras and lidar, adaptively partitioning the image using depth information, segmenting the boundary using a geodesic frame model, and optimizing the image processing through fuzzy theory and low-rank matrix reduction algorithms, this multimodal data fusion technology, combining the advantages of event cameras and lidar, overcomes the limitations of single-modal data and provides a more comprehensive and rich information foundation for product quality inspection. The application of adaptive partitioning and the geodesic frame model enables more accurate image segmentation and more complete boundary processing, effectively avoiding detection errors caused by inaccurate image segmentation in traditional methods. Fuzzy theory and low-rank matrix reduction algorithms further optimize the image processing process, enhancing the algorithm's interference tolerance and computational efficiency. This overcomes the low accuracy, low efficiency, and weak interference tolerance that plague traditional product quality inspection methods. Compared with existing technologies, it achieves more accurate, efficient, and reliable product quality inspection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0052] Figure 1 A flowchart of the first embodiment of the multimodal fusion product quality detection method of this application is provided;
[0053] Figure 2 A schematic diagram of a group boundary low-rank mechanism model detection network structure provided in Example 1 of the multimodal fusion product quality detection method of this application;
[0054] Figure 3 A flowchart of the second embodiment of the multimodal fusion product quality detection method of this application is provided;
[0055] Figure 4 A schematic diagram of a product multimodal acquisition and imaging network structure provided in Example 2 of the multimodal fusion product quality detection method of this application;
[0056] Figure 5 A schematic diagram of a simplified process of the multimodal fusion product quality detection method provided in Example 2 of this application;
[0057] Figure 6This is a schematic diagram of the module structure of the multimodal fusion product quality inspection device according to an embodiment of the present application;
[0058] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the multimodal fusion product quality detection method in the embodiment of the present application.
[0059] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0060] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0061] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0062] The main solution of the embodiment of the present application is: obtaining image data and depth information of the area to be detected collected by the event camera and the lidar; partitioning the image data according to the depth information to obtain a partition set; segmenting the boundary of the area to be detected based on the geodesic frame model to obtain multiple loop regions; optimizing the partition image of the partition set according to fuzzy theory, low-rank matrix reduction algorithm and the multiple loop regions to obtain a target area image; comparing the target area image with the standard quality rule paradigm, and completing product quality inspection based on the comparison results.
[0063] In this embodiment, for ease of description, the following description is made by taking a product quality inspection device for identifying multimodal fusion as the execution subject.
[0064] Due to the low accuracy and low efficiency of industrial product quality inspection caused by incomplete image segmentation and insufficient pixel relationship mining in the existing technology, this application provides a solution. By adopting the multimodal data fusion technology of event camera and lidar, adaptive partitioning is performed in combination with depth information, the boundary is segmented using the geodesic frame model, and the image is optimized through fuzzy theory and low-rank matrix reduction algorithm. The multimodal data fusion technology combines the advantages of event camera and lidar, overcomes the limitations of single-modal data, and provides a more comprehensive and rich information foundation for product quality inspection. The application of adaptive partitioning and geodesic frame model makes image segmentation more accurate and boundary processing more complete, effectively avoiding the detection errors caused by inaccurate image segmentation in traditional methods. Fuzzy theory and low-rank matrix reduction algorithm further optimize the image processing process and enhance the algorithm's anti-interference ability and computational efficiency. It solves the problems of low detection accuracy, low efficiency and weak anti-interference ability in traditional product quality inspection. Compared with the existing technology, it achieves more accurate, efficient and reliable product quality inspection results.
[0065] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of performing the above functions, multimodal fusion product quality inspection equipment, etc. The following uses multimodal fusion product quality inspection equipment as an example to illustrate this embodiment and the following embodiments.
[0066] Based on this, the embodiment of the present application provides a product quality detection method of multimodal fusion, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the multimodal fusion product quality detection method of this application.
[0067] In this embodiment, the multimodal fusion product quality detection method includes steps S10 to S50:
[0068] Step S10, acquiring image data and depth information of the area to be detected collected by the event camera and the lidar;
[0069] It's important to note that an event camera is a visual acquisition device that operates differently from traditional cameras. While traditional cameras capture images at a fixed frame rate, event cameras only record pixel changes when the brightness of a pixel exceeds a set threshold. This approach offers the advantages of low scanning latency and no motion blur, but the collected pixel data is sparse and incomplete. For example, when inspecting the surface quality of fast-moving industrial parts, images captured by traditional cameras may be blurred due to the movement of the object, but event cameras can clearly capture real-time changes in pixels on the component's surface.
[0070] LiDAR (LiDAR) is a device that uses the time difference between emitting a laser beam and receiving the reflected light to accurately measure the distance to a target object. It can obtain highly accurate depth information, which can be used to construct a three-dimensional model of an object, helping people understand its spatial structure and shape.
[0071] Furthermore, image data refers to the collection of visual information about the inspection area captured by the event camera, recorded and stored in pixel form. Image data includes appearance characteristics such as color and texture of the inspection area, and serves as the foundation for subsequent product quality testing.
[0072] Additionally, depth information, acquired by the LiDAR, represents the distance between each point on an object within the detection area and the LiDAR. This adds a third dimension to the image, providing a more intuitive understanding of the object's spatial position and shape, enabling more accurate analysis of surface features and defects.
[0073] Understandably, while event cameras offer low scanning latency and zero motion blur, their pixel data is sparse and incomplete. Combining them with lidar to obtain depth information allows for 3D multi-dimensional holographic imaging, removing distortion caused by slow response times. This complements the event camera, enabling multimodal fusion to characterize surface features from different angles, capturing complete and detailed surface information about the product under inspection for subsequent processing. The event camera and lidar collaborate to collect information about the inspection area. The event camera captures pixel-level brightness changes, generating image data reflecting surface details. The lidar simultaneously acquires the spatial depth of each point in the area, generating a depth map containing this spatial information. For example, in the inspection of automotive metal exteriors, the event camera can capture surface details such as scratches and paint peeling, while the lidar measures the height fluctuations of deformed areas. Fusion of these two data types provides comprehensive input for subsequent image segmentation and boundary analysis.
[0074] Step S20, partitioning the image data according to the depth information to obtain a partition set;
[0075] It should be noted that partitioning refers to the process of dividing an entire image into several regions according to certain rules. These rules can be based on pixel grayscale, texture, depth, edge features, and other characteristics. The purpose is to group image blocks with similar characteristics for easier processing and analysis.
[0076] Furthermore, a partition set refers to a collection of multiple sub-regions obtained after the image is segmented. Each partition represents the local information of a certain area in the image and is relatively independent. Multiple partitions are combined to form the entire original image.
[0077] As you can understand, the image data is processed using depth information acquired from the LiDAR to divide the image into multiple partitions. Because depth information can reflect the spatial differences between different regions, using changes in depth values as a segmentation basis can effectively separate parts of the image at different spatial levels or structures. Ultimately, a set of partitions containing multiple subregions is formed, paving the way for subsequent boundary modeling and optimization.
[0078] In a feasible implementation, step S20 may include steps S21 to S24:
[0079] Step S21, determining the initial number of partitions, the initial number of super pixels, and the image gradient according to the depth information and the image data;
[0080] It should be noted that the number of initial partitions refers to the number of regions preset when the original image is initially divided. This number determines the coarseness of the image segmentation in the initial stage and is the basis for subsequent image refinement processing.
[0081] Furthermore, the initial number of superpixels refers to the number of superpixel blocks into which the image is divided. Superpixels are larger than traditional pixels and have consistent color and texture. They are used to improve image processing efficiency and segmentation quality. A larger number indicates finer image segmentation, which better captures edges and details.
[0082] It is understood that the parameters for the initial image segmentation are calculated using the combined features of depth information and image data. The number of initial partitions is set by analyzing the local variation range of depth in the image so that each region has a relatively consistent spatial hierarchy.
[0083] It's also important to note that image gradients are the speed and direction of grayscale changes in an image, and are used to measure the strength of edges or contours within an image. High gradient values typically correspond to object edges or areas of texture change, making them an important basis for image segmentation and boundary detection.
[0084] It's understandable that using image texture and color information to calculate image gradients and estimate the number of superpixels allows for more refined processing of pixels at the edge. For example, when inspecting a part with a complex surface, areas with distinct depth levels are divided into multiple partitions, while contours with clear edges are enhanced through denser superpixel extraction to enhance segmentation sensitivity.
[0085] Step S22, adjusting the initial number of partitions and the initial number of super pixels by using a morphological closing reconstruction algorithm and the image gradient to obtain a target number of initial partitions and a target number of super pixels;
[0086] It's important to note that the morphological closure reconstruction algorithm is a structural restoration method based on image morphology. It fills small holes in an image through dilation and erosion operations while preserving the original edge shape. This algorithm can eliminate noise and enhance boundaries when processing images, and is widely used for partition contour correction and structural cleanup.
[0087] Furthermore, the target initial partition number is the corrected partition number under the guidance of morphological processing and image gradient. Compared with the initial value, it is more consistent with the actual structural changes of the image and has higher structural consistency and boundary clarity.
[0088] In addition, the target super-pixel number is the number of refined pixel units determined based on image gradient and morphological analysis, which is more suitable for processing complex contour areas in the image and makes the division more reasonable.
[0089] It's understandable that the morphological closure reconstruction algorithm introduces structural corrections to the initial partitions set, eliminating partition anomalies caused by image noise or depth errors. Simultaneously, it incorporates image gradient information to further adjust the number of superpixels to better reflect the actual complexity of the image's edge regions. This avoids both computational redundancy caused by overly detailed segmentation and erroneous partitioning caused by blurred boundaries. For example, when analyzing an image containing multiple seams between mechanical parts, the algorithm can effectively close boundary breaks caused by grayscale discontinuities and re-define the appropriate number of partitions.
[0090] Furthermore, adaptive reconstruction partitioning refers to the process of image partitioning, combining the image's depth information, texture features and gradient changes, dynamically adjusting the number of partitions and the number of super pixels, and introducing fuzzy logic and morphological reconstruction algorithms to perform layered refinement, structural correction and boundary flexible division on the image, thereby generating a regional division result that is more in line with the actual structure of the target. This method is different from the fixed parameter or single feature driven division method. It can adaptively adjust the partitioning strategy according to different image contents to achieve more accurate and robust image area expression, and provide a structural optimization input basis for subsequent boundary modeling and target recognition. The adaptive reconstruction partitioning function is defined as follows:
[0091]
[0092] Where, is the morphological closing reconstruction algorithm, To mark the closed structure, is the gradient operation, b i It is an embedded structural element, i is the range parameter of the structural element, and s and m in the subscript of V are both natural numbers N + , g is the acquired image gradient, f is the morphological extension reconstruction relationship f = εb obtained from gi (g) The input of the morphological closure reconstruction algorithm is as follows:
[0093]
[0094] The calculation method of the reconstruction area F is In ψ(g, s, m), s and m are in a dynamic balance with each other. When the number of clusters m is relatively large, the range of each clustering area s is relatively small. On the contrary, when the range of the clustering area s is relatively large and the number of clusters m is relatively small, the increase of s will reduce the segmented area, but the accuracy of segmentation will decrease. Therefore, it is necessary to reasonably grasp the number of cluster sets to be planned for partition reconstruction, and extract superpixels from the clustering areas of pixels with similar features in adjacent positions. The superpixels are extracted based on the ARP algorithm, which enhances the anisotropic features to strengthen the edges and improves the anti-interference performance.
[0095] Step S23, performing fuzzy processing on the target initial partition number and the target super pixel number to obtain a fuzzy member set and a group center point set;
[0096] It should be noted that fuzzification involves converting originally precise numerical parameters into uncertain representations under a membership function. Its essence is to enhance the system's tolerance for fuzzy boundaries, making the model more robust when processing natural images.
[0097] In addition, the fuzzy member set is a set of probabilities of each image element belonging to a certain partition obtained through fuzzy calculation. It reflects the possibility that a pixel belongs to multiple partitions and is used to describe uncertain boundary areas.
[0098] Furthermore, the group center point set refers to the center position coordinates of each partition or superpixel determined under fuzzy processing. These center points serve as a reference for the clustering of each fuzzy member and determine the center of gravity and range of the final region division.
[0099] As can be understood, fuzzy logic conversion is performed on the target initial number of partitions and the number of super pixels to generate membership information for each pixel relative to multiple regions, forming a fuzzy membership set. Based on these membership relationships, the cluster centers of each region are extracted to form a group center set, which is used to guide the natural transition of partition boundaries.
[0100] For fuzzy data processing, it is necessary to cluster the data of the relevant subspaces. The clustering can be as follows: X = x 1, x 2, ...x i ...x n, xi ∈R n , R n is a dataset of image pixels, and n is the number of pixels in the image. Clustering can be calculated as follows:
[0101]
[0102] In the formula, U is defined as Relative to the cluster center point v j The fuzzy member x j A set of, each member x in the set j , define its weight as 0≤u ij ≤1, assuming m is the index of matrix U, and V is the set of cluster center points.
[0103] Step S24: partitioning is performed according to the fuzzy member set and the group center point set to obtain a partition set.
[0104] As you can understand, the final image region segmentation is performed based on the fuzzy member set and the group center point set. Each pixel is assigned to the most likely target region based on its membership. Some edge pixels can be shared by multiple partitions based on the fuzzy weight, forming a boundary transition zone. This approach preserves structural integrity while reducing partition jumps. For example, when inspecting complex industrial parts, areas with polishing marks on the edges can be flexibly divided into multiple related regions, improving subsequent inspection sensitivity.
[0105] In the image clustering calculation formula, U and V are dynamically balanced. The fewer the number of elements in the cluster center point set V, the greater the cumulative sum of the weight relationship U. Conversely, the more elements in the center point set V, the smaller U. The calculation of fuzzy members is as follows:
[0106]
[0107] The cluster center point calculation formula is as follows:
[0108]
[0109] Partition U and V by iterating them t times to obtain U and V that are both in the minimum state at the same time.
[0110] Step S30, segmenting the boundary of the area to be detected based on the geodesic frame model to obtain multiple loop areas;
[0111] It's important to note that the geodesic frame model is a geometric modeling method based on geodesic distance. Geodesic distance refers to the shortest distance between two points on an object's surface along a path along the surface. In image analysis, this model is used to more naturally simulate object contours and is particularly well-suited for dealing with curved or irregularly shaped boundaries.
[0112] Additionally, boundary segmentation is the process of breaking complex edge structures into several line or curve segments, allowing for more detailed analysis of each segment's shape, orientation, and structural properties. This approach can simplify subsequent computations and enhance the semantic representation of boundaries.
[0113] It is understandable that the geodesic framework model is used to extract the boundaries of the partition set, and the boundaries are subjected to curve modeling and segmentation operations. The geodesic modeling method can more realistically express the structural characteristics of the target edge, such as curvature and undulation, and is particularly suitable for scenes with irregular edges or severe deformations. Ultimately, multiple closed boundary areas are formed, namely multi-segment loop areas, which serve as the core processing unit for the next stage of image optimization. For example, when detecting metal welds, the geodesic model is used to divide the continuous raised or collapsed boundary segments to form the welding forming area to be judged.
[0114] In a feasible implementation, step S30 may include steps S31 to S34:
[0115] Step S31, obtaining a boundary threshold according to a geodesic frame model;
[0116] It should be noted that the boundary threshold is a value set according to the geodesic distance or the degree of structural change, and is used to determine whether a certain area in the image constitutes a boundary segment or a boundary inflection point.
[0117] It is understood that a geodesic framework model is established based on image structural information, and a threshold parameter for boundary determination is calculated by analyzing changes in image surface features and path deviation. This boundary threshold is used to subsequently determine which areas in the image meet the boundary conditions. In this embodiment, the boundary threshold is a real number.
[0118] Step S32, dividing the area to be detected to obtain boundary segments, and determining the curve neighborhood of the boundary segments according to the boundary threshold;
[0119] It should be noted that boundary segments are continuous edge segments demarcated from image edges through path analysis. Each segment represents an edge region with consistent orientation or geometric features within the image structure. Decomposing boundary segments helps break down complex boundary structures into manageable sub-components.
[0120] Furthermore, a curve neighborhood is a spatial region defined around a boundary segment, capturing pixel or feature information near the boundary. Its range can be defined based on geodesic length, pixel gradient, or other edge attributes, reflecting the boundary's impact area.
[0121] It can be understood that edge analysis is performed on the area to be detected within the range of the boundary threshold, and the continuous edge is divided into multiple boundary segments. Subsequently, a curve neighborhood is set for each boundary segment, that is, the area range generated around the boundary path, to collect gradient and structural change information related to the boundary segment. i , use V i As a representative, find its curve neighborhood, where S j With S i adjacent,
[0122] According to the geodesic framework model, an adaptive geodesic distance map can be constructed The curve neighborhood is calculated as follows:
[0123] ψ i ={x∈Ω|U i ≤ζ}
[0124] Where, the curve neighborhood ψ i Is to build S i Curve neighborhood, set a real boundary threshold as the threshold ζ∈R + To estimate the geodesic distance map U i To construct a curve neighborhood Use ψ seg Construct its neighborhood, positive real number ε∈R + is a constant, and each geodesic distance is calculated as follows:
[0125]
[0126] It is understandable that each geodesic distance must comply with the formula U i , calculate S i The geodesic neighborhood U i Through multiple iterations until for point x, U i (x)>ζ, then this pixel does not belong to the curve neighbor S j , thus dividing the curve neighborhood.
[0127] Step S33, determining a target path adjacent to the boundary based on the boundary segment and the curve neighborhood;
[0128] It should be noted that a target path is a continuous path that originates from a boundary segment in the image structure, passes through its curved neighborhood, and connects to adjacent boundary segments. This path reflects the structural correlation between boundary segments and is used to infer whether structures belong to the same closed region. Adjacent boundaries are boundary segments that are spatially close or morphologically related in the image structure. These boundary segments are typically located at the ends of the same structural unit or form a partially closed curve, and are key nodes for boundary wrapping and closure analysis.
[0129] Based on the curve neighborhood of each boundary segment, we search for boundary segments that are spatially or structurally adjacent to it and calculate the optimal connecting path from this segment to the adjacent segments to form the target path. This path does not simply pursue the shortest distance but also takes into account factors such as boundary direction and texture continuity.
[0130] In a feasible implementation, step S33 may include steps S331 to S333:
[0131] Step S331, calculating a weighted distance based on the boundary segment and the curve neighborhood;
[0132] It should be noted that weighted distance is a weighted path distance measurement that comprehensively considers multi-dimensional features such as image gradient, texture continuity, and depth variation. It is used to represent the connection cost between two regions. The smaller the weight, the more natural the connection path between the two boundaries.
[0133] It can be understood that for each boundary segment and its corresponding curve neighborhood area, the weighted distance between it and other possible adjacent boundaries is calculated. This calculation process comprehensively considers multiple factors such as edge direction consistency, image gradient continuity, and depth change smoothness, and integrates them into a single value in the form of a weighted factor to represent the difficulty of establishing a connection between any two boundary segments. The weighted distance between edge paths is calculated to calculate the shortest distance between two neighboring edges. Use G = (ν, ε) to represent the boundary, ν is the set of pixel points, and ε is the edge e connecting the points. i,j ∈ε,w i,j Represents its weighted distance. The geodesic distance between points is the shortest distance between pixels in the gradient image. The weighted distance is calculated as follows:
[0134]
[0135] Where, The minimization path model defines a curved path γ, γ:[0,1]→Ω, which is the gradient path from a starting point to an end point in the image domain. The distance of the gradient path can be calculated by integrating the pixel gradients, where γ′ = dγ / dμ is the first-order derivative of r. The weighted distance between the source point s and the end point x can be calculated using the weighted distance formula to calculate the minimum value of the possible path L.
[0136] Step S332, calculating the geodesic distance according to the weighted distance;
[0137] It's important to note that geodesic distance is the shortest path length between two points on an image or structure surface within a constrained space. This distance accounts for the true shape of image structures, such as curvature and occlusion, ensuring that paths more closely follow the actual edge direction, making it a more realistic and effective path metric.
[0138] It can be understood that based on the weighted distance matrix, the shortest path network in the image structure space is constructed, and the geodesic distance between boundary segments is further calculated. The geodesic distance not only considers the geometric position between boundaries, but also integrates weight information, making the path more consistent with the actual connection relationship of the image structure. The calculation of the geodesic distance is as follows:
[0139]
[0140] Where g s,x is the geodesic distance between s and x, S i and S h For the boundary segment.
[0141] Step S333: determining a connection domain according to the geodesic distance, and determining a target path of adjacent boundaries according to the connection domain.
[0142] It should be noted that the connection domain refers to the spatial connection range constructed by boundary segments whose geodesic distance meets the set conditions. This domain indicates the potential connection between two or more boundary segments in the image and is a key component in establishing a closed structure.
[0143] It can be understood that the calculated geodesic distance is used to filter out all boundary segment pairs with geodesic distances below a threshold to construct a connection domain. Subsequently, the optimal path in each connection domain is selected as the target path to achieve structural connection between adjacent boundaries.
[0144] For the boundary segment S i and S j , taking the geodesic distance as the parameter to represent its distance, define a domain segmentation and connection domain F:M×R n (n=2,3). Further, for the starting point boundary S i→0, end point S j →1, its weight distance can be obtained by L F calculate, S can be calculated by the following optimal path calculation formula i to S j The best path between them is obtained by minimizing L F To find the shortest path between boundaries. The best path calculation formula is as follows:
[0145]
[0146] S i As the baseline, the geodesic distance Find a path that fixes the neighborhood boundaries.
[0147] Furthermore, a structure-adaptive boundary secant is constructed. The adaptive boundary secant is to first set a marker point R in the target area, then connect the point R and the target boundary to form a secant line, and intersect the boundary only once. str (x): = φ(x)δ(x) Calculate the intersection line. In the process of calculating the intersection line, a structure-adaptive edge intersection line ψ: = ψ can be calculated based on the image edge information provided by the boundary solution in the previous step. str .
[0148] In the delivery line calculation formula,
[0149] Where φ: Is the edge symbol, δ:Ω→{1,∞}, is a set symbol of edge points. The best edge intersection point can be obtained by calculation. The geodesic distance path end point b is calculated as follows:
[0150]
[0151] Where, is the end point of the shortest gradient path from the marked point to the target boundary, and U can be obtained p The minimum value of (x) is obtained.
[0152] Step S34 , determining a gradient path according to the target path, and determining neighborhood segments according to the gradient path to obtain multiple loop regions.
[0153] It's important to note that a gradient path is a continuous path of pixels along the target path, extracted based on the image's grayscale or color change trends. The pixels along this path have relatively consistent or continuous gradient changes, which is an important basis for determining edge direction and segmenting regions.
[0154] Furthermore, neighborhood segmentation is a structural unit formed by dividing the surrounding neighborhood around the gradient path, reflecting the gradient characteristics of the structure, material, or color in the image. This segmentation is often used to construct continuous boundaries of complex structures.
[0155] A multi-segment loop region is a closed or nearly closed structure formed by combining multiple gradient paths and their neighborhood segments. It is used to delineate areas in an image that may contain targets and has spatial coherence in structure.
[0156] It can be understood that based on the target path, the gradient path is extracted by analyzing the gradient changes of the image along the path. Subsequently, the surrounding area is divided with this path as the main axis to form neighborhood segments. After connecting multiple neighborhood segments, closed or nearly closed multi-segment loop regions are generated, which are subsequently used as target detection areas.
[0157] Step S40, optimizing the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain a target region image;
[0158] It's important to note that fuzzy theory is a mathematical tool for processing ambiguous and uncertain information. It uses membership functions to describe the state of an object probabilistically. In image processing, it can be used to describe areas with unclear edges and slow grayscale transitions, making boundary analysis more flexible and robust.
[0159] Low-rank matrix reduction algorithms, which convert image matrices into low-rank structures with higher information density, are primarily used for denoising, compression, and feature extraction. They preserve the image's primary structural information while reducing background noise. Common methods include singular value decomposition and robust principal component analysis.
[0160] Furthermore, the target area image refers to a target object image with high quality, clear structure and clear boundaries extracted after image processing, which serves as an input basis for subsequent quality judgment.
[0161] As can be understood, fuzzy theory and a low-rank matrix reduction algorithm are used to optimize the image quality of each segment based on multiple loop regions. Fuzzy theory improves the responsiveness of boundary regions, ensuring that blurred edges are not missed; while the low-rank algorithm removes noise and redundant data, preserving the target's salient structure. The optimization results focus on the target region, creating an image of the target region with complete structure and minimal interference. For example, applying this processing to a defective wheel hub significantly improves the clarity of the crack boundary and eliminates the interference of complex background textures.
[0162] In a feasible implementation, step S40 may include steps S41 to S43:
[0163] Step S41, performing fuzzy equivalent partitioning on the partitions of the partition set according to fuzzy theory, and constructing a fuzzy relationship matrix;
[0164] It's important to note that fuzzy theory is a mathematical theory used to describe and address uncertainty and ambiguity. Its core concept is "degree of membership." By expanding the notion of an element's attribution to a property from "belongs" or "does not belong" to a continuous value between 0 and 1, it's suitable for image processing tasks with unclear boundaries or significant transition regions.
[0165] Furthermore, fuzzy equivalence partitioning divides pixels or regions with similar characteristics in an image into fuzzy sets based on fuzzy relationships. Each pixel no longer belongs to a single partition, but rather to multiple sets to varying degrees. This flexible image region partitioning method adapts to the natural transition characteristics of image structure.
[0166] In addition, the fuzzy relationship matrix is a matrix that represents the similarity between any two image elements. Each value in the matrix represents the strength of the fuzzy relationship between the two elements, that is, the degree of similarity between them in the sense of fuzzy equivalence, and the value range is 0 to 1.
[0167] As can be understood, fuzzy theory is introduced to perform fuzzy equivalence partitioning on the partition set. By comparing the similarities between partitions in terms of texture, grayscale, depth, and other aspects, fuzzy membership is calculated and a fuzzy relationship matrix is constructed. This matrix quantifies the degree of fuzzy connectivity between different image regions, providing a flexible expression foundation for subsequent structural optimization.
[0168] Step S42, performing rank reduction processing on the fuzzy relationship matrix according to a low rank matrix reduction algorithm to generate a low rank pixel matrix and a spatial relationship matrix;
[0169] It should be noted that the low-rank pixel matrix is the main pixel composition information retained after structural compression of the original image information. It represents the key areas and features in the image, eliminates redundancy and noise interference, and is the core abstraction of the image visual content.
[0170] In addition, the spatial relationship matrix is used to represent the strength of spatial connections between different regions or pixels in the image. This matrix is generated in conjunction with the low-rank pixel matrix during the rank reduction process, preserving the spatial proximity and topological information of the image structure and assisting in boundary preservation and structure reconstruction.
[0171] It can be understood that the generated fuzzy relationship matrix is used as input and a low-rank matrix reduction algorithm is used to perform information compression and feature extraction. Through matrix decomposition, a low-rank pixel matrix and a spatial relationship matrix are obtained. The former extracts the core visual information of the image, while the latter preserves the topological connections between spatial structures.
[0172] Based on fuzzy theory, the relationship between pixels is clustered in descending order, and fuzzy theory is combined with price reduction processing. The relationship between similar pixels is linear, pixel boundaries are weakened, and the internal connection between pixels is strengthened. Starting from the global structure of pixel data, according to the CRA (Change Rank Adapter) algorithm, the internal correlation between data is strengthened through the fuzzy regularization paradigm to reasonably determine the minimized rank. By processing the relationship between data, the matrix rank is set to minimize and establish the objective function. From the perspective of global structure, the edge coupling degree is eliminated, and a pixel data regression matrix A is constructed in combination with the cluster center V to reduce the rank of the pixel matrix. The calculation of the low-rank pixel matrix is as follows:
[0173]
[0174] Where A is the data space formed by linear expansion, and the spatial relationship matrix U is constructed by expanding the member relationship. By optimizing the relationship between pixel matrices and continuously iterating, the best low-order structure is calculated to obtain the smallest optimized target area. The matrix of the target area is extracted to segment the target area.
[0175] Step S43: Optimizing the partition images of the partition set according to the low-rank pixel matrix, the spatial relationship matrix, and the multiple loop regions to obtain a target region image.
[0176] It can be understood that by integrating the core image structure represented by the low-rank pixel matrix and the spatial relationship matrix, and using multiple loop regions as boundary constraints, the original partitioned image is fully optimized. This optimization process combines the core image content with the boundary topology, removes redundant pixels, strengthens the target edges, and ultimately outputs an image of the target area with clear boundaries and concentrated content.
[0177] Step S50 , comparing the target area image with the standard quality rule paradigm, and completing product quality inspection based on the comparison result.
[0178] It should be noted that the standard quality rule paradigm is a benchmark reference for product quality. It can be a high-quality image sample annotated by experts, a typical boundary model of qualified products, or a quality threshold range generated through historical data, which defines the qualified standards in terms of product appearance, shape or size.
[0179] It is understood that the target area image is compared item by item with the standard quality rule paradigm. Comparison methods may include image template matching, edge overlap analysis, and texture similarity calculation. By analyzing the structural differences between the two, it is determined whether the target has quality issues such as deformation, defects, and errors. By establishing a final inspection mechanism that compares with the standard image, automatic product quality identification and judgment can be achieved, avoiding the subjectivity and fatigue of manual inspection, and improving product consistency and inspection efficiency.
[0180] In a feasible implementation, step S50 may include steps S51 to S55:
[0181] Step S51, comparing the target area image with the standard quality rule paradigm, and calculating a matching score based on the comparison result;
[0182] It should be noted that the matching score is a numerical indicator that measures the degree of similarity between the target image and the standard template through methods such as image feature comparison, boundary overlap analysis, and texture consistency calculation. A higher score indicates a closer match to the standard template, while a lower score indicates a greater deviation.
[0183] It is understood that the target area image is compared against the standard quality rule paradigm for structural features and visual content. The comparison process may include edge coincidence calculation, grayscale / color histogram matching, and local texture difference analysis. The comparison result is quantified into a matching score, reflecting the degree of consistency between the detected object and the standard template.
[0184] Step S52: when the matching score is less than a preset matching threshold, determining that the product is unqualified;
[0185] It should be noted that the preset matching threshold refers to the minimum score requirement for image matching set by the system or quality standard. If the score of the target image compared with the paradigm is lower than this threshold, it means that the structural difference is too large and does not meet the quality requirements.
[0186] It is understood that when the calculated matching score falls below a set threshold, the product is deemed unqualified. This determination means that the detected image deviates significantly from the standard pattern in terms of structural position, boundary shape, or image texture.
[0187] Step S53, when the matching score is greater than or equal to the preset matching threshold, determining whether the target area image matches the standard quality rule paradigm;
[0188] It can be understood that when the matching score between the target area image and the standard quality rule paradigm is greater than or equal to the preset matching threshold, it means that the image has passed the preliminary quality assessment. In order to ensure the rigor of the judgment and avoid misjudgment due to scoring errors or neglect of local details, this step does not immediately conclude that the product is qualified, but rather needs to determine whether the image matches the standard quality rule paradigm.
[0189] Step S54, when the target area image does not match the standard quality rule paradigm, generating a repair suggestion report according to the matching result;
[0190] It is understandable that if the image of the target area is judged to have met the score requirements, but there is a structural inconsistency with the standard quality rule paradigm, it will be deemed as a mismatch. In this case, to ensure that subsequent processing is timely and accurate, this step will generate a repair recommendation report based on the specific differences. The report content may include: marking inconsistent areas, describing the type of problem, such as offset, insufficient size, edge jaggedness, etc., and proposing correction methods based on historical data or expert rules. For example, when inspecting a plastic shell, although the overall shape is consistent, the size of a certain interface is slightly smaller. This step will report the problem location and recommend re-molding or trimming.
[0191] Step S55 , when the target area image matches the standard quality rule pattern, it is determined that the product is qualified and the product quality inspection is completed.
[0192] It is understandable that when the target area image is confirmed to be completely matched with the standard quality rule paradigm, it means that the product has correct structure, clear boundaries, and no abnormal texture in key parts, and meets all quality standard requirements. At this time, the product is judged to be qualified and the image inspection process of the current product is terminated. The inspection results will be recorded synchronously in the database and can be used for subsequent statistical analysis or quality tracking. Directly connecting the standard matching judgment with production decision-making can improve the efficiency and consistency of quality control and provide data support for rapid release and full process traceability on the production line. Figure 2 , Figure 2 This is a schematic diagram of the network structure of the group boundary low-rank mechanism model detection method of the first embodiment of the multimodal fusion product quality detection method of this application.
[0193] like Figure 2As shown in the figure, image acquisition and preprocessing are performed, including obtaining image data and depth information from the camera and lidar, and then multimodal feature fusion is performed on the image, the target area to be detected is marked, and it is processed according to adaptive reconstruction partitioning. Then, super pixels in the area are extracted and blurred. After feature extraction, regional pixel cohesion is performed, a geodesic framework model is constructed to determine the proximal curve neighborhood and perform domain boundary segmentation, a structurally adaptive boundary secant is constructed, the optimal path is found according to the geodesic model, the marked points in the domain are connected to the target boundary to ensure the integrity of the boundary segmentation, the clustered pixel relationship is strengthened based on fuzzy theory, and the image pixel vector matrix is adapted to the reduced order structure. Continuous iteration and convergence are performed to extract the target area, matching the standard product quality rule paradigm to determine whether the product surface characteristics meet the quality standards, and real-time feedback is provided to adjust the production and processing process. The entire process includes model area construction, super pixel extraction, fuzzy processing, clustering, geodesic framework model construction, curve neighborhood calculation, boundary secant construction, pixel vector matrix reduced order adaptation, product quality rule paradigm matching and product qualification evaluation, and finally outputs the product qualification evaluation result.
[0194] This embodiment provides a multimodal fusion product quality inspection method. By adopting technical means such as multimodal data fusion of event cameras and lidars, adaptive partitioning based on depth information, boundary segmentation processing of geodesic frame models, and image optimization processing using fuzzy theory and low-rank matrix reduction algorithms, it solves the technical problems of low detection accuracy, low efficiency, and weak anti-interference ability caused by incomplete image segmentation and insufficient pixel relationship mining in traditional product quality inspection, achieves the beneficial effects of improving detection accuracy, enhancing anti-interference ability, and improving detection efficiency, and provides an efficient, accurate, and reliable solution for industrial product quality inspection.
[0195] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 Step S10 of the multimodal fusion product quality detection method includes steps S11 to S13:
[0196] Step S11, obtaining a depth image signal of the area to be detected collected by the event camera, and obtaining a high-frequency pixel signal of the area to be detected collected by the laser radar;
[0197] It should be noted that the depth image signal is a signal that combines time information and pixel response and is indirectly generated by the event camera. It reflects the positional relationship between the pixels in the image in the event response sequence and can be used to estimate the edge, shape and relative spatial position of the object.
[0198] High-frequency pixel signals, on the other hand, refer to areas of an image where grayscale changes dramatically. These typically correspond to object edges, areas with significant texture changes, or areas with sudden local structural changes. High-frequency pixel signals collected by LiDAR can be used for detail recognition and spatial contour analysis.
[0199] As you can see, the two types of sensors generate complementary image information. The event camera captures depth-based image signals representing brightness changes in the area being detected, reflecting object deformation or motion boundaries. The lidar collects high-frequency pixel signals from its spatial structure, providing detailed data on areas with high depth variations.
[0200] Step S12, fusing the depth image signal and the high-frequency pixel signal to generate a multimodal image;
[0201] It should be noted that multimodal images refer to images generated by the fusion of multiple perceptual modalities such as visual modality, spatial modality, and temporal modality. They usually contain richer feature dimensions, such as grayscale, edge, depth, and texture information, and perform better in structural perception and detail recognition.
[0202] It is understood that the event camera depth image signal obtained in step S11 is jointly processed with the high-frequency pixel signal from the lidar, and the two types of image data are fused into a multimodal image in the same coordinate system through spatiotemporal registration and feature alignment. The fusion process can use methods such as feature map overlay, weighted fusion, or channel merging to ensure that the output image takes into account the advantages of both types of data in terms of spatial resolution and texture expression.
[0203] Step S13: performing distortion correction on the multimodal image to obtain image data and depth information.
[0204] It should be noted that distortion correction is the process of repairing geometric image distortion caused by the camera's optical system, including barrel distortion and pincushion distortion. The corrected image maintains the proportions of the real-world spatial structure, which is a prerequisite for precise positioning and measurement.
[0205] It can be understood that the optical distortion model is applied to the fused multimodal image to align the image structure with the actual physical proportions. After correction, the output is image data containing 2D visual features and corresponding 3D depth information, which serves as the raw input for subsequent image analysis and target modeling.
[0206] Reference Figure 4 , Figure 4 This is a schematic diagram of the product multimodal acquisition and imaging network structure of the first embodiment of the multimodal fusion product quality detection method of this application.
[0207] The left side shows the combined modal light imaging component. This component uses laser illumination and structured modulation to trigger an event camera to capture high-frequency frame images, recording the temporal fluctuations in brightness to form high-frequency template image information. The acquired image signals are fed into the multi-dimensional stacking imaging module, where 3D depth information is generated through multi-dimensional scanning, denoising, and holographic imaging. The 3D image information then enters the surface feature fusion stage. The image sequentially passes through multiple perspective processing modules, designated SurfacePerspective1 through SurfacePerspective5, to extract surface information from different angles. The perspective information is combined in the fusion layer to generate a multi-modal feature fused image. The fused image then enters the annotation processor module, which performs feature recognition and quality assessment for annotations 1 through n. The final output is fully annotated image data, which can be used for product surface inspection or subsequent decision-making. The overall process integrates structured light imaging, high-frequency feature extraction, 3D reconstruction, multi-angle fusion, and intelligent annotation.
[0208] This embodiment provides a multimodal fusion product quality inspection method. By adopting technical means such as the fusion processing of the depth image signal of the event camera and the high-frequency pixel signal of the lidar, as well as distortion correction of the multimodal image, it solves the technical problems of structural distortion and detail loss caused by insufficient single modal information and image distortion in traditional image acquisition and fusion, and achieves the beneficial effects of improving the integrity and accuracy of image data, enhancing detail recognition capabilities, and improving the accuracy of subsequent image analysis and target modeling, providing a higher-quality image data foundation for product quality inspection.
[0209] For example, in order to help understand the implementation process of the multimodal fusion product quality detection method obtained by combining this embodiment with the above embodiment 1, please refer to Figure 5 , Figure 5 A brief flowchart of a multimodal fusion product quality inspection method is provided. Specifically:
[0210] Industrial products are captured using visible light cameras and lidar sensors. Event cameras capture product images within their visual range, acquiring high-frequency frame image information. Lidar imaging captures a point cloud image of the product, further yielding 3D depth image information. Multimodal feature fusion of the event camera image information and the lidar depth information is performed. The fused image then enters the fuzzy cohesion and order reduction adaptation phase. After processing, the product is determined to be qualified. If the determination is negative, the feedback reconstruction production process begins. If the determination is positive, the process ends. The dashed line on the right represents the image region analysis process. First, the areas to be detected and identified are annotated. Then, adaptive reconstruction partitioning and fuzzy superpixel cohesion processing are performed. Close neighborhoods and boundary profiles are determined. Correlations between image pixel data are enhanced, and order reduction adaptation is performed. Image features are extracted and compared with quality rule paradigm templates to complete the determination.
[0211] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the multimodal fusion product quality detection method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0212] This application also provides a multi-modal fusion product quality detection device, please refer to Figure 6 , the multimodal fusion product quality detection device includes:
[0213] An acquisition module 10 is used to acquire image data and depth information of the area to be detected collected by the event camera and the lidar;
[0214] A partitioning module 20, configured to partition the image data according to the depth information to obtain a partition set;
[0215] A segmentation module 30 is used to segment the boundary of the area to be detected based on a geodesic frame model to obtain multiple loop areas;
[0216] An optimization module 40 is configured to optimize the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain a target region image;
[0217] The detection module 50 is used to compare the target area image with the standard quality rule paradigm and perform product quality detection based on the comparison result.
[0218] The multimodal fusion product quality inspection device provided in this application, which utilizes the multimodal fusion product quality inspection method described in the aforementioned embodiments, can address the technical issues of low accuracy and inefficiency in industrial product quality inspections caused by incomplete image segmentation and insufficient pixel relationship mining. Compared to the prior art, the beneficial effects of the multimodal fusion product quality inspection device provided in this application are the same as those of the multimodal fusion product quality inspection method described in the aforementioned embodiments. Other technical features of the multimodal fusion product quality inspection device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0219] In one embodiment, the partitioning module 20 is further used to determine the initial number of partitions, the initial number of super pixels and the image gradient based on the depth information and the image data; adjust the initial number of partitions and the initial number of super pixels through a morphological closing reconstruction algorithm and the image gradient to obtain a target initial number of partitions and a target number of super pixels; fuzzy the target initial number of partitions and the target number of super pixels to obtain a fuzzy member set and a group center point set; and partition according to the fuzzy member set and the group center point set to obtain a partition set.
[0220] In one embodiment, the segmentation module 30 is further used to obtain a boundary threshold based on a geodesic frame model; divide the area to be detected to obtain boundary segments, and determine the curve neighborhood of the boundary segment based on the boundary threshold; determine the target path of the adjacent boundary based on the boundary segment and the curve neighborhood; determine the gradient path based on the target path, and determine the neighborhood segmentation based on the gradient path to obtain multiple loop areas.
[0221] In one embodiment, the segmentation module 30 is further used to calculate a weighted distance based on the boundary segment and the curve neighborhood; calculate a geodesic distance based on the weighted distance; determine a connection domain based on the geodesic distance, and determine a target path of an adjacent boundary based on the connection domain.
[0222] In one embodiment, the optimization module 40 is also used to perform fuzzy equivalent division on the partitions of the partition set according to fuzzy theory to construct a fuzzy relationship matrix; reduce the fuzzy relationship matrix according to a low-rank matrix reduction algorithm to generate a low-rank pixel matrix and a spatial relationship matrix; optimize the partition image of the partition set according to the low-rank pixel matrix, the spatial relationship matrix and the multiple loop regions to obtain a target area image.
[0223] In one embodiment, the detection module 50 is further used to compare the target area image with the standard quality rule paradigm, and calculate a matching score based on the comparison result; when the matching score is less than a preset matching threshold, the product is determined to be unqualified; when the matching score is greater than or equal to the preset matching threshold, it is determined whether the target area image and the standard quality rule paradigm match; when the target area image and the standard quality rule paradigm do not match, a repair suggestion report is generated based on the matching result; when the target area image and the standard quality rule paradigm match, the product is determined to be qualified, and the product quality inspection is completed.
[0224] In one embodiment, the acquisition module 10 is also used to obtain the depth image signal of the area to be detected collected by the event camera, and obtain the high-frequency pixel signal of the area to be detected collected by the lidar; fuse the depth image signal and the high-frequency pixel signal to generate a multimodal image; and perform distortion correction on the multimodal image to obtain image data and depth information.
[0225] The present application provides a multimodal fusion product quality detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the multimodal fusion product quality detection method in the above-mentioned embodiment 1.
[0226] Reference below Figure 7 , which shows a schematic structural diagram of a product quality inspection device suitable for implementing the multimodal fusion of the embodiments of the present application. The multimodal fusion product quality inspection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The multimodal fusion product quality inspection device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0227] like Figure 7As shown, the multimodal fusion product quality inspection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the multimodal fusion product quality inspection device are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the multimodal fusion product quality inspection device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a multimodal fusion product quality inspection device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have instead.
[0228] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0229] The multimodal fusion product quality inspection device provided in this application, which utilizes the multimodal fusion product quality inspection method described in the above-mentioned embodiment, can address the technical issues of low accuracy and inefficiency in industrial product quality inspections caused by incomplete image segmentation and insufficient pixel relationship mining. Compared to the prior art, the beneficial effects of the multimodal fusion product quality inspection device provided in this application are the same as those of the multimodal fusion product quality inspection method described in the above-mentioned embodiment. The other technical features of this multimodal fusion product quality inspection device are the same as those disclosed in the above-mentioned embodiment and are not further elaborated here.
[0230] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0231] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0232] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the multimodal fusion product quality detection method in the above-mentioned embodiment.
[0233] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0234] The computer-readable storage medium may be included in the multimodal fusion product quality inspection device; or it may exist independently without being assembled into the multimodal fusion product quality inspection device.
[0235] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the multimodal fusion product quality inspection device, the multimodal fusion product quality inspection device enables the following: to obtain image data and depth information of the area to be inspected collected by the event camera and the lidar; to partition the image data according to the depth information to obtain a partition set; to segment the boundary of the area to be inspected based on the geodesic frame model to obtain multiple loop regions; to optimize the partition image of the partition set according to fuzzy theory, low-rank matrix reduction algorithm and the multiple loop regions to obtain a target area image; to compare the target area image with the standard quality rule paradigm, and to complete product quality inspection based on the comparison result.
[0236] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0237] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0238] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0239] The computer-readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned multimodal fusion product quality detection method. This computer-readable storage medium can address the technical issues of low accuracy and inefficiency in industrial product quality detection due to incomplete image segmentation and insufficient pixel relationship mining. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the multimodal fusion product quality detection method provided in the aforementioned embodiment, and are not further elaborated here.
[0240] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the multimodal fusion product quality detection method as described above.
[0241] The computer program product provided in this application can address the technical issues of low accuracy and inefficiency in industrial product quality inspection caused by incomplete image segmentation and insufficient pixel relationship mining. Compared to the prior art, the beneficial effects of the computer program product provided in this application are similar to those of the multimodal fusion product quality inspection method provided in the aforementioned embodiments, and are not further elaborated here.
[0242] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A multimodal fusion product quality detection method, characterized in that: The method comprises: Obtain image data and depth information of the area to be detected collected by the event camera and lidar; Partitioning the image data according to the depth information to obtain a partition set; Segmenting the boundary of the area to be detected based on a geodesic frame model to obtain multiple loop areas; Optimizing the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain a target region image; The target area image is compared with the standard quality rule paradigm, and product quality detection is completed based on the comparison result.
2. The method according to claim 1, wherein The step of partitioning the image data according to the depth information to obtain a partition set includes: determining an initial number of partitions, an initial number of super pixels, and an image gradient according to the depth information and the image data; Adjusting the initial partition number and the initial super pixel number by using a morphological closing reconstruction algorithm and the image gradient to obtain a target initial partition number and a target super pixel number; Performing fuzzy processing on the target initial partition number and the target super pixel number to obtain a fuzzy member set and a group center point set; Partitioning is performed according to the fuzzy member set and the group center point set to obtain a partition set.
3. The method according to claim 1, wherein The step of segmenting the boundary of the area to be detected based on the geodesic frame model to obtain multiple loop areas includes: Obtain boundary thresholds based on the geodesic frame model; Dividing the area to be detected to obtain boundary segments, and determining a curve neighborhood of the boundary segments according to the boundary threshold; determining a target path adjacent to the boundary based on the boundary segment and the curve neighborhood; A gradient path is determined according to the target path, and neighborhood segments are determined according to the gradient path to obtain multiple loop regions.
4. The method according to claim 3, wherein The step of determining a target path adjacent to a boundary based on the boundary segment and the curve neighborhood includes: Calculating a weighted distance based on the boundary segment and the curve neighborhood; Calculating geodesic distance according to the weighted distance; A connection domain is determined according to the geodesic distance, and a target path of adjacent boundaries is determined according to the connection domain.
5. The method according to claim 1, wherein The step of optimizing the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain the target region image comprises: Performing fuzzy equivalence partitioning on the partition set according to fuzzy theory, and constructing a fuzzy relationship matrix; According to a low-rank matrix reduction algorithm, the fuzzy relationship matrix is reduced to generate a low-rank pixel matrix and a spatial relationship matrix; The partition images of the partition set are optimized according to the low-rank pixel matrix, the spatial relationship matrix, and the multiple loop regions to obtain a target region image.
6. The method according to claim 1, wherein The step of comparing the target area image with the standard quality rule paradigm and completing product quality inspection according to the comparison result includes: Comparing the target area image with the standard quality rule paradigm, and calculating a matching score based on the comparison result; When the matching score is less than a preset matching threshold, determining that the product is unqualified; When the matching score is greater than or equal to the preset matching threshold, determining whether the target area image matches the standard quality rule paradigm; When the target area image does not match the standard quality rule paradigm, generating a repair suggestion report according to the matching result; When the target area image matches the standard quality rule pattern, the product is determined to be qualified and the product quality inspection is completed.
7. The method according to claim 1, wherein The step of obtaining image data and depth information of the area to be detected collected by the event camera and the laser radar includes: Acquire the depth image signal of the area to be detected collected by the event camera, and acquire the high-frequency pixel signal of the area to be detected collected by the lidar; generating a multimodal image by fusing the depth image signal and the high-frequency pixel signal; Distortion correction is performed on the multimodal image to obtain image data and depth information.
8. A multimodal fusion product quality detection device, characterized in that: The device comprises: An acquisition module is used to obtain image data and depth information of the area to be detected collected by the event camera and lidar; A partitioning module, configured to partition the image data according to the depth information to obtain a partition set; A segmentation module, configured to segment the boundary of the area to be detected based on a geodesic frame model to obtain multiple loop regions; An optimization module is used to optimize the partition images of the partition set according to fuzzy theory, a low-rank matrix reduction algorithm, and the multiple loop regions to obtain a target region image; The detection module is used to compare the target area image with the standard quality rule paradigm and complete product quality detection based on the comparison result.
9. A multimodal fusion product quality inspection device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multimodal fusion product quality detection method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the multimodal fusion product quality detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Real-time detection method and system for scrap steel impurities
CN121074056A