Vacuum cup weld joint positioning detection system based on vision

By generating weld feature maps through multimodal visual perception and feature enhancement networks, and combining texture contour fusion and spatiotemporal coordinate alignment, parameters are dynamically adjusted to solve the problems of large positioning errors and poor adaptability in the detection of thermos cup welds. This achieves high-precision and rapid weld positioning, improving the adaptability and efficiency of the detection system.

CN121353221AInactive Publication Date: 2026-01-16ZHEJIANG JINGJIANG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511493560.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-01-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for inspecting weld seams in thermos cups suffer from problems such as large positioning errors, poor adaptability, low inspection efficiency, cumbersome parameter adjustments, and the inability to dynamically optimize, making it difficult to meet the high precision and adaptability requirements of modern production lines.

Method used

A multimodal visual perception layer is used to acquire multimodal image information, a weld feature map is generated through a feature enhancement network, and three-dimensional contour data is processed by combining a texture contour fusion layer and a spatiotemporal coordinate alignment layer. Parameters are dynamically adjusted and the detection strategy is optimized in real time through a feedback optimization execution layer to achieve high-precision and highly adaptable weld positioning and detection.

Benefits of technology

It achieves high-precision, rapid, and flexible positioning of welds, reduces the impact of interference factors such as reflection and scratches, improves the adaptability of the detection system and the reliability of the detection results, reduces manual intervention, and improves production efficiency and quality stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353221A_ABST
    Figure CN121353221A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual inspection, and discloses a visual-based vacuum cup weld joint positioning detection system. The system comprises a multi-modal visual perception layer, a texture contour fusion layer, a space-time coordinate alignment layer, a dynamic parameter adjustment layer and a feedback optimization execution layer. The multi-modal visual perception layer is used for acquiring multi-modal image information and generating a surface welding seam feature map; the texture contour fusion layer generates a weld contour feature map through structured light and a phase unwrapping algorithm; the space-time coordinate alignment layer fuses the two feature maps and generates a positioning confidence score; the dynamic parameter adjustment layer converts the score into a detection instruction and issues the detection instruction; and the feedback optimization execution layer monitors the image quality change, generates a detection strategy validity index and dynamically optimizes the detection strategy. The system can improve the accuracy and adaptability of welding seam positioning, cope with the surface interference of the vacuum cup and the change of the production environment, and guarantee the stable and reliable detection effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of visual detection, in particular to a visual-based thermos welding seam positioning detection system. BACKGROUND

[0002] In the production process of thermos, the welding seam quality directly affects the sealing performance and service life of the product, so the welding seam positioning detection is an important link to ensure the production quality of thermos. At present, the positioning detection of thermos welding seam mainly relies on the combination of traditional visual detection method and manual sampling inspection, which has many technical limitations. The traditional visual detection method mainly uses a single camera to collect two-dimensional images, and identifies the welding seam position through a preset gray threshold or edge detection algorithm. However, there are many interference factors on the surface of the thermos, such as reflection, scratches, stains, etc., and it is difficult to accurately distinguish the welding seam from these interference areas in two-dimensional images, resulting in large positioning error. Especially when the welding seam has a small deformation or surface oxidation, the recognition degree of two-dimensional features will decrease significantly, and it is easy to miss or misjudge. At the same time, the imaging range of a single field of view is fixed, and for thermos of different specifications and different curvatures, it is necessary to frequently change the lens or adjust the equipment parameters, which is tedious and has poor adaptability, and it is difficult to meet the rapid switching requirements of flexible production lines. In the aspect of three-dimensional profile detection, some detection systems introduce structure light imaging technology, but the projection angle and density of the structure light stripe in the existing technology are fixed, and cannot be dynamically adjusted according to the actual position of the welding seam. When the welding seam is in different areas of the cup surface, the distortion degree of the structure light stripe is different, and error accumulation is easy to occur in the phase unwrapping process, resulting in insufficient accuracy of three-dimensional profile data. In addition, the fusion of two-dimensional visual information and three-dimensional profile data mainly uses simple coordinate superposition method, which lacks accurate alignment in time and space dimensions, and the feature correlation of the two kinds of data is weak, making it difficult to form effective complementation and affecting the reliability of the positioning result. Manual sampling inspection can compensate for the shortcomings of machine detection to some extent, but it has the problems of low efficiency and high cost. Manual detection depends on the experience and responsibility of the operator, and the judgment standards of different personnel are different, so the consistency of the detection results is poor. In the mass production scene, the coverage rate of manual sampling inspection is limited, and it is difficult to fully control the welding seam quality of each product, which may easily cause unqualified products to flow into the market. The parameter adjustment of the existing detection system is mainly static setting, which cannot be dynamically optimized according to the real-time detection situation. When the thermos model is changed or the production environment changes, the technical personnel need to manually recalibrate the equipment parameters, which not only prolongs the production preparation time, but also may affect the detection effect due to improper parameter setting. At the same time, the detection system lacks effective feedback mechanism, and cannot quantitatively evaluate the effectiveness of the detection strategy, making it difficult to continuously improve the detection precision and stability. With the continuous growth of market demand for thermos cups, the requirements for product quality and production efficiency are increasing. Traditional testing methods can no longer meet the needs of modern production lines, and a weld seam positioning and testing system that can adapt to complex working conditions and achieve high precision and high adaptability is needed. Summary of the Invention

[0003] The purpose of this invention is to provide a vision-based weld seam positioning and detection system for thermos cups, in order to solve the problems mentioned in the background art.

[0004] To achieve the above objectives, the present invention provides a vision-based weld seam positioning and detection system for thermos cups, the system comprising: The multimodal vision perception layer acquires multimodal image information based on a dynamic field-of-view adjustment mechanism, applies a feature enhancement network to the multimodal image information to generate a feature map of the weld seam on the surface of the thermos cup, and performs regional saliency division on the weld seam feature map. The texture contour fusion layer deploys linear array imaging components according to the defined significant regions, emits structured light stripes to cover the surface of the cup, and calculates three-dimensional contour data through a phase unwrapping algorithm to generate a contour feature map of the weld seam of the thermos cup. A spatiotemporal coordinate alignment layer is used to establish a spatial mapping between vision and contour. A time synchronization protocol is used to calibrate the clock. The surface weld feature map and weld contour feature map are fused through a feature association network to generate a positional reliability score. The dynamic parameter adjustment layer converts the location reliability score into executable detection instructions based on a multi-objective decision-making algorithm, and sends the executable detection instructions to the imaging controller in real time through a communication protocol. The feedback optimization execution layer monitors the changes in image quality after the imaging controller executes the detection command in real time, calculates the deviation between the image quality change and the preset image quality change threshold, generates a detection strategy effectiveness index, and dynamically optimizes the visual detection strategy until the detection strategy effectiveness index reaches its optimal value.

[0005] Preferably, the method for acquiring multimodal image information based on a dynamic field-of-view adjustment mechanism includes: The thermos cup weld positioning and detection system deploys three vision devices: a linear array camera, an infrared thermal imager, and a structured light projector, to collect multimodal image information of the cup surface. The multimodal image information includes texture distribution data, surface temperature distribution data, and surface three-dimensional contour data of the cup surface. The non-uniform field of view mechanism of the simulation variable focus lens dynamically adjusts the device resolution and the sampling frequency; based on the design model of the vacuum cup, the cup body ring weld, the longitudinal weld and the interface weld area are recorded as the high-resolution focus area; the remaining area of the cup body surface is recorded as the low-resolution peripheral area; the resolution function is dynamically defined, and the cup body surface is divided into the high-resolution focus area and the low-resolution peripheral area; the high-resolution focus area adopts continuous scanning sampling, and the low-resolution peripheral area adopts interval scanning sampling; through the dynamic adjustment of the control instruction device parameters, the final multi-modal image information is obtained.

[0006] Preferably, the method of applying the feature enhancement network to the multi-modal image information to generate the vacuum cup surface weld feature map comprises: The multi-modal image information is subjected to mean normalization processing, and the normalized multi-modal image information is stacked into a four-dimensional data block in the channel dimension, the dimensions of the four-dimensional data block including height, width, channel number and time sequence frame; a feature enhancement network structure is constructed, which includes a forward feature extraction path, a backward detail recovery path and a cross-layer connection; The multi-modal image information is input into the feature enhancement network structure, in the forward feature extraction path, a convolutional neural network is used to extract features from the input multi-modal data, generate feature maps of different scales, and extract finer detail information layer by layer, the feature maps of different scales including bottom edge feature maps, middle texture feature maps and high-level structure feature maps; in the backward detail recovery path, the high-level structure feature map is gradually recovered in spatial resolution through upsampling; the cross-layer connection is established between the forward feature extraction path and the backward detail recovery path, and the feature maps of the same scale are fused; A 3×3 convolution is applied on each output layer of the feature enhancement network structure to adjust the feature dimension, the feature maps of different scales output by the feature enhancement network structure are fused to generate the vacuum cup surface weld feature map.

[0007] Preferably, the method of dividing the weld feature map into regions of saliency comprises: The weld saliency score of each region in the weld feature map is calculated, the texture distribution data, the surface temperature distribution data and the surface three-dimensional contour data of the cup body surface included in the multi-modal image information are weighted and summed to obtain the weld saliency score; The preset weld seam saliency score first threshold and the weld seam saliency score second threshold are compared with the weld seam saliency score, if the weld seam saliency score is less than the weld seam saliency score first threshold, the region corresponding to the score is marked as gray; if the weld seam saliency score is greater than the weld seam saliency score first threshold and less than the weld seam saliency score second threshold, the region corresponding to the score is marked as yellow; if the weld seam saliency score is greater than the weld seam saliency score second threshold, the region corresponding to the score is marked as red. The weld seam feature map is divided into regions of saliency using different colors, different colors representing different saliency regions, the gray region being defined as a non-weld seam region, the yellow region being defined as a suspected weld seam region, and the red region being defined as a salient weld seam region.

[0008] Preferably, the method for generating the thermos cup weld seam contour feature map comprises: According to the divided salient regions, different density structure light projection units are deployed on the surface of the cup; the structure light projection unit is a linear array projection assembly, different projection strategies are adopted for different salient regions, the structure light projection unit emits sinusoidal stripes to the surface of the cup for initial global scanning, the projection angle is dynamically adjusted according to the salient regions of the weld seam feature map, wherein the projection angle interval of the salient weld seam region is smaller than that of the suspected weld seam region, and the projection angle interval of the suspected weld seam region is smaller than that of the non-weld seam region; an angle adjustment algorithm is applied to dynamically adjust the projection angle, a surface reflectivity correction model is introduced, the parameters in the reflectivity model are continuously corrected through an iterative phase solving algorithm, and the process is stopped when the maximum number of iterations is reached; a contour calculation matrix is established based on the principle of triangulation, the surface contour is solved by the least square method, and finally a thermos cup weld seam contour feature map with salient markers is generated.

[0009] Preferably, the method for establishing the spatial mapping of vision and contour comprises: A unified global coordinate system is defined with the center of the mouth of the thermos cup as the origin, the X-axis and the Y-axis being parallel to the surface of the cup, and the Z-axis being perpendicular to the plane of the cup mouth; a calibration board is used to calibrate the linear array camera to obtain the intrinsic and extrinsic parameters of the linear array camera; the pixel coordinates of the weld seam feature map are converted to camera coordinates through the intrinsic parameters of the linear array camera, and then the camera coordinates are converted to global coordinates through the extrinsic parameters of the linear array camera, thereby obtaining the weld seam feature map in the global coordinate system; the origin of the global coordinate system is used as a reference point to calibrate the structure light projection unit to obtain the extrinsic parameters of the structure light projection unit; The weld contour feature map is converted into a global coordinate system by an external parameter of the structured light projection unit, and then a weld contour feature map in the global coordinate system is obtained; in the global coordinate system, the weld feature map and the weld contour feature map are spatially registered to establish a spatial mapping of the vision and the contour; a time source connected to the linear array camera is set as a master clock, a time source connected to the structured light projection unit is set as a slave clock, time calibration is performed through a time synchronization protocol, and a time mapping of the vision and the contour is established.

[0010] Preferably, the method for fusing the surface weld feature map and the weld contour feature map through the feature association network comprises: According to the surface weld feature map and the weld contour feature map in the global coordinate system, an association graph is constructed, a multi-layer graph convolution structure based on a multi-head attention mechanism is adopted to build a feature association network, and the surface weld feature map and the weld contour feature map in the global coordinate system are fused; the feature association network comprises an input layer, a feature mapping layer, a graph attention layer, a cross-modal interaction layer and an output layer; the association graph is input into the input layer of the feature association network, and a positional confidence score is generated through the output layer.

[0011] Preferably, the method for constructing the association graph comprises: Each salient region in the surface weld feature map in the global coordinate system is taken as a vision node, and a feature vector of each salient region is extracted as a vision node feature; each contour edge region of the weld contour feature map in the global coordinate system is taken as a contour node, and a three-dimensional feature vector of each contour edge region is extracted as a contour node feature; all vision nodes and contour nodes are collected to obtain a node set; All vision nodes are traversed, and the Euclidean distance between any two vision nodes in the global coordinate system is calculated; a preset vision distance threshold is set, if the Euclidean distance between any two vision nodes in the global coordinate system is less than the preset vision distance threshold, a bidirectional edge is added between the corresponding two vision nodes; if the Euclidean distance between any two vision nodes in the global coordinate system is greater than or equal to the preset vision distance threshold, no bidirectional edge is added. All contour nodes are traversed, and the Euclidean distance between any two contour nodes in the global coordinate system is calculated; a preset contour distance threshold is set, if the Euclidean distance between any two contour nodes in the global coordinate system is less than the preset contour distance threshold, a bidirectional edge is added between the corresponding two contour nodes; if the Euclidean distance between any two contour nodes in the global coordinate system is greater than or equal to the preset contour distance threshold, no bidirectional edge is added; the nearest neighbor search is performed to find the nearest contour node for each vision node, and a bidirectional edge between the vision node and the corresponding nearest contour node is added; all bidirectional edges are collected to obtain an edge set; and the association graph is constructed based on the obtained node set and edge set.

[0012] Preferably, the method for converting the positioning confidence score into executable detection instructions based on the multi-objective decision algorithm comprises: The surface of the cup is divided into rectangular grid units, and each unit records the current profile deviation value and the positioning confidence score; meanwhile, all adjustable detection parameters are listed, including the adjustment range of each parameter, the energy consumption cost and the influence degree on the positioning accuracy; Three decision objectives are established, including the first positioning objective, the second positioning objective and the energy efficiency objective; two types of constraint conditions are set, including the hard constraint and the elastic constraint; A multi-objective decision algorithm is used to randomly generate m sets of detection parameter optimization schemes, the completion of the three decision objectives is evaluated for each scheme, the completion score of the three decision objectives is obtained, the completion scores of the three decision objectives are added together to obtain a comprehensive score, and one scheme with the highest comprehensive score is selected from the m sets of schemes as the final optimization scheme; the selected final optimization scheme is converted into actual executable detection instructions; the executable detection instructions include the first adjustment instruction, the second adjustment instruction and the monitoring and early warning instruction; The method for obtaining the completion score of the three decision objectives comprises: The positioning accuracy improvement percentage of the cup weld is observed as the completion score of the first positioning objective; the dispersion of the positioning confidence scores in different regions is calculated as the completion score of the second positioning objective; and the total energy consumption caused by the adjustment of all detection parameters is calculated as the completion score of the energy efficiency objective.

[0013] Preferably, the method for generating the detection strategy effectiveness index comprises: The detection instruction parameters executed by the imaging controller are collected, the image quality change is synchronously monitored, the deviation of the image quality change from the preset image quality change threshold is calculated, and the standardized detection strategy effectiveness index d with a value range of 0 to 1 is generated; the preset standardized detection strategy effectiveness index first threshold d1 and the standardized detection strategy effectiveness index second threshold d2 are set; if d is greater than d2, it is judged that the detection strategy is effective; if d1 is less than d and d is less than d2, it is judged that the detection strategy still needs to be observed; and if d is less than d1, the detection strategy is invalid, the visual detection strategy is optimized, and early warning information is immediately generated Compared with the prior art, the present application has the following advantages: Through the dynamic field of view adjustment mechanism of the multi-modal visual perception layer, various interference factors on the surface of the vacuum cup can be flexibly coped with. The dynamic field of view adjustment can adjust the imaging range and the focal length in real time according to the actual position of the weld and the change of the surrounding environment, process the multi-modal image information by combining the feature enhancement network, effectively highlight the weld features, reduce the influence of interference such as reflection and scratches on the detection result, and make the generated weld feature map more clear and accurate. The texture profile fusion layer deploys a linear array imaging component according to the salient region, and emits a structured light stripe in a targeted manner. The three-dimensional profile data is calculated by combining a phase unwrapping algorithm. The generated weld profile feature map can accurately present the three-dimensional morphology of the weld. This method avoids the profile data loss or distortion caused by unreasonable stripe projection in traditional structured light imaging, and fully displays the three-dimensional features of the weld, which effectively complements the surface weld feature map. The space-time coordinate alignment layer establishes the spatial mapping of vision and profile, and calibrates the clock through a time synchronization protocol, ensuring the consistency of the surface weld feature map and the weld profile feature map in the space-time dimension. The feature correlation network integrates the advantages of two-dimensional surface information and three-dimensional profile information, and the generated positioning confidence score can more truly reflect the actual position of the weld, reducing the positioning deviation caused by single feature information. The dynamic parameter adjustment layer converts the positioning confidence score into executable detection instructions based on a multi-objective decision algorithm, and real-time issues the detection instructions to the imaging controller, realizing the dynamic adjustment of the detection parameters. This dynamic adjustment mechanism can quickly respond to various changes in the detection process, and can make the imaging equipment always in the best working state without manual intervention, improving the adaptability of the system to different specifications of the thermos and complex production environments. The feedback optimization execution layer continuously optimizes the detection strategy by monitoring the image quality changes in real time, calculating the deviation and generating a detection strategy effectiveness index. This process enables the system to automatically discover and correct problems in the detection process, gradually improve the detection performance, adapt to various complex situations that may occur in the production process, and reduce quality problems caused by improper detection strategies. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 The working principle diagram of the vision-based thermos weld positioning and detection system described in the present application; Figure 2 The flowchart for acquiring multi-modal image information by the dynamic field of view adjustment mechanism; Figure 3 The flowchart for generating a weld feature map by the feature enhancement network; Figure 4 The flowchart for generating a thermos weld profile feature map; Figure 5 The flowchart for establishing the spatial mapping of vision and profile. DETAILED DESCRIPTION

[0015] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the protection scope of the present application.

[0016] Please refer to Figures 1-5 The present application provides a visual-based thermos welding seam positioning detection system, which comprises: The multi-modal visual perception layer obtains multi-modal image information based on a dynamic field of view adjustment mechanism. First, a plurality of visual devices are deployed to collect multi-modal data on the surface of the thermos. Then, a feature enhancement network is applied to process the data to generate a thermos surface welding seam feature map. Subsequently, the feature map is divided into regions of saliency to determine the saliency of the welding seams in different regions.

[0017] The texture contour fusion layer deploys a line array imaging component based on the salient regions divided by the multi-modal visual perception layer, emits a structured light stripe to cover the cup body, calculates three-dimensional contour data through a phase unwrapping algorithm, and then generates a thermos welding seam contour feature map to accurately depict the welding seam contour.

[0018] The spatio-temporal coordinate alignment layer establishes a spatial mapping relationship between the vision and the contour, calibrates the related clocks using a time synchronization protocol, and fuses the surface welding seam feature map and the welding seam contour feature map through a feature association network to finally generate a positioning confidence score, which provides a basis for the generation of subsequent detection instructions.

[0019] The dynamic parameter adjustment layer converts the positioning confidence score generated by the spatio-temporal coordinate alignment layer into executable detection instructions based on a multi-objective decision algorithm, and transmits these instructions to the imaging controller in real time through a preset communication protocol to dynamically adjust the detection process.

[0020] The feedback optimization execution layer monitors the image quality changes after the imaging controller executes the detection instructions in real time, calculates the deviation of the image quality changes from the preset threshold, generates a detection strategy effectiveness index, dynamically optimizes the visual detection strategy according to the index, and ensures the efficient and stable operation of the entire detection system until the detection strategy effectiveness index reaches an optimal state.

[0021] In the process of obtaining multi-modal image information based on the dynamic field of view adjustment mechanism, the system deploys three visual devices: a linear array camera, an infrared thermal imager, and a structured light projector. The linear array camera captures texture details at different positions by continuously scanning the surface of the cup, forming texture distribution data that reflects the ups and downs and material differences of the cup surface. The infrared thermal imager senses the infrared radiation of the cup surface to generate surface temperature distribution data, and the temperature differences in different areas may be related to the presence of the weld. The structured light projector projects a specific pattern onto the surface of the cup, and combined with the feedback from the imaging device, it initially obtains the surface three-dimensional profile data, providing a basis for subsequent detailed processing. The data collected by the three devices together constitute multi-modal image information, covering various physical characteristics of the cup surface.

[0022] The non-uniform field of view mechanism of the simulated variable focus lens is simulated, and according to the design model of the thermos, the areas that need to be focused on are determined. The ring weld, longitudinal weld, and interface weld areas on the cup are set as high-resolution focus areas, which are the core of detection and require more detailed image information. Other parts of the cup surface are set as low-resolution peripheral areas, which can appropriately reduce the data acquisition density while ensuring that the basic information is not lost. By dynamically defining the resolution function, the boundaries of the high-resolution focus area and the low-resolution peripheral area are accurately divided, ensuring that the division of the two areas meets the actual detection requirements. For the high-resolution focus area, continuous scanning sampling is used, the linear array camera scans at a higher frequency, the infrared thermal imager increases the density of temperature measurement points, and the structured light projector enhances the density of the projected pattern, thereby obtaining continuous and detailed data. The low-resolution peripheral area uses interval scanning sampling, appropriately reduces the scanning frequency and data acquisition density, while reducing the amount of data and maintaining a grasp of the overall surface condition. By sending control instructions, the parameters of the three visual devices are adjusted in real time, such as the scanning speed of the linear array camera, the sampling interval of the infrared thermal imager, and the pattern density of the structured light projector, ultimately obtaining multi-modal image information that meets the detection requirements.

[0023] When applying the feature enhancement network to generate the thermos surface weld feature map from the multi-modal image information, first, mean normalization is performed. The mean and standard deviation of each modal image information are calculated, the data is converted to the same numerical range, and the influence of the differences in dimension and numerical range between different modal data is eliminated, making the subsequent network processing more stable. The normalized multi-modal image information is stacked according to the channel dimension to form a four-dimensional data block, where the four dimensions correspond to the height, width, channel number, and time sequence frame of the image, the channel number corresponds to different modal, and the time sequence frame records the acquisition data at different time points, facilitating the network to capture dynamic change information.

[0024] The constructed feature enhancement network structure includes a forward feature extraction path, a backward detail recovery path, and a cross-layer connection. The forward feature extraction path is composed of multiple convolutional layers, each of which processes the input data through a convolution kernel of different size. The initial convolutional layer extracts a bottom edge feature map, capturing the basic features such as lines and edges on the surface of the cup; the middle convolutional layer further processes to generate a middle texture feature map, reflecting the texture patterns and local details on the surface; and the deep convolutional layer extracts a high-level structure feature map, integrating global information to form an abstract representation of the overall structure of the cup. As the number of network layers increases, the spatial resolution of the feature map gradually decreases, but the semantic information contained becomes more rich.

[0025] The backward detail recovery path is symmetrical to the forward feature extraction path and gradually recovers the spatial resolution of the feature map through upsampling operations. During upsampling, methods such as interpolation are used to supplement pixel information, allowing the high-level structure feature map to be restored to a size similar to the bottom layer feature map, facilitating feature fusion. Cross-layer connections are established between the forward feature extraction path and the backward detail recovery path to fuse feature maps of the same scale. For example, the bottom edge feature map of a certain layer in the forward path is superimposed with the feature map of the corresponding scale in the backward path, so that the fused feature map contains both high-level structure information and bottom-level detail features, enhancing the expression ability of the features.

[0026] At each output layer of the feature enhancement network structure, 3x3 convolution is applied for feature dimension adjustment. The 3x3 convolution kernel can extract local features while keeping the spatial size of the feature map relatively stable. By adjusting the number of convolution kernels, the feature maps of different output layers are converted to the same feature dimension, preparing for subsequent fusion operations. Different scale feature maps after dimension adjustment are fused by pixel-by-pixel addition or splicing, integrating the information of features at each layer, and finally generating a surface weld feature map of the thermos cup that can clearly reflect the weld features.

[0027] When dividing the weld feature map into regions of saliency, the weld saliency score of each region is calculated. According to the importance of texture distribution data, surface temperature distribution data, and surface three-dimensional contour data in weld detection, different weights are assigned, and the comprehensive score of each region is obtained by weighted summation. In the texture distribution data, the weld region usually has a different texture pattern from the surrounding area; in the surface temperature distribution data, the weld region may have temperature differences due to the influence of the welding process; and in the surface three-dimensional contour data, the weld may appear as a protrusion or a depression. These information jointly participate in the score calculation.

[0028] The first threshold and the second threshold of the preset weld saliency score are determined according to a large amount of actual detection data and typical manifestations of weld features. The weld saliency score of each region is compared with the two thresholds: the region with a score less than the first threshold, which does not conform to the typical features of the weld in terms of texture, temperature and contour features, is marked in gray and defined as a non-weld region; the region with a score greater than the first threshold and less than the second threshold, which has some features close to the weld features but not obvious enough, is marked in yellow and defined as a suspected weld region; the region with a score greater than the second threshold, which has obvious weld features, is marked in red and defined as a salient weld region. Through the marking of different colors, the region saliency of the weld feature map is clearly divided, providing clear region guidance for subsequent weld positioning and detection.

[0029] In the process of generating the mug weld contour feature map, different density of structured light projection units are deployed on the surface of the cup body according to the salient regions divided by the multi-modal visual perception layer. Different projection strategies are adopted for different salient regions. The structured light projection units emit sinusoidal fringes to the surface of the cup body. First, an initial global scan is performed to obtain the overall contour information of the cup body surface. Then, the projection angle is dynamically adjusted according to the salient regions of the weld feature map. The projection angle interval of the salient weld region is set to be smaller than that of the suspected weld region, and the projection angle interval of the suspected weld region is set to be smaller than that of the non-weld region, so as to realize differential scanning coverage of different regions.

[0030] The angle adjustment algorithm is applied to dynamically adjust the projection angle. In this process, a surface reflectivity correction model is introduced. This model takes into account the differences in the reflection characteristics of the structured light fringes due to the material of the cup body surface. By analyzing the reflected light intensity and phase changes in different regions, the parameters in the model are corrected. An iterative phase solving algorithm is used to continuously optimize the parameters of the reflectivity model. Each iteration adjusts based on the deviation between the current measurement data and the model prediction value. When the maximum number of iterations is reached, the parameter correction process is stopped.

[0031] A contour calculation matrix is established based on the principle of triangulation. This matrix includes the relative position relationship between the structured light projection unit and the imaging device, the phase information of the projected fringes, and the reflection characteristics of the cup body surface. By using the least squares method to solve this matrix, the three-dimensional coordinates of each point on the cup body surface are calculated, and a complete surface contour is constructed. When generating the mug weld contour feature map, the salient weld region, the suspected weld region and the non-weld region are presented in different marking ways, so that the contour feature map can clearly reflect the contour differences of different regions, and correspond to the region division of the previously generated surface weld feature map.

[0032] In specific operation, the deployment density of the structured light projection unit is adjusted according to the division of the significant area, the deployment density of the projection unit in the significant weld area is higher than that in the suspected weld area, and the deployment density of the projection unit in the suspected weld area is higher than that in the non-weld area, so as to ensure that the key area can obtain more dense structured light stripe coverage and improve the accuracy of the profile data. The adjustment of the projection angle is realized by driving the projection unit by the stepping motor, and the step length of the angle adjustment is determined according to the significance of the area, and a smaller step length is used in the significant weld area to realize more fine angle adjustment.

[0033] In the phase solving process, the collected stripe image is preprocessed to remove noise and interference signals, and clear phase information is extracted. The wrapped phase is converted into absolute phase by the phase unwrapping algorithm, and the three-dimensional coordinates are calculated by combining the triangulation principle. In the calculation process, the curved surface characteristics of the cup body are fully considered, and the profile data is fitted to the curved surface, so that the generated profile feature map can more accurately reflect the actual shape of the thermal cup.

[0034] The finally generated thermal cup weld profile feature map contains three-dimensional profile information of the cup body surface, and different significant areas are clearly distinguished by marking, providing accurate profile data support for subsequent space-time coordinate alignment and feature fusion.

[0035] In the establishment of the spatial mapping of vision and profile, a unified global coordinate system is first defined, taking the center of the mouth of the thermal cup as the origin, the X and Y axes being parallel to the cup surface, and the Z axis being perpendicular to the cup mouth plane. The spatial position of any point on the cup surface is determined by the three coordinate axes. The line array camera is calibrated using a calibration board to obtain the intrinsic and extrinsic parameters of the line array camera. The intrinsic parameters include the focal length, pixel size and distortion coefficient of the camera, and the extrinsic parameters include the translation vector and rotation matrix of the camera in the global coordinate system.

[0036] The pixel coordinates of the weld feature map are converted to camera coordinates through the intrinsic parameters of the line array camera. This conversion process needs to consider the imaging model of the camera to map the two-dimensional pixel coordinates to the three-dimensional camera coordinate system. Then the camera coordinates are converted to the global coordinates through the extrinsic parameters of the line array camera. The rotation matrix in the extrinsic parameters is used to describe the directional relationship between the camera coordinate system and the global coordinate system, and the translation vector is used to describe the positional relationship between the origins of the two coordinate systems. After the two conversion steps, the weld feature map in the global coordinate system can be obtained.

[0037] The origin of the global coordinate system is taken as the reference point to calibrate the structured light projection unit and obtain the extrinsic parameters of the structured light projection unit, including the position and attitude parameters of the projection unit in the global coordinate system. The weld profile feature map is converted to the global coordinate system through the extrinsic parameters of the structured light projection unit. The conversion method is similar to the conversion of the line array camera coordinates, and the projection model and spatial attitude of the projection unit need to be considered. Finally, the weld profile feature map in the global coordinate system is obtained.

[0038] In the global coordinate system, the weld feature map and the weld contour feature map are spatially coordinated, the corresponding feature points in the two feature maps are found, the coordinate transformation parameters are calculated, and the two feature maps are accurately aligned in the same coordinate system, thereby establishing the spatial mapping of vision and contour.

[0039] In terms of time synchronization, the time source connected to the line array camera is set as the master clock, and the time source connected to the structured light projection unit is set as the slave clock. The two clocks are calibrated using a time synchronization protocol. The master clock sends a time synchronization signal, and the slave clock adjusts its clock phase and frequency after receiving the signal, so that the time deviation of the two clocks is controlled within a preset range. In this way, the image acquisition of the line array camera and the projection of the structured light projection unit are kept in time synchronization, and the time mapping of vision and contour is established.

[0040] In the coordinate conversion process, the conversion from pixel coordinates to camera coordinates can be represented as:

[0041]

[0042] where (x, y) are pixel coordinates, (u, v, f) are camera coordinates, f is the camera focal length, and (cx, cy) are pixel coordinates of the image center. , , , ,

[0043] Through the above steps, the accurate alignment of vision and contour in space and time is realized, and a unified space-time reference is provided for subsequent feature fusion and position confidence score generation. In actual operation, the calibration process needs to be performed multiple times to reduce the influence of calibration error on coordinate conversion accuracy. The calibration frequency of time synchronization is determined according to the running speed and detection accuracy requirements of the system to ensure that the time deviation is always within the allowable range during the entire detection process. After spatial coordinate registration, the registration result needs to be verified to check whether the position deviation of the corresponding areas in the two feature maps is within the acceptable range. If the deviation is too large, the registration operation needs to be performed again.

[0044] ​​​Example 4: When fusing surface weld feature maps and weld contour feature maps using a feature association network, an association map is first constructed based on the surface weld feature maps and weld contour feature maps in the global coordinate system. Each salient region in the surface weld feature map is treated as an independent visual node. These regions include salient weld regions, suspected weld regions, and non-weld regions. For each visual node, its feature vector is extracted. The feature vector contains details of the texture distribution of the region, the surface temperature change trend, and its position information in the global coordinate system. This information is extracted and integrated from the previously processed multimodal image data. Simultaneously, each contour edge region in the weld contour feature map is treated as a contour node. Each contour edge region corresponds to a contour line or a contour surface on the cup surface. A three-dimensional feature vector is extracted for each contour node. This vector contains the spatial coordinate parameters of the contour, the curvature change of the contour line, and the angular relationship between adjacent contours. This data comes from the three-dimensional contour calculation results.

[0045] Collect all visual nodes and contour nodes to form a complete node set. Traverse all visual nodes and calculate the Euclidean distance between any two visual nodes in the global coordinate system, which is the straight-line distance obtained by taking the square root of the sum of the squares of the coordinate differences. A preset visual distance threshold is set based on the size of the thermos cup and the possible distribution range of the weld seam. When the Euclidean distance between two visual nodes is less than this threshold, it indicates that the two regions are spatially close and may be related; therefore, a bidirectional edge is added between the two visual nodes. If the distance is greater than or equal to the threshold, the two are considered to be weakly related, and no edge is added.

[0046] Contour nodes are processed in the same way. All contour nodes are traversed, and the Euclidean distance between any two contour nodes in the global coordinate system is calculated. A preset contour distance threshold is used, which is set with reference to the density of contour lines and the typical length of weld contours. When the distance between two contour nodes is less than the contour distance threshold, a bidirectional edge is added between them; otherwise, no edge is added. Furthermore, a nearest neighbor search algorithm is used to find the nearest contour node for each visual node among all contour nodes, and a bidirectional edge is added between these two nodes to establish a direct association between visual features and contour features. All bidirectional edges generated through the above steps are collected to form an edge set. Based on the node set and the edge set, an association graph reflecting the relationship between visual nodes and contour nodes is constructed.

[0047] The constructed feature association network includes an input layer, a feature mapping layer, a graph attention layer, a cross-modal interaction layer, and an output layer. The association graph is input into the input layer, which converts the node feature vectors and edge association information into a format that meets the network processing requirements. The feature mapping layer processes the node features through multiple fully connected layers, converting the feature vectors to a higher dimensional space while preserving key feature information, making the features of different types of nodes comparable in the same dimensional space.

[0048] The graph attention layer uses a multi-head attention mechanism to calculate the weights of each node and its connected edges. For each node, different attention weights are assigned to each connected node based on their feature similarity and spatial distance. The higher the weight, the greater the influence of the connected node on the current node. In this way, the information of important nodes is highlighted, and the interference of irrelevant nodes is weakened, making the network pay more attention to key features related to weld positioning.

[0049] The cross-modal interaction layer is responsible for deep fusion of visual node features and contour node features. This layer uses a cross-attention mechanism to allow visual nodes and contour nodes to interact and transfer feature information. For example, the temperature feature of a visual node can affect the weight distribution of related contour nodes, and the spatial feature of a contour node can also adjust the feature expression of a visual node. This interactive fusion process is implemented through multiple layers of neural networks, each layer performing nonlinear transformation on the interacted features to enhance their expression ability.

[0050] The output layer integrates all the fused node features, aggregates them into a comprehensive feature vector through global pooling, and then processes them through fully connected layers and activation functions to generate a positioning confidence score. The numerical value of this score reflects the reliability of the current mug weld positioning result, and its calculation takes into account the consistency of visual features, the accuracy of contour features, and the closeness of their association.

[0051] During the entire network processing process, each layer optimizes its parameters through backpropagation algorithms, allowing the network to continuously learn how to better associate visual features and contour features. The final positioning confidence score accurately reflects the actual situation of the weld positioning, providing a basis for subsequent dynamic parameter adjustment. The training data for the network comes from a large number of mug sample images and their corresponding actual weld position information, and through multiple rounds of iterative training, the network achieves stable processing results.

[0052] In the embodiment 5, when the position confidence score is converted into executable detection instructions based on the multi-objective decision algorithm, the surface of the cup is first divided into multiple rectangular grid units according to the preset size, and the size of each grid unit is determined according to the diameter and height of the cup to ensure that the entire surface of the cup is covered and not overlapped. The current profile deviation value and the position confidence score of each grid unit are recorded separately, the profile deviation value reflects the deviation of the actual profile in the unit from the design model, and the position confidence score comes from the result of the previous fusion processing. At the same time, all adjustable detection parameters are listed, including the scanning frequency of the line array camera, the sampling interval of the infrared thermal imager, the stripe density of the structured light projector, the exposure time of the imaging controller, etc., and the adjustment range of each parameter is specified, such as the minimum and maximum values of the scanning frequency; the energy consumption cost is calculated, that is, the energy consumption corresponding to the adjustment of the parameter to different values; and the historical influence of the parameter on the positioning accuracy is recorded, that is, the change of the positioning accuracy after adjusting the parameter in the past.

[0053] Three decision objectives are established, the first positioning objective focuses on the improvement of the positioning accuracy of the cup weld, the second positioning objective focuses on the consistency of the position confidence score in different regions, and the energy efficiency objective focuses on the energy consumption control in the detection process. Two types of constraint conditions are set, the hard constraint includes that the parameter adjustment cannot exceed the physical limit of the equipment, such as the scanning frequency of the line array camera cannot exceed the maximum allowed value; the elastic constraint includes that the positioning accuracy needs to reach a basic range, which can be adjusted according to the actual detection requirements.

[0054] When using the multi-objective decision algorithm, m groups of detection parameter optimization schemes are randomly generated, and the number of m is determined according to the complexity of the parameters and the computing power. For each scheme, the completion of the three decision objectives is evaluated. When evaluating the first positioning objective, the positioning accuracy improvement percentage of the cup weld is observed, that is, the accuracy improvement percentage of the positioning result under the scheme compared with the previous result; when evaluating the second positioning objective, the dispersion of the position confidence score in different regions is calculated, and the smaller the dispersion, the more consistent the confidence scores of the regions; when evaluating the energy efficiency objective, the total energy consumption caused by the adjustment of all detection parameters is calculated, that is, the sum of the energy consumption of each parameter. The completion of the three objectives is converted into a score of 0-100, and the three scores are added to obtain a comprehensive score, and the scheme with the highest comprehensive score is selected from the m schemes as the final optimization scheme, which is converted into executable detection instructions. The executable detection instructions include the first adjustment instruction for adjusting the sampling parameters of the imaging device, the second adjustment instruction for adjusting the related parameters of the structured light projection, and the monitoring and warning instruction for triggering the corresponding monitoring mechanism when the detection is abnormal.

[0055] When generating the detection strategy effectiveness index, the parameters of the detection instructions executed by the imaging controller are collected in real time, including the actual adjusted scanning frequency, sampling interval, etc., and at the same time, the image quality change is monitored through the image quality evaluation module, including the changes of the clarity, contrast, noise level and other indicators. The deviation of these image quality changes from the preset image quality change threshold is calculated. The preset threshold is determined according to the quality indicators of the historical high-quality detection images. According to the deviation size, a standardized detection strategy effectiveness index d is generated, which has a value range of 0 to 1, and the larger the value is, the higher the effectiveness of the detection strategy is. Two thresholds d1 and d2 are preset, d1 is less than d2. When d is greater than d2, it is judged that the current detection strategy is effective, and the execution continues according to the strategy; when d1 is less than d and less than d2, it is judged that the effect of the detection strategy is in an intermediate state, and the image quality change still needs to be continuously observed; when d is less than d1, it is judged that the detection strategy is invalid, at this time the optimization process of the visual detection strategy is started, the related parameters are adjusted to generate a new scheme, and immediately a warning information is generated, which is sent to the control center through the communication module of the system, reminding the operator to pay attention to the abnormal situation.

[0056] In the whole process, the conversion of the detection instructions needs to comply with the communication protocol of the equipment, so as to ensure that the imaging controller can accurately parse and execute the instructions. The image quality monitoring and parameter collection are kept synchronous, so as to avoid inaccurate deviation calculation caused by time difference. The optimization process of the detection strategy takes targeted measures according to different invalid reasons, such as adjusting the energy efficiency related parameters in priority when the invalidity is caused by too high energy consumption, and focusing on optimizing the parameters related to the positioning target when the invalidity is caused by insufficient positioning accuracy.

[0057] It should be noted that, in the present text, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment.

[0058] Although the embodiments of the present application have been shown and described, it can be understood by those of ordinary skill in the art that various changes, modifications, replacements and variations can be made to these embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A visual-based thermos-weld joint positioning detection system, characterized in that, The method comprises the following steps: A multi-modal visual perception layer obtains multi-modal image information based on a dynamic field of view adjustment mechanism, applies a feature enhancement network to the multi-modal image information, generates a thermos cup surface weld feature map, and performs regional saliency division on the weld feature map; A texture contour fusion layer deploys a line array imaging component according to the divided salient regions, emits a structured light stripe to cover the surface of the cup, calculates three-dimensional contour data through a phase unwrapping algorithm, and generates a thermos cup weld contour feature map; A space-time coordinate alignment layer establishes a spatial mapping of vision and contour, adopts a time synchronization protocol to calibrate the clock, fuses the surface weld feature map and the weld contour feature map through a feature correlation network, and generates a positioning confidence score; A dynamic parameter adjustment layer converts the positioning confidence score into an executable detection instruction based on a multi-objective decision algorithm, and delivers the executable detection instruction to an imaging controller in real time through a communication protocol; A feedback optimization execution layer monitors the image quality change after the imaging controller executes the detection instruction in real time, calculates the deviation of the image quality change from a preset image quality change threshold, generates a detection strategy effectiveness index, dynamically optimizes the visual detection strategy, and stops until the detection strategy effectiveness index reaches the optimal value.

2. The visual-based positioning detection system for the welding seam of a vacuum cup according to claim 1, characterized in that, The method for obtaining multi-modal image information based on a dynamic field of view adjustment mechanism comprises the following steps: A line array camera, an infrared thermal imager, and a structured light projector are deployed on a thermos cup weld positioning detection system, and three kinds of visual equipment are used to collect multi-modal image information of the surface of the cup; the multi-modal image information includes texture distribution data, surface temperature distribution data, and surface three-dimensional contour data of the surface of the cup; A non-uniform field of view mechanism of a variable focus lens is simulated to dynamically adjust the device resolution and sampling frequency; based on the design model of the thermos cup, the cup girth weld, longitudinal weld, and interface weld regions are recorded as high-resolution focus areas; the remaining regions on the surface of the cup are recorded as low-resolution peripheral areas; a resolution function is dynamically defined to divide the surface of the cup into high-resolution focus areas and low-resolution peripheral areas; the high-resolution focus areas are continuously scanned and sampled, the low-resolution peripheral areas are intermittently scanned and sampled, and the device parameters are dynamically adjusted through control instructions to obtain the final multi-modal image information.

3. The visual-based positioning detection system for the welding seam of a vacuum cup according to claim 2, characterized in that, The method for applying a feature enhancement network to the multi-modal image information to generate a thermos cup surface weld feature map comprises the following steps: The multi-modal image information is subjected to mean normalization, and the normalized multi-modal image information is stacked into a four-dimensional data block according to the channel dimension; the four-dimensional data block includes height, width, channel number, and time sequence frame; a feature enhancement network structure is constructed, which includes a forward feature extraction path, a backward detail recovery path, and a cross-layer connection; The multi-modal image information is input into the feature enhancement network structure, in the forward feature extraction path, the input multi-modal data is extracted by using a convolutional neural network to generate feature maps of different scales, and more fine detailed information is extracted layer by layer, and the feature maps of different scales include bottom edge feature maps, middle texture feature maps and high structure feature maps; in the backward detail recovery path, the high structure feature map is recovered step by step by upsampling the spatial resolution; the cross-layer connection is established between the forward feature extraction path and the backward detail recovery path, and the feature maps of the same scale are fused; 3×3 convolution is applied to each output layer of the feature enhancement network structure to adjust the feature dimension, and the feature maps of different scales output by the feature enhancement network structure are fused to generate the thermos cup surface weld feature map.

4. The visual-based positioning and detection system for the welding seam of a vacuum cup according to claim 3, characterized in that, The method for dividing the weld feature map into regions of significance includes: The weld significance score of each region in the weld feature map is calculated, the texture distribution data, surface temperature distribution data and surface three-dimensional contour data of the cup body surface included in the multi-modal image information are weighted and summed to obtain the weld significance score; The first threshold value and the second threshold value of the weld significance score are preset, and the weld significance score is compared with the first threshold value and the second threshold value of the weld significance score respectively, if the weld significance score is less than the first threshold value of the weld significance score, the region corresponding to the score is marked as gray; if the weld significance score is greater than the first threshold value of the weld significance score and less than the second threshold value of the weld significance score, the region corresponding to the score is marked as yellow; if the weld significance score is greater than the second threshold value of the weld significance score, the region corresponding to the score is marked as red; Different colors are used to divide the weld feature map into regions of significance, different colors represent different significant regions, the gray region is defined as a non-weld region, the yellow region is defined as a suspected weld region, and the red region is defined as a significant weld region.

5. The visual-based positioning and detection system for the welding seam of a vacuum cup according to claim 4, characterized in that, The method for generating the thermos cup weld contour feature map includes: According to the divided significant region, different density structure light projection units are deployed on the cup body surface; the structure light projection unit is a linear array projection component, different projection strategies are adopted for different significant regions, the structure light projection unit emits sinusoidal stripes to the cup body surface for initial global scanning, the projection angle is dynamically adjusted according to the significant region of the weld feature map, wherein the projection angle interval of the significant weld region is less than that of the suspected weld region, and the projection angle interval of the suspected weld region is less than that of the non-weld region; an angle adjustment algorithm is applied to dynamically adjust the projection angle, a surface reflectivity correction model is introduced, parameters in the reflectivity model are continuously corrected by an iterative phase solving algorithm, and the process is stopped when the maximum iteration number is reached; a contour calculation matrix is established based on the triangulation principle, the surface contour is solved by the least square method, and finally the thermos cup weld contour feature map with significant markers is generated.

6. The visual-based positioning and detection system for the welding seam of a vacuum cup according to claim 5, characterized in that, The method for establishing the spatial mapping of vision and contour includes: A global coordinate system is defined with the center of the mouth of the vacuum cup as the origin, the X-axis and the Y-axis being parallel to the surface of the cup, and the Z-axis being perpendicular to the plane of the cup mouth; a calibration board is used to calibrate the linear array camera to obtain the intrinsic and extrinsic parameters of the linear array camera; the pixel coordinates of the weld feature map are converted to camera coordinates through the intrinsic parameters of the linear array camera, and the camera coordinates are converted to the global coordinates through the extrinsic parameters of the linear array camera, so as to obtain the weld feature map in the global coordinate system; the origin of the global coordinate system is taken as a reference point to calibrate the structured light projection unit to obtain the extrinsic parameters of the structured light projection unit; The weld contour feature map is converted to the global coordinate system through the extrinsic parameters of the structured light projection unit, so as to obtain the weld contour feature map in the global coordinate system; in the global coordinate system, the weld feature map and the weld contour feature map are spatially registered to establish a spatial mapping between the vision and the contour; the time source connected to the linear array camera is set as a master clock, the time source connected to the structured light projection unit is set as a slave clock, time calibration is performed through a time synchronization protocol, and a time mapping between the vision and the contour is established.

7. The visual-based positioning and detection system for the welding seam of a vacuum cup according to claim 6, characterized in that, The method for fusing the surface weld feature map and the weld contour feature map through the feature association network comprises: According to the surface weld feature map and the weld contour feature map in the global coordinate system, an association graph is constructed, a multi-layer graph convolution structure based on a multi-head attention mechanism is used to build a feature association network, and the surface weld feature map and the weld contour feature map in the global coordinate system are fused; the feature association network comprises an input layer, a feature mapping layer, a graph attention layer, a cross-modal interaction layer and an output layer; the association graph is taken as the input of the input layer of the feature association network, and a positional confidence score is generated through the output layer.

8. The visual-based positioning and detection system for the welding seam of a vacuum cup according to claim 7, characterized in that, The method for constructing the association graph comprises: Significant regions in the surface weld feature map in the global coordinate system are taken as vision nodes, contour edge regions of the weld contour feature map are taken as contour nodes, feature vectors of the two types of nodes are extracted as node features, all vision nodes and contour nodes are collected, and a node set is obtained; The node features comprise vision node features and contour node features; The nearest neighbor search is performed to find the nearest contour node for each vision node, and a bidirectional edge between the vision node and the corresponding nearest contour node is added; all bidirectional edges are collected to obtain an edge set; and the association graph is constructed based on the obtained node set and edge set.

9. The visual-based positioning and detection system for the welding seam of a vacuum cup according to claim 8, characterized in that, The method for converting the positional confidence score into an executable detection instruction based on the multi-objective decision algorithm comprises: The surface of the cup is divided into rectangular grid units, and each unit records a current contour deviation value and a positional confidence score; meanwhile, all adjustable detection parameters are listed, and the detection parameters include the adjustment range of each parameter, the energy consumption cost and the influence degree on the positioning accuracy; Three decision objectives are established, and the decision objectives include a first positioning objective, a second positioning objective and an energy efficiency objective; two types of constraint conditions are set, and the constraint conditions include a hard constraint and a flexible constraint; The multi-objective decision algorithm is used to randomly generate m groups of detection parameter optimization schemes, to evaluate the completion of three decision objectives for each scheme, to obtain the completion score of the three decision objectives, to add the completion scores of the three decision objectives, to obtain a comprehensive score, and to select one scheme with the highest comprehensive score from the m groups of schemes as a final optimization scheme; the selected final optimization scheme is converted into an actual executable detection instruction; the executable detection instruction includes a first adjustment instruction, a second adjustment instruction, and a monitoring and early warning instruction; The method for obtaining the completion score of the three decision objectives includes: The positioning accuracy improvement percentage of the cup body weld is observed as the completion score of the first positioning objective; the dispersion of the positioning reliability scores of different regions is calculated as the completion score of the second positioning objective; and the total energy consumption caused by the adjustment of all detection parameters is calculated as the completion score of the energy efficiency objective.

10. The visual-based positioning and detection system for the welding seam of a vacuum cup according to claim 9, characterized in that, The method for generating the detection strategy effectiveness index includes: The detection instruction parameters executed by the imaging controller are collected, the image quality change is synchronously monitored, the deviation of the image quality change from a preset image quality change threshold is calculated, a standardized detection strategy effectiveness index d with a value range of 0 to 1 is generated, a first threshold d1 of the standardized detection strategy effectiveness index and a second threshold d2 of the standardized detection strategy effectiveness index are preset, if d is greater than d2, it is determined that the detection strategy is effective, if d1 is less than d and d is less than d2, it is determined that the detection strategy still needs to be observed, and if d is less than d1, the detection strategy is invalid, the vision detection strategy is optimized, and early warning information is immediately generated.