Central point detection method based on diffusion model and task context information
By applying a center point detection method based on diffusion model and task context information in industrial detection and intelligent manufacturing, the problem of insufficient critical position detection accuracy in complex shapes and noise images is solved, and high-precision key position detection is achieved, improving the accuracy and reliability of the system.
Patent Information
- Application Number
- CN202510265188.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-13
AI Technical Summary
In complex shapes and noise images, it is difficult for the prior art to achieve high-precision critical position detection, resulting in increased robot operation inaccuracy and production costs.
The center point detection method based on diffusion model and task context information is adopted, and the key point position of the workpiece is obtained through the existing key point detector, the intermediate area is intercepted as the center area of the key point to be identified, and the image is processed through the diffusion model, combined with Gaussian distribution and task text information, denoising and prediction are performed, and high-precision key point coordinates are finally obtained.
Improve the accuracy of workpiece key position detection, enhance the accuracy, yield and reliability of visually guided industrial inspection and intelligent manufacturing systems, and avoid the need to add additional equipment and sensors.
Smart Images

Figure CN120147406A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial inspection and robot intelligent manufacturing, and particularly relates to a center point detection method based on a diffusion model and task context information. Background Art
[0002] In production and manufacturing, it is often necessary to visually locate the key positions (points) to be processed and guide the robot to complete specific tasks. During the manufacturing process, especially in precision manufacturing, the accuracy of visually locating the key positions (points) is crucial. Once the key positions (points) are not accurately located, the robot may make mistakes when performing related operations. Taking vision-guided automatic welding as an example, the position of the workpiece to be welded is obtained through sensors such as cameras, and then the key points in the image are located through computer vision algorithms. Then, according to the position calibration relationship between the robot coordinate system and the camera coordinate system, the positions of the key points in the robot coordinate system are calculated, and then the robot is guided to carry the welding tool to the desired position to carry out relevant welding work. A typical process is as Figure 1 shown.
[0003] However, in practical applications, the shapes of some workpieces are complex and variable, with some edges being blurred and features not being obvious. It is difficult to accurately detect them relying on traditional computer vision methods, and it is also difficult to achieve the expected accuracy and stability using deep learning-based key point detection methods. Currently, the industry uses sensors with higher accuracy and larger data volume, such as depth cameras, radars, etc., and uses existing key point detection methods to identify the key positions (points) on the workpiece, such as the YOLO system detector and its variants. However, the input of this kind of key point detector is a scaled original image, and it is difficult for the detector to comprehensively utilize all local features for accurate detection. Therefore, the accuracy is limited, and it often lacks good detection ability in complex scenes and noisy images. As Figure 2 shown, in this welding center point detection task, these detectors can detect the welding center point of the workpiece, but the accuracy is not high enough. In the figure, A represents the welding center point detected by the detector, while B represents the true value manually marked. In practical applications, usually the four detected points are connected diagonally and the center is determined as the welding key point for welding. It can be seen that there is a certain deviation between the welding center point detected by the deep learning-based detector and the true value.
[0004] Although the above methods improve the imaging accuracy and quality, there is no good perception method or solution to improve the positioning accuracy of key positions (points). And using more advanced sensors will also increase the production cost. Therefore, how to use the existing sensor data to find a high-precision key position (point) detection method has become an urgent need in the field of industrial inspection and robot intelligent manufacturing based on vision and other sensors. Summary of the Invention
[0005] To solve the problems mentioned in the above background art; the purpose of the present invention is to provide a center point detection method based on a diffusion model and task context information.
[0006] The center point detection method of the present invention based on a diffusion model and task context information has the following detection method: Use an existing key point detector to obtain the key point position coordinates of the workpiece, and intercept the middle area as the central area of the key point to be recognized through the positions of these points; and regard the marked key point coordinates (x, y) as a standard Gaussian distribution, with the key points (x, y) as the centers of the Gaussian distribution respectively, and at the same time set the variance of the Gaussian distribution, and process the above-mentioned recognized central area of the key point to be recognized through a diffusion model. When processing, combine the previous task text information as context information to guide the diffusion model to diffuse the currently noisy image with a certain deviation to obtain a heat map with the above Gaussian distribution, and use the central value of the finally predicted position distribution as the finally predicted position coordinates (x', y'), and transfer it to the original image coordinate system to obtain the key point coordinates of the workpiece; thereby improving the accuracy of workpiece key point detection.
[0007] Preferably, the diffusion model is a deep learning-based model. Using the local region of interest as the input, using task information and step information as prompt conditions to denoise the region of interest through the diffusion model, and training the diffusion model with the Gaussian distribution heat map of the key point coordinates of the manually marked workpiece, so as to obtain optimized diffusion model parameters, making it have better denoising ability, and being able to complete the prediction of the heat map of the Gaussian distribution of the center point structure under the prompt and guidance of task information and step information, so as to obtain the coordinates of the key points to be detected.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: First, improve the accuracy of detecting the key positions (points) of the workpiece, thereby improving the accuracy, yield and reliability of vision-guided industrial inspection and intelligent manufacturing system processes.
[0009] Second, compared with directly detecting the key points of the object to be processed, it has a more accurate reasoning process and is sensitive to the requirements of the task.
[0010] Third, there is no need to add additional equipment and sensors, which has no impact and burden on the process in related applications and can meet the actual use requirements. Brief Description of the Drawings
[0011] For ease of explanation, the present invention will be described in detail by the following specific embodiments and accompanying drawings.
[0012] Figure 1Schematic diagram of the vision - guided intelligent welding process of the prior art.
[0013] Figure 2 Typical corner - point detection example diagram of the prior art; Figure 3 Flowchart of the present invention; Figure 4 Flowchart for obtaining the position coordinates of an object in an image in this specific embodiment. Specific embodiment
[0014] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be described below through specific examples shown in the drawings. However, it should be understood that these descriptions are merely exemplary and do not limit the scope of the present invention. The structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the implementation conditions of the present invention. Therefore, they do not have technical substantive significance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the objectives that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention. In addition, in the following description, the descriptions of well - known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0015] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0016] As Figure 3 shown, this specific embodiment adopts the following technical solution: using the image information of the area near the detected key points of the workpiece as the region of interest (ROI) for extracting the key points of the object to be processed. Denoising and predicting this region through a diffusion model to construct a heat map with a Gaussian distribution centered on the object's center position, and then processing the heat map and detecting its center position, thereby improving the detection accuracy of key positions (points). Embodiment
[0017] This embodiment takes the vision - guided robot automatic terminal welding as an application case: 1. Obtain the coordinates of the key positions (points) of the object in the image. As Figure 4 shown, all the terminals (Pins) and their key points in the figure can be detected through a key - point detector, as Figure 4 shown; II. For one of the terminals, taking the detected key points as a reference, select the region of interest for key point extraction of the object to be processed. Through a Denosing UNet network model, denoise it and restore it to a heat map composed of a Gaussian distribution with a certain radius centered on the expected point to be processed, as shown in the following formula:
[0018] where Y xy represents the image pixel value of the Gaussian distribution region, x and y are the x and y direction coordinates of the pixel point from the center of the Gaussian distribution, p x and p y are the center coordinates of the Gaussian distribution, is the variance.
[0019] III. According to the size of the region of interest for key point extraction of the object to be processed, construct a heat map by forming a Gaussian distribution h k with the coordinates of the key positions (points) detected by the detector, and the pixel values of the remaining parts are 0. Then, use the diffusion model to process and denoise the region of interest for key point extraction of the object to be processed. Finally, obtain the distribution h 0 of the key positions (points) after denoising, and extract the mean value in h 0 to obtain the predicted value of the final key position (point).
[0020] IV. Use the text of the task and the diffusion step size as prompt conditions to guide the model to perform diffusion. Train the diffusion model with a certain number of datasets to make it have the ability to accurately predict the heat map. And train the model with text prompts for different task applications, so that the model can distinguish the characteristics of the key points to be detected for different tasks. The heat map predicted by the trained model and the heat map constructed according to the true value are optimized by the following cross-entropy loss function:
[0021] where L hm is the loss function of the heat map, Y xy is the predicted gray value at the coordinate point (x,y), is the true gray value at the coordinate point (x,y).
[0022] V. Detect the center position of the Gaussian distribution through the predicted heat map, thereby detecting the key position (point) value, calculate its position in the robot coordinate system, and guide the robot to complete the processing at this position.
[0023] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention.
[0024] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A center point detection method based on a diffusion model and task context information, characterized by: Its detection method is as follows: using the existing key point detector to obtain the key point position coordinates of the workpiece, and intercepting the middle area through the positions of these points as the central area of the key points to be identified; and treating the marked key point coordinates (x, y) as a standard Gaussian distribution, with the key points (x, y) as the centers of the Gaussian distribution, and setting the variance of the Gaussian distribution, and processing the central area of the key points to be identified through a diffusion model. During the processing, the previous task text information is combined as context information to guide the diffusion model to diffuse the current image with a certain deviation of the noise distribution to obtain a heat map with the above Gaussian distribution, and the center value of the final predicted position distribution is used as the final predicted position coordinate (x', y'), and it is transferred to the original image coordinate system to obtain the key point coordinates of the workpiece, thereby improving the accuracy of the workpiece key point detection.
2. The center point detection method based on diffusion model and task context information according to claim 1 is characterized in that: The diffusion model is a model based on deep learning. It takes the local region of interest as input, uses task information and step information as prompt conditions to denoise the region of interest through a diffusion model, and trains the diffusion model with the Gaussian distribution heat map of the key point coordinates of the manually annotated workpiece, thereby obtaining optimized diffusion model parameters, so that it has better denoising ability, and can complete the prediction of the Gaussian distribution heat map constructed by the center point under the prompt and guidance of task information and step information, thereby obtaining the coordinates of the key points that need to be detected.