Manipulator cup receiving position visual positioning method and system based on automatic soft drink making and selling machine
By adopting a lightweight object detection model with dynamic anchor box optimization and adaptive weighted fusion technology in automatic soft drinks, the problem of insufficient accuracy in complex environments is solved, and efficient and flexible visual positioning and grasping effects are achieved.
Patent Information
- Application Number
- CN202510552310.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-22
AI Technical Summary
The traditional cup positioning method has insufficient accuracy when facing changes in different sizes and environments, the robot motion control is not flexible enough, and the visual positioning algorithm has poor accuracy in complex environments, which affects the stability and user experience of automatic soft drinks and vending machines.
A lightweight object detection model based on dynamic anchor box optimization is adopted to combine the deep feature pyramid to perform multi-scale feature fusion, and a high confidence cup detection result is generated through adaptive size matching. The real-time posture data at the end of the robot arm is combined with the visual positioning results to generate the final grab coordinates, and the clamping force is monitored through the force sensor, and the coordinate conversion model parameters are dynamically adjusted to improve positioning accuracy.
Achieve accurate and fast visual positioning of cup-attached positions in complex environments, reducing dependence on unstable visual data, improving positioning robustness and reliability, and ensuring accurate capture of robotics.
Smart Images

Figure CN120347740A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automated intelligent manufacturing, and particularly relates to a visual positioning method and system for a manipulator cup receiving position based on an automatic soft drink vending machine. Background Art
[0002] Traditional cup receiving position positioning methods mostly use mechanical limits or simple photoelectric sensors. Although these methods meet the automation requirements to a certain extent, they are often not flexible and accurate enough when faced with cups of different sizes or cup position adjustments. In addition, when the mechanical structure wears or the sensor ages, the positioning error will gradually increase, affecting the stability of the vending machine and the user experience.
[0003] When performing visual positioning, problems such as unstable image quality will also be faced. For example, environmental factors such as light changes, cup reflection, or shadows may cause the image quality to decline, affecting the accuracy of visual positioning. In a complex actual environment, the image may be interfered by various noises, such as electromagnetic interference and light scattering. These noises will affect the clarity of the image and the accuracy of feature extraction. When the cup position changes greatly or the cup appearances are similar, existing visual positioning algorithms may be difficult to accurately identify and locate the cup receiving position. Different automatic soft drink vending machines may have different cup shapes, sizes, and colors, and existing algorithms may perform poorly when faced with these changes. The automatic soft drink vending machine needs to quickly respond to user needs, and the visual positioning algorithm needs to complete the positioning task within a short time, which poses a high requirement for the real-time performance of the algorithm. The movement accuracy of the manipulator directly affects the success rate of cup receiving. If the movement control of the manipulator is not precise enough, it may lead to cup receiving failure. The manipulator and the visual system need to work closely together. If the communication and control between the two are not smooth enough, it may affect the performance of the entire system. At different time periods and different environmental conditions, the light intensity and color may change, which poses a challenge to the adaptability of the visual positioning system. The automatic soft drink vending machine may need to handle various different types of cups, such as plastic cups, paper cups, glass cups, etc. These cups have differences in appearance and material, which may affect the accuracy of visual positioning.
[0004] Therefore, it is particularly important to develop an efficient, flexible, and accurate visual positioning method for the cup receiving position. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a visual positioning method and system for a manipulator cup receiving position based on an automatic soft drink vending machine; The object of the present invention can be achieved by the following technical solutions: During the operation of the automatic soft drink vending machine, after the robotic arm receives the cup receiving task, it sends a trigger signal to the automatic soft drink vending machine through the motion controller. The visual positioning module of the automatic soft drink vending machine drives the global camera to perform visual recognition on the cup dropping position through a multi-modal trigger synchronization mechanism after receiving the trigger signal, and obtains a real-time input image; Perform object detection on the real-time input image through a lightweight object detection model optimized based on dynamic anchor boxes, generate object candidate regions using adaptive size matching, and achieve multi-scale feature fusion by combining a deep feature pyramid to output a high-confidence cup detection result; Convert the cup feature points in the image coordinate system to the robotic arm base coordinate system according to the cup detection result to obtain the cup coordinates; combine the real-time pose data of the robotic arm end with the cup coordinates, and generate the final target grasping coordinates based on an adaptive weighted fusion model for real-time error evaluation.
[0006] The robotic arm performs a grasping action according to the target coordinates, and at the same time monitors the clamping force through a force sensor; if the deviation between the actual clamping position and the expected coordinates exceeds the threshold, trigger the visual unit repositioning process, and dynamically adjust the parameters of the coordinate transformation model based on historical error data; when the positioning fails 3 times continuously, the system switches to the infrared-assisted positioning mode and starts the self-diagnosis module to analyze the fault source.
[0007] Specifically, the visual positioning module includes a global camera, an embedded AI computing module, and a dynamic environment perception module; the embedded AI computing module is used to run the object detection model, and the dynamic environment perception module is used to run the adaptive weighted fusion model.
[0008] Specifically, the multi-modal trigger synchronization mechanism connects the robotic arm controller and the visual positioning module through a hardware synchronization signal line; synchronizes the multi-camera clocks using the precise time protocol and shares the time stamp with the robotic arm motion trajectory planner; embeds a motion prediction module in the visual positioning module to pre-generate a motion area based on the robotic arm joint angular velocity and perform real-time error evaluation.
[0009] The object detection model realizes the precise positioning and pose estimation of the cup through a multi-stage processing flow: a lightweight convolutional network optimized based on dynamic anchor boxes completes real-time object detection: uses a multi-image stitching enhancement technique to preprocess the input image, generates object candidate regions through an adaptive size matching mechanism, and realizes multi-scale feature fusion using a deep feature pyramid structure, and finally outputs a high-confidence cup detection result; subsequently, introduce an image segmentation strategy with local brightness adaption in the interference suppression stage: calculate the gray statistical characteristics of the pixel neighborhood through a sliding window, dynamically generate a regional binary threshold matrix, and optimize the edge smoothness by combining the Gaussian weighted mean to effectively separate the cup main body from the environmental noise; In the feature extraction stage, a vision-based positioning framework with multi-sensor fusion is adopted to construct three-dimensional space features: feature tracking, local map optimization, and loop detection are synchronously processed through a parallel thread architecture. A fast binary descriptor is used to construct multi-level feature matching relationships, and inertial measurement data is combined to achieve sub-pixel edge feature point positioning. The matching accuracy is improved by a probability-driven outlier iterative screening mechanism: a random subset sampling is used to establish an initial geometric constraint model, the spatial consistency score of the matching point pairs is calculated, and outlier data points deviating from the main model are removed through multiple rounds of weighted voting mechanism. Finally, a cup pose feature descriptor with topological consistency is generated.
[0010] According to the cup detection result, the cup feature points in the image coordinate system are transformed into the robotic arm base coordinate system; a dynamic weight Kalman filtering algorithm is introduced to fuse the real-time pose data of the robotic arm end and the vision-based positioning result to generate the final target grasping coordinates. The adaptive weighted fusion model combines the real-time pose data of the robotic arm end and the cup coordinates obtained by vision-based positioning to generate the final grasping coordinates: , where, P grasp is the grasping coordinate, P visual is the cup coordinate obtained by vision-based positioning, P end-effector is the real-time position of the end effector of the robotic arm, c is the confidence score of vision detection, and e represents the error evaluation coefficient; The confidence score of the vision detection is obtained from the output of the target detection model and is bound to the corresponding cup coordinates. The error evaluation coefficient is determined based on the motion prediction module embedded in the vision-based positioning module and is used to reflect the prediction error during the motion of the robotic arm.
[0011] Specifically, the transformation of the vision-based positioning cup coordinates is realized through a homogeneous transformation matrix: the cup feature points in the image coordinate system obtained by the global camera are transformed using the camera intrinsic matrix and the camera extrinsic matrix to obtain the cup coordinates relative to the camera coordinate system. Combining the relative position relationship between the robotic arm base coordinate system and the camera coordinate system of the automatic soft drink vending machine, the cup coordinates are further transformed into the robotic arm base coordinate system through the homogeneous transformation matrix to obtain the final cup coordinates.
[0012] Specifically, the dynamic environment perception module dynamically adjusts the target object characteristics and weights according to the confidence level output by the target detection model. When the confidence level is lower than the preset imaging blur threshold, the weight of the vision data is reduced, the dependence on the high-reflection area is reduced, and at the same time, the weight of the force sensor is increased by combining the robotic arm motion planning data.
[0013] The vision-based positioning system for the robotic arm cup receiving position of the automatic soft drink vending machine includes a motion trigger unit, a target detection unit, and a fusion positioning unit. The motion trigger unit is used to, during the operation of the automatic soft drink vending machine, after the manipulator receives the cup receiving task, send a trigger signal to the automatic soft drink vending machine through the motion controller. The visual positioning module of the automatic soft drink vending machine drives the global camera to perform visual recognition on the cup dropping position through a multi-modal trigger synchronization mechanism after receiving the trigger signal, and obtains a real-time input image; The target detection unit performs target detection on the real-time input image through a lightweight target detection model based on dynamic anchor box optimization, generates target candidate regions by using adaptive size matching, combines deep feature pyramids to achieve multi-scale feature fusion, and outputs a cup detection result with high confidence; The fusion positioning unit converts the cup feature points in the image coordinate system to the manipulator base coordinate system according to the cup detection result to obtain the cup coordinates; combining the real-time pose data of the manipulator end with the cup coordinates, an adaptive weighted fusion model based on real-time error evaluation generates the final target grasping coordinates.
[0014] The beneficial effects of the present invention are as follows: By adopting the above technical solutions, the present invention can achieve accurate and fast visual positioning of the cup receiving position in a complex dynamic environment. Under adverse conditions such as light changes, cup reflection or shadows, by dynamically adjusting the target object characteristics and weights, the present invention can reduce the dependence on unstable visual data and improve the robustness of positioning. At the same time, combining the manipulator motion planning data and the force sensor weights further improves the accuracy and reliability of positioning. In addition, the multi-modal trigger synchronization mechanism of the present invention ensures the fast response of the trigger signal and the precise synchronization of the camera clock, providing strong support for real-time image processing. In summary, the present invention shows significant advantages in the visual positioning of the cup receiving position of the manipulator of the automatic soft drink vending machine, and makes a positive contribution to the development of automated intelligent manufacturing technology. Description of the Drawings
[0015] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.
[0016] Figure 1 It is a flow schematic diagram of the visual positioning method for the cup receiving position of the manipulator based on the automatic soft drink vending machine of the present invention. Detailed Embodiments
[0017] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific embodiments, structures, features and their effects according to the present invention.
[0018] Please refer to Figure 1, A visual positioning method and system for the manipulator cup receiving position based on an automatic soft drink vending machine, specifically including: Trigger signal and visual recognition: During the operation of the automatic soft drink vending machine, after the manipulator receives the cup receiving task, it sends a trigger signal to the automatic soft drink vending machine through the motion controller. After receiving the trigger signal through the multi-modal trigger synchronization mechanism, the visual positioning module of the automatic soft drink vending machine drives the global camera to perform visual recognition on the cup dropping position and obtains the real-time input image.
[0019] Object detection and feature extraction: Perform object detection on the real-time input image through a lightweight object detection model optimized based on dynamic anchor boxes, use adaptive size matching to generate target candidate regions, and combine deep feature pyramids to achieve multi-scale feature fusion, and output a cup detection result with high confidence.
[0020] Coordinate transformation and grasping coordinate generation: According to the cup detection result, convert the cup feature points in the image coordinate system to the manipulator base coordinate system to obtain the cup coordinates. Combining the real-time pose data of the manipulator end with the cup coordinates, an adaptive weighted fusion model based on real-time error evaluation generates the final target grasping coordinates.
[0021] The visual positioning module includes: a global camera: used to collect image data of the cup dropping position in real time. An embedded AI computing module: used to run the lightweight object detection model, using the NVIDIA Jetson TX2 platform, and performing model quantization acceleration through TensorRT. The inference latency optimization formula is: , where N ops is the computational amount of the model, f GPU is the main frequency of the GPU, and Q is the quantization factor.
[0022] In this embodiment, a lightweight YOLOv5s model is used to detect the cup body in real-time images, and an adaptive threshold segmentation algorithm is combined to eliminate environmental interference; the improved ORB-SLAM3 framework is used to extract the feature points of the cup edge, and the RANSAC algorithm is used to eliminate abnormal matching points to generate a high-precision cup pose feature descriptor; by combining the real-time performance of the lightweight network with the dynamic anchor box optimization, a processing ability of 30FPS for 608×608 resolution images is achieved with a model size of 27.3MB; by adopting a local binary pattern strategy with adjustable neighborhood block size, the segmentation algorithm can still maintain an effective area recognition rate of more than 90% under the condition of ±50% light change; by improving the feature matching thread scheduling mechanism, the ORB feature extraction speed is increased by 40% while maintaining the matching recall rate of the BRIEF descriptor; by designing an iterative termination condition based on probability density estimation, the computational complexity of the abnormal point elimination process is reduced from O(n³) to O(nlog n).
[0023] The multi-modal trigger synchronization mechanism connects the robotic arm controller and the vision system through a hardware synchronization signal line to ensure that the trigger delay < 1ms; the PTP (Precision Time Protocol) is used to synchronize the clocks of multiple cameras and share the time stamps with the robotic arm motion trajectory planner; a motion prediction module is embedded in the vision algorithm to pre-generate the ROI area based on the joint angular velocity of the robotic arm, reducing the image processing time by more than 20%.
[0024] During the execution stage of the grasping action, the robotic arm accurately moves to the specified position according to the generated target grasping coordinates. At the same time, the force sensor continuously monitors the clamping force to ensure that the cup will not be damaged or slip during the grasping process. If there is a deviation between the actual clamping position and the expected coordinates, and this deviation exceeds the preset threshold, the system will immediately trigger the repositioning process of the vision unit. In this process, the vision positioning module will re-identify the cup visually and dynamically adjust the parameters of the coordinate transformation model based on the historical error data to improve the accuracy of subsequent positioning.
[0025] When the positioning fails three times in a row, the system will automatically switch to the infrared-assisted positioning mode. In this mode, the infrared sensor will play a key role in determining the position of the cup by detecting the infrared reflection signal of the cup. At the same time, the system starts the self-diagnosis module to diagnose and analyze possible fault sources in order to repair the problems in time and ensure the continuous and stable operation of the system.
[0026] In addition, the system of the present invention also has a high degree of flexibility and scalability. It can adapt to different types of cups, such as plastic cups, paper cups, glass cups, etc., as well as cups of different shapes, sizes and colors. By adjusting the parameters of the target detection model and the coordinate transformation model, the system can easily handle various The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may, for example, but not be limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component.
[0027] The computer-readable signal media may include data signals propagated in a baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component.
[0028] The program code contained on the computer-readable media may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0029] The above are only the preferred embodiments of the present invention and do not impose any formal limitations on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content without departing from the technical solution of the present invention. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A visual positioning method for the cup receiving position of a manipulator based on an automatic soft drink vending machine, characterized in that, Including: During the operation of the automatic soft drink vending machine, after the robotic arm receives the cup receiving task, it sends a trigger signal to the automatic soft drink vending machine through the motion controller. The visual positioning module of the automatic soft drink vending machine drives the global camera to perform visual recognition on the cup dropping position through a multi-modal trigger synchronization mechanism after receiving the trigger signal, and obtains a real-time input image; Performs object detection on the real-time input image through a lightweight object detection model optimized based on dynamic anchor boxes, generates object candidate regions using adaptive size matching, combines deep feature pyramids to achieve multi-scale feature fusion, and outputs a cup detection result with high confidence; Converts the cup feature points in the image coordinate system to the robotic arm base coordinate system according to the cup detection result to obtain the cup coordinates; combines the real-time pose data of the robotic arm end with the cup coordinates, and generates the final target grasping coordinates based on an adaptive weighted fusion model for real-time error evaluation.
2. The method according to claim 1, characterized in that, The visual positioning module includes a global camera, an embedded AI computing module, and a dynamic environment perception module; the embedded AI computing module is used to run the object detection model, and the dynamic environment perception module is used to run the adaptive weighted fusion model.
3. The method according to claim 1, characterized in that The multi-modal trigger synchronization mechanism connects the robotic arm controller and the visual positioning module through a hardware synchronization signal line; synchronizes the multi-camera clocks using the Precision Time Protocol and shares timestamps with the robotic arm motion trajectory planner; embeds a motion prediction module in the visual positioning module to pre-generate a motion area based on the robotic arm joint angular velocity and perform real-time error evaluation.
4. The method according to claim 1, characterized in that The object detection model realizes the precise positioning and pose estimation of the cup through a multi-stage processing process, including: A lightweight convolutional network optimized based on dynamic anchor boxes performs real-time object detection: preprocesses the input image using a multi-image stitching enhancement technique, generates object candidate regions through an adaptive size matching mechanism, and performs multi-scale feature fusion using a deep feature pyramid structure, and outputs a cup detection result with high confidence; In the interference suppression stage, an image segmentation strategy with local brightness adaption is introduced: calculates the gray statistical characteristics of the pixel neighborhood through a sliding window, dynamically generates a regional binary threshold matrix, and combines the Gaussian weighted mean to optimize the edge smoothness to separate the cup main body from the environmental noise.
5. The method according to claim 1, wherein The adaptive weighted fusion model combines the real-time pose data of the robotic arm end with the cup coordinates obtained by visual positioning to generate the final grasping coordinates: , Among them, P grasp is the grasping coordinate, P visual is the cup body coordinate for visual positioning, P end-effector is the real-time position of the end effector of the robotic arm, c is the confidence score of visual detection, and e represents the error evaluation coefficient; The confidence score of the visual detection is obtained through the output of the object detection model and is bound to the corresponding cup coordinates; the error evaluation coefficient is determined based on the motion prediction module embedded in the visual positioning module and is used to reflect the prediction error during the motion of the robotic arm.
6. The method according to claim 1, characterized in that, The conversion of the visual positioning cup body coordinates is achieved through a homogeneous transformation matrix: the cup body feature points in the image coordinate system obtained by the global camera are subjected to coordinate conversion using the camera internal parameter matrix and the camera external parameter matrix to obtain the cup body coordinates relative to the camera coordinate system; combined with the relative position relationship between the robotic arm base coordinate system and the camera coordinate system of the automatic soft drink vending machine, the cup body coordinates are further converted to the robotic arm base coordinate system through the homogeneous transformation matrix to obtain the final cup body coordinates.
7. The method according to claim 2, wherein The dynamic environment perception module dynamically adjusts the target object characteristics and weights according to the confidence level output by the target detection model. When the confidence level is lower than the preset imaging blur threshold, it reduces the weight of visual data, reduces the dependence on high-reflection areas, and at the same time combines the robotic arm motion planning data to increase the weight of the force sensor.
8. A visual positioning system for the cup receiving position of the manipulator based on an automatic soft drink vending machine, which is used to execute the method described in claims 1-7, characterized in that, It includes a motion trigger unit, a target detection unit, and a fusion positioning unit; The motion trigger unit is used to, during the operation of the automatic soft drink vending machine, after the robotic arm receives the cup receiving task, send a trigger signal to the automatic soft drink vending machine through the motion controller. The visual positioning module of the automatic soft drink vending machine drives the global camera to perform visual recognition on the cup dropping position through a multi-modal trigger synchronization mechanism after receiving the trigger signal, and obtains a real-time input image; The target detection unit performs target detection on the real-time input image through a lightweight target detection model based on dynamic anchor box optimization, generates target candidate regions using adaptive size matching, and combines a deep feature pyramid to achieve multi-scale feature fusion, and outputs a high-confidence cup body detection result; The fusion positioning unit converts the cup body feature points in the image coordinate system to the robotic arm base coordinate system according to the cup body detection result to obtain the cup body coordinates; combined with the real-time pose data of the robotic arm end and the cup body coordinates, an adaptive weighted fusion model based on real-time error evaluation generates the final target grasping coordinates.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that , when the processor executes the computer program, it implements the visual positioning method for the cup receiving position of the robotic arm based on the automatic soft drink vending machine as described in any one of claims 1-7.
10. A storage medium, on which a computer program is stored, characterized in that , when this program is executed by the processor, it implements the visual positioning method for the cup receiving position of the robotic arm based on the automatic soft drink vending machine as described in any one of claims 1-7.