Intelligent factory voice and visual cooperative control method and system
By using a voice and vision collaborative control method in a smart factory, efficient grasping of complex objects and interference scenarios has been achieved, solving the problem of low grasping efficiency in traditional methods and improving system response speed and grasping success rate.
Patent Information
- Application Number
- CN202511186926.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-25
AI Technical Summary
Traditional smart factory robotic arm grasping systems have low grasping efficiency when handling complex objects or interference scenarios, making it difficult to achieve efficient grasping.
The system employs a voice and vision-based collaborative control method. It wakes up the device with a wake word, recognizes voice commands to acquire images of the target object, determines feature points and the type of grasping difficulty, calculates the robotic arm's grasping posture, processes complex objects using single-angle or multi-angle coordinate parameters, eliminates interference or corrects shooting conditions, and achieves precise positioning and grasping.
It significantly improves system response speed and grasping success rate, reduces traditional operation delays, and improves the grasping accuracy and efficiency of complex objects through quantitative difficulty assessment and feature point processing.
Smart Images

Figure CN120715909B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of industrial automation and machine vision technology, and in particular to a voice and vision cooperative control method and system for intelligent factory. BACKGROUND
[0002] In the traditional intelligent factory mechanical arm grabbing system, a fixed programming path or a single vision guiding method is usually used to control the movement of the mechanical arm. The prior art dynamically adjusts the illuminator according to the comparison result to dynamically adjust the light intensity around the product on the production line, real-time acquires and analyzes image data from the multispectral camera to identify defects on the surface of the product, and removes the product with defects. However, the traditional method uses single-view multi-point coordinate calculation, which is difficult to handle complex objects or interference scenes, resulting in low grabbing efficiency.
[0003] Chinese patent application No. CN202411173365.4 discloses an intelligent factory production line vision detection system and method, which comprises: an environment perception module that acquires the light intensity around the product on the production line in real time; a system integrated control module that receives the light intensity around the product in real time and compares it with a preset threshold; a dynamic adjustment illuminator that dynamically adjusts the light intensity around the product on the production line according to the comparison result; a multispectral camera module that acquires images of the product that meets the light intensity in real time; an intelligent identification and processing module that acquires and analyzes image data from the multispectral camera in real time to identify defects on the surface of the product; if the identified product surface has defects, the product with defects is removed by an automatic removal module and placed in a defective product placement area; if the products placed in the defective product placement area exceed two-thirds, the remote interactive control panel module is fed back, and the production personnel are notified by the warning module to detect the products in the defective product placement area to determine whether there is a false detection.
[0004] However, the prior art has the following problems:
[0005] The traditional method uses single-view multi-point coordinate calculation, which is difficult to handle complex objects or interference scenes, resulting in low grabbing efficiency. SUMMARY
[0006] Therefore, the present application provides an intelligent factory voice and vision cooperative control method and system to overcome the problem of low grabbing efficiency caused by the traditional method of using single-view multi-point coordinate calculation to handle complex objects or interference scenes.
[0007] To achieve the above purpose, the present application provides an intelligent factory voice and vision cooperative control method, which comprises:
[0008] Step S1, according to the preset wake-up word, the device is woken up, the voice instruction is recognized through the preset instruction word, and the target object image is collected when the voice recognition is successful;
[0009] Step S2, determining the feature points of the target object based on the target object image, and determining the grabbing difficulty tendency type according to a ratio of the number of feature points located in the feature region to the total number of feature points;
[0010] Step S3, determining the calculation parameters of the target object grabbing posture based on the determined grabbing difficulty tendency type and the shape complexity degree of the target object, including:
[0011] taking the single-angle target object space three-dimensional coordinate parameter as the calculation sub-parameter of the target posture three-dimensional coordinate of the mechanical arm;
[0012] taking the multi-angle target object space three-dimensional coordinate parameter as the calculation sub-parameter of the target posture three-dimensional coordinate of the mechanical arm, analyzing whether there is a conflict of the target object space three-dimensional coordinate parameters under different angles, analyzing whether there is a local space three-dimensional coordinate distortion phenomenon of the target object, and determining the interference coordinate to be excluded or the shooting condition to be corrected based on the analysis result;
[0013] Step S4, determining the grabbing posture of the mechanical arm according to the determined grabbing posture calculation sub-parameter, and grabbing the target object.
[0014] Further, in the step S2, the grabbing difficulty tendency type is determined according to the ratio of the number of feature points located in the feature region to the total number of feature points, including:
[0015] calculating the ratio of the number of feature points located in the feature region to the total number of feature points to obtain a grabbing difficulty representation,
[0016] if the grabbing difficulty representation is less than or equal to a preset grabbing difficulty representation, it is determined that the grabbing difficulty tendency type of the target object is strong grabbing difficulty tendency;
[0017] if the grabbing difficulty representation is greater than the preset grabbing difficulty representation, it is determined that the grabbing difficulty tendency type of the target object is weak grabbing difficulty tendency.
[0018] Further, in the step S3, the calculation parameters of the target object grabbing posture are determined based on the determined grabbing difficulty tendency type and the shape complexity degree of the target object, including:
[0019] if the grabbing difficulty tendency type of the target object is weak grabbing difficulty tendency and the shape complexity degree of the target object is low, the single-angle target object space three-dimensional coordinate parameter is taken as the calculation sub-parameter of the target posture three-dimensional coordinate of the mechanical arm;
[0020] If the target object has a strong grasping difficulty tendency or a high shape complexity, the multi-angle target object space three-dimensional coordinate parameters are used as the calculation sub-parameters of the mechanical arm target posture three-dimensional coordinates, and it is analyzed whether there is a conflict of the target object space three-dimensional coordinate parameters at different angles, and whether there is a local space three-dimensional coordinate distortion phenomenon of the target object, and based on the analysis result, the interference coordinates are excluded or the shooting conditions are corrected.
[0021] Further, in the step S3, the contour feature of the target object is extracted based on the target object image, the contour curvature is calculated, and the shape complexity of the target object is determined based on the contour curvature.
[0022] Further, the analysis of whether there is a conflict of the target object space three-dimensional coordinate parameters at different angles includes:
[0023] Convert the coordinates to the same reference system,
[0024] Align the data by hand-eye calibration,
[0025] Select the key points of the target object,
[0026] Calculate the coordinates of the same key point at different angles,
[0027] Calculate the Euclidean distance deviation,
[0028] Determine the key points with a Euclidean distance deviation greater than or equal to a preset Euclidean distance deviation, and obtain abnormal key points,
[0029] Count the number of abnormal key points,
[0030] If the number of abnormal key points is greater than or equal to a preset number, it is determined that there is a conflict of the target object space three-dimensional coordinate parameters at different angles;
[0031] If the number of abnormal key points is less than the preset number, it is determined that there is no conflict of the target object space three-dimensional coordinate parameters at different angles.
[0032] Further, the analysis of whether there is a local space three-dimensional coordinate distortion phenomenon of the target object includes:
[0033] Divide the target object into several local regions,
[0034] Calculate the variance of the coordinates of a single local region,
[0035] Calculate the average variance of the coordinates of each local region,
[0036] Calculate the difference between the variance and the average variance of the coordinates of a single local region, and obtain a distortion representation,
[0037] counting a number of local regions with distortion characterization greater than or equal to a preset distortion characterization, to obtain a distortion region number,
[0038] if the distortion region number is greater than or equal to a preset distortion region number, determining that there is a local spatial three-dimensional coordinate distortion phenomenon of the target object;
[0039] if the distortion region number is less than the preset distortion region number, determining that there is no local spatial three-dimensional coordinate distortion phenomenon of the target object.
[0040] Further, the determining of the excluded interference coordinates or the correction of the shooting conditions based on the analysis result comprises:
[0041] if there is a conflict of the spatial three-dimensional coordinate parameters of the target object under different angles and the local spatial three-dimensional coordinate distortion phenomenon of the target object, determining to exclude the interference coordinates;
[0042] if there is only a conflict of the spatial three-dimensional coordinate parameters of the target object under different angles or only the local spatial three-dimensional coordinate distortion phenomenon of the target object, determining to correct the shooting conditions.
[0043] Further, the excluding of the interference coordinates comprises:
[0044] eliminating the abnormal key points and the local regions with distortion characterization greater than or equal to a preset distortion characterization.
[0045] Further, the correcting of the shooting conditions comprises:
[0046] increasing a shooting angle to increase the spatial three-dimensional coordinate parameters of the target object, wherein the increase amount of the shooting angle is positively correlated with the distortion region number.
[0047] The application provides an intelligent factory voice and vision cooperative control system, comprising:
[0048] Further, a voice instruction recognition module is used to recognize a wake-up instruction based on a preset wake-up word and recognize a voice instruction based on a preset instruction word.
[0049] An image acquisition module is connected with the voice instruction recognition module and used to acquire a target object image.
[0050] A feature point recognition module is connected with the voice instruction recognition module and the image acquisition module respectively and used to determine a feature point of the target object and determine a grasping difficulty tendency type of the target object based on the feature point.
[0051] An analysis module connected with the image acquisition module and the feature point recognition module respectively, used to determine single-angle target object space three-dimensional coordinate parameters or multi-angle target object space three-dimensional coordinate parameters as the calculation sub-parameters of the mechanical arm target posture three-dimensional coordinates based on the target object's grasping difficulty tendency type, analyze whether there is a conflict of target object space three-dimensional coordinate parameters under different angles, analyze whether there is a local target object space three-dimensional coordinate distortion phenomenon, and determine the interference coordinates to be excluded or the shooting conditions to be corrected based on the analysis result;
[0052] A calculation module connected with the image acquisition module, the feature point recognition module and the analysis module, used to calculate the grasping pose of the mechanical arm based on the calculation sub-parameters of the determined mechanical arm target posture three-dimensional coordinates.
[0053] Compared with the prior art, the beneficial effects of the present application are that in the present application, contactless control is realized by pre-setting the wake-up word, the operator can quickly start the grasping process through natural language processing instructions, significantly reducing the time delay of traditional button or touch screen operation, the voice instruction is seamlessly connected with visual acquisition, avoiding the manual confirmation link, improving the system response speed, the grasping difficulty classification is based on the feature point distribution ratio (feature area ratio), the single-angle or multi-angle coordinate calculation mode is automatically selected according to the object characteristics, and the multi-angle coordinate calculation mode is used to improve the grasping success rate for target objects with greater grasping difficulty.
[0054] Further, in the present application, the ratio of the number of feature points in the feature area to the total number of feature points (grasping difficulty representation) is calculated to realize the qualitative to quantitative change of difficulty evaluation, for low-texture objects, even if the total number of feature points is small, but the uniform distribution is still determined as weak difficulty, avoiding excessive conservative strategy, for local high-feature objects, the feature points are concentrated, the strong difficulty is accurately identified, high-precision calculation is triggered, abnormal point clouds are automatically shielded in the light reflection area, and the grasping success rate of curved surface objects is improved.
[0055] Further, in the present application, the shape complexity of the target object is converted into a quantitative index (such as curvature distribution, high curvature area ratio, etc.) by extracting the contour and calculating the curvature, avoiding subjective judgment errors, the present application does not depend on the specific type of the target object (such as regular geometric body, irregular object), the curvature distribution can be used to distinguish simple shapes (such as cylindrical body) and complex shapes (such as gear, blade), providing a basis for subsequent classification, matching or path planning, through hand-eye calibration and coordinate unification, the spatial alignment problem of multi-sensor (such as camera) or multi-angle measurement data is solved, and the positioning deviation caused by non-uniform coordinate system is avoided.
[0056] Further, the present application can accurately locate the local distortion in three-dimensional coordinates (such as local dense or sparse point cloud, local geometric structure deformation) by dividing the target into local areas and calculating the variance difference, which is more sensitive to subtle abnormalities than global analysis. Through distortion detection and intelligent decision-making, the influence of noise, outliers and coordinate conflicts on the three-dimensional model is significantly reduced, and the reconstruction accuracy and stability are improved. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 Flow chart of the intelligent factory voice and visual cooperative control method of the present application;
[0058] Figure 2 Flow chart of determining the grasping difficulty tendency type of the present application;
[0059] Figure 3 Flow chart of determining the calculation parameters of the grasping posture for the target of the present application;
[0060] Figure 4 Flow chart of determining whether there is a conflict of the spatial three-dimensional coordinate parameters of the target under different angles of the present application. DETAILED DESCRIPTION
[0061] In order to make the objects and advantages of the present application more clear and apparent, the present application will be further described below in conjunction with the embodiments; it should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the present application.
[0062] It should be pointed out that the data in the present embodiment are obtained by comprehensive analysis and evaluation of the historical data and the corresponding historical determination results of the present application in the past 6 months before the present determination. Those skilled in the art can understand that the determination method of the present application for a single parameter can be to select the value with the highest proportion as the preset standard parameter according to the data distribution, to use weighted summation to obtain the value as the preset standard parameter, to substitute each historical data into a specific formula and to obtain the value by using the formula as the preset standard parameter, or other selection methods, as long as the present application can clearly define different specific situations in the single determination process by the obtained value.
[0063] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application, and are not intended to limit the protection scope of the present application.
[0064] It should be noted that in the description of the present application, the terms of direction or position relationship indicated by "upper", "lower", "left", "right", "inner", "outer" and the like are based on the direction or position relationship shown in the drawings, which is only for the convenience of description, and does not indicate or imply that the device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0065] In addition, it should be further pointed out that in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected, it can be mechanically connected, or it can be electrically connected, it can be directly connected, or indirectly connected through an intermediate medium, it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0066] Please refer to Figure 1 The intelligent factory voice and vision cooperative control method provided by the embodiment of the present application,
[0067] The intelligent factory voice and vision cooperative control method provided by the embodiment of the present application,
[0068] Step S1, completing device wake-up according to a preset wake-up word, identifying a voice instruction through a preset instruction word, and collecting a target object image when voice recognition is successful;
[0069] Step S2, determining feature points of the target object based on the target object image, and determining a grasping difficulty tendency type according to a ratio of the number of feature points located in a feature region to the total number of feature points;
[0070] Step S3, determining a calculation parameter of a grasping posture for the target object based on the determined grasping difficulty tendency type and the shape complexity of the target object, including:
[0071] Taking a single-angle target object space three-dimensional coordinate parameter as a calculation sub-parameter of a mechanical arm target posture three-dimensional coordinate;
[0072] Taking a multi-angle target object space three-dimensional coordinate parameter as a calculation sub-parameter of a mechanical arm target posture three-dimensional coordinate, analyzing whether there is a conflict of the target object space three-dimensional coordinate parameter under different angles, analyzing whether there is a local space three-dimensional coordinate distortion phenomenon of the target object, and determining to exclude interference coordinates or correct shooting conditions based on the analysis result;
[0073] Step S4, determining a grasping posture of the mechanical arm according to the determined grasping posture calculation sub-parameter, and grasping the target object.
[0074] Specifically, in this embodiment, the feature points refer to unique pixel points extracted from the target object image by a feature point detection algorithm, which can represent the local texture, contour, edge or geometric structure of the target object. Such feature points have local uniqueness and can be used for shape recognition and spatial positioning of the target object through their position, gray scale change, gradient direction and other characteristic parameters. They are the basic units for describing the surface features of the target object.
[0075] Specifically, in this embodiment, the feature region refers to a specific region on the surface of the target object that is suitable for stable grasping by the mechanical arm gripper. This region needs to meet the structural stability (such as a smooth surface without significant protrusions or depressions), gripping reliability (such as having a certain friction or geometric constraints to avoid the target object from slipping during grasping), and other characteristics. It is a key region for evaluating the feasibility of grasping the target object. The distribution of feature points in the feature region directly reflects the grasping adaptability of the region, and the proportion of the number of feature points in the entire target object is used to quantitatively represent the grasping difficulty of the target object.
[0076] In this application, the pre-set wake-up word is used to realize contactless control. The operator can quickly start the grasping process through natural language processing instructions, significantly reducing the time delay of traditional button or touch screen operation. The voice command seamlessly connects with visual acquisition, avoiding manual confirmation steps, improving system response speed, and automatically selecting single-angle or multi-angle coordinate calculation mode based on the feature point distribution ratio (feature region proportion) of the grasping difficulty classification. For target objects with high grasping difficulty, the multi-angle coordinate calculation mode is used to improve the success rate of grasping.
[0077] Specifically, in this embodiment, the grasping pose of the mechanical arm can be determined in the following way:
[0078] Twenty-five sets of calibration board data are collected by calling camera resources through OpenCV and RealSense SDK, and the corresponding mechanical arm pose (including: mechanical arm position parameters (X, Y, Z) and mechanical arm attitude parameters (RX, RY, RZ)) of each set of calibration board data is recorded.
[0079] The size of the checkerboard is adjusted according to the number and size of the checkerboards in the calibration board, the number of checkerboards in the horizontal coordinate axis and the number of checkerboards in the vertical coordinate axis parameters, the camera intrinsic parameters and distortion coefficients are obtained, the calibration program is run, and the hand-eye matrix between the current camera and the end of the mechanical arm is obtained.
[0080] The photographing pose of the mechanical arm is adjusted, and the position parameters and attitude parameters (X, Y, Z, RX, RY, RZ) of the mechanical arm when taking pictures are recorded.
[0081] The YOLO visual recognition program is run, combined with the three-dimensional coordinate calculation API of the depth camera, to obtain the type of the target object and the spatial three-dimensional coordinates of the target object under the current camera.
[0082] According to the hand-eye matrix, the spatial three-dimensional coordinates of the target object, and the pose parameters (RX, RY, RZ) of the robot arm, the position parameters (X, Y, Z) of the robot arm are calculated, and finally the complete motion parameters (X, Y, Z, RX, RY, RZ) of the robot arm motion are obtained;
[0083] For the target object with irregular shape or inconvenient placement for grabbing, the object positioning and pose estimation are performed through GraspNet, and the motion parameters (X, Y, Z, RX, RY, RZ) of the robot arm are calculated according to the hand-eye matrix, the spatial three-dimensional coordinates of the target object, and the pose estimation result of GraspNet;
[0084] According to the determined motion data (X, Y, Z, RX, RY, RZ) of the robot arm, the robot arm is controlled to move to the grabbing pose;
[0085] The driving gripper completes the grabbing of the target object and places the target object to the specified position.
[0086] Specifically, in this embodiment, the conversion from the pixel coordinates in the camera to the spatial three-dimensional coordinates is obtained through the camera SDK.
[0087] Specifically, in this embodiment, the related parameters are adjusted according to the number and size of the checkerboards in the calibration board, wherein the parameters include: the size of the checkerboard, the number of checkerboards in the horizontal coordinate axis, and the number of checkerboards in the vertical coordinate axis.
[0088] Specifically, in step S2, the grabbing difficulty tendency type is determined according to the ratio of the number of feature points in the feature region to the total number of feature points, including:
[0089] The multi-target recognition model is used to complete the recognition of the target objects in the camera field of view,
[0090] The feature extraction algorithm is used to complete the feature extraction of the target objects, and the ratio of the number of feature points in the feature region to the total number of feature points is calculated to obtain the grabbing difficulty representation.
[0091] If the grabbing difficulty representation is less than or equal to the preset grabbing difficulty representation, it is determined that the grabbing difficulty tendency type of the target object is strong grabbing difficulty tendency.
[0092] If the grabbing difficulty representation is greater than the preset grabbing difficulty representation, it is determined that the grabbing difficulty tendency type of the target object is weak grabbing difficulty tendency.
[0093] Specifically, in this embodiment, the preset grasping difficulty representation can be determined by the following method: 100 typical part samples are collected and divided into two categories: easy-to-grasp samples (50): parts with a successful grasping rate > 95% (such as square parts with smooth surfaces); difficult-to-grasp samples (50): parts with a grasping failure rate > 30% (such as small cylindrical parts, flexible objects); feature points of the target object image are extracted using SIFT / SURF and other feature point detection algorithms, the number of feature points located in the feature region and the total number of feature points are counted, the ratio of the number of feature points to the total number of feature points is calculated, and the R value distribution of the samples is analyzed: the R value of the easy-to-grasp sample is concentrated in the interval of 0.65~0.95, and the average value is 0.82; the R value of the difficult-to-grasp sample is concentrated in the interval of 0.20~0.55, and the average value is 0.38; the dividing point (such as 0.55) of the R value distribution of the two types of samples is selected as the preset grasping difficulty representation.
[0094] Specifically, in step S3, the calculation parameters of the target object grasping posture are determined based on the determined grasping difficulty tendency type and the shape complexity of the target object, including:
[0095] If the grasping difficulty tendency type of the target object is weak grasping difficulty tendency and the shape complexity of the target object is low, the single-angle target object space three-dimensional coordinate parameter is used as the calculation sub-parameter of the mechanical arm target posture three-dimensional coordinate;
[0096] If the grasping difficulty tendency type of the target object is strong grasping difficulty tendency or the shape complexity of the target object is high, the multi-angle target object space three-dimensional coordinate parameter is used as the calculation sub-parameter of the mechanical arm target posture three-dimensional coordinate, whether there is a conflict of the target object space three-dimensional coordinate parameter under different angles is analyzed, whether there is a local space three-dimensional coordinate distortion phenomenon of the target object is analyzed, and the interference coordinate is excluded or the shooting condition is corrected based on the analysis result.
[0097] In the present application, by calculating the ratio of the number of feature points in the feature region to the total number of feature points (grasping difficulty representation), the qualitative to quantitative change of difficulty evaluation is realized. For low-texture objects, even if the total number of feature points is small, uniform distribution is still determined as weak difficulty, avoiding excessive conservative strategy. For local high-feature objects, the strong difficulty is accurately identified when the feature points are concentrated, triggering high-precision calculation, automatically shielding abnormal point cloud in the light reflection area, and improving the success rate of grasping curved surface objects.
[0098] Specifically, in step S3, the contour feature of the target object is extracted based on the target object image, the contour curvature is calculated, and the shape complexity of the target object is determined based on the contour curvature.
[0099] Specifically, the analysis of whether there is a conflict of the target object space three-dimensional coordinate parameter under different angles includes:
[0100] transform the coordinates to the same reference system,
[0101] align the data using hand-eye calibration,
[0102] select the key points of the target object,
[0103] calculate the coordinates of the same key point at different angles,
[0104] calculate the Euclidean distance deviation,
[0105] determine the key points with a Euclidean distance deviation greater than or equal to a preset Euclidean distance deviation, obtaining abnormal key points,
[0106] count the number of abnormal key points,
[0107] if the number of abnormal key points is greater than or equal to a preset number, it is determined that there is a conflict in the spatial three-dimensional coordinate parameters of the target object at different angles;
[0108] if the number of abnormal key points is less than the preset number, it is determined that there is no conflict in the spatial three-dimensional coordinate parameters of the target object at different angles.
[0109] Specifically, in this embodiment, the preset Euclidean distance deviation can be determined by the following method: calibrate the sensor accuracy: for binocular cameras, measure a calibration board with a known size in the workspace, and count the measurement error distribution to determine that the standard deviation is ±0.15mm; for depth cameras, measure a fixed point multiple times to calculate the standard deviation of the depth error, which is ±0.2mm; according to the error propagation law, the theoretical error after fusion of the two sensors is: A ≈0.25mm; determine the positioning accuracy of the mechanical arm: the repeatability of the robot is ±0.08mm (manufacturer's specification), and the positioning error in the working area in actual test is within ±0.1mm; considering the superposition of sensor error and mechanical arm error, the maximum allowable deviation is: B=A+mechanical arm error=0.25+0.1=0.35mm; measure the key point coordinates in 100 conflict-free scenarios, calculate the Euclidean distance deviation between the two sensors, and obtain the distribution: 95% of the deviation values are <0.3mm, and the maximum value is 0.32mm; in the test with a known 0.5mm artificial offset, the system successfully identifies all abnormal points, and the preset Euclidean distance deviation is set to 0.3mm based on the theoretical calculation and actual measurement data.
[0110] Specifically, in this embodiment, the number of key points directly affects the proportion threshold of abnormal points, and the preset number is usually set in proportion (such as 10%-20%) to avoid misjudgment due to differences in the total number of key points.
[0111] In the present application, by extracting the contour and calculating the curvature, the complexity of the shape of the target object can be converted into a quantitative indicator (such as curvature distribution, high curvature area proportion, etc.), avoiding subjective judgment errors. The present application does not depend on the specific type of the target object (such as regular geometric body, irregular object), and can distinguish simple shapes (such as cylinder) from complex shapes (such as gear, blade) through curvature distribution, providing a basis for subsequent classification, matching or path planning. Through hand-eye calibration and coordinate unification, the spatial alignment problem of multi-sensor (such as camera) or multi-angle measurement data is solved, avoiding positioning deviation caused by non-uniform coordinate system.
[0112] Specifically, the analysis of whether there is a local spatial three-dimensional coordinate distortion phenomenon of the target object includes:
[0113] dividing the target object into a plurality of local regions,
[0114] calculating the variance of the coordinates of a single local region,
[0115] calculating the average variance of the coordinates of each local region,
[0116] calculating the difference between the variance and the average variance of the coordinates of a single local region to obtain a distortion representation,
[0117] counting the number of local regions with distortion representation greater than or equal to a preset distortion representation to obtain a distortion region number,
[0118] if the distortion region number is greater than or equal to a preset distortion region number, it is determined that there is a local spatial three-dimensional coordinate distortion phenomenon of the target object;
[0119] if the distortion region number is less than the preset distortion region number, it is determined that there is no local spatial three-dimensional coordinate distortion phenomenon of the target object.
[0120] Specifically, in this embodiment, the number of preset distortion regions can be determined by the following method: set the target object as a cube with an edge length of 10 cm, divide it into 20 local regions, and obtain 100 three-dimensional coordinate points for each local region by shooting with a binocular camera; divide the six faces of the cube into four rectangular regions (for example, divide each face into a 2x2 grid), and the total number of local regions = 6x4 = 24; assuming that the grabbing task allows a maximum of 10% of the regions to have distortion, then the number of preset distortion regions = 4x10% = 2.4, rounded up to 3; if the number of statistical distortion regions is greater than or equal to 3, it is determined that there is a three-dimensional coordinate distortion phenomenon; calculate the variance of each region and the average variance: assuming that the coordinate variance of each local region is: [0.12, 0.08, 0.56, 0.23, 0.15, …, 0.31] (a total of 24 data); calculate the average variance: σ = 0.25, and the standard deviation: σ = 0.15; take the average variance + 1 times the standard deviation as the critical value of the distortion representation: 0.25 + 0.15 = 0.4; assuming that there are 5 regions with a variance greater than or equal to 0.4, then the number of preset distortion regions can be set to 5 (that is, the threshold is dynamically determined according to the actual distribution).
[0121] Specifically, the determination of the excluded interference coordinates or the correction of the shooting conditions based on the analysis result comprises:
[0122] If there is a conflict between the three-dimensional coordinate parameters of the target object at different angles and there is a local three-dimensional coordinate distortion phenomenon of the target object, it is determined to exclude the interference coordinates.
[0123] If there is only a conflict between the three-dimensional coordinate parameters of the target object at different angles or only a local three-dimensional coordinate distortion phenomenon of the target object, it is determined to correct the shooting conditions.
[0124] By dividing the target object into local regions and calculating the variance difference, the local distortion in the three-dimensional coordinates (such as local dense or sparse point clouds, local geometric structure deformation) can be accurately positioned. Compared with global analysis, it is more sensitive to subtle abnormalities. Through distortion detection and intelligent decision-making, the influence of noise, outliers and coordinate conflicts on the three-dimensional model is significantly reduced, and the reconstruction accuracy and stability are improved.
[0125] Specifically, the excluded interference coordinates comprise:
[0126] The abnormal key points and the local regions with a distortion representation greater than or equal to the preset distortion representation are removed.
[0127] Specifically, the correction of the shooting conditions comprises:
[0128] Increase the shooting angle to increase the three-dimensional coordinate parameters of the target object, wherein the increase of the shooting angle is positively correlated with the number of distortion regions.
[0129] Specifically, in this embodiment, increasing the shooting angle to increase the target object space three-dimensional coordinate parameters can be adjusted in the following manner: divided into 30 local areas according to the curvature distribution of the curved surface (12 high curvature areas and 18 low curvature areas); calculate the coordinate variance of each local area, the average variance is 0.35 mm2, and the standard deviation is 0.18 mm2; the distortion representation critical value is 0.35+0.18=0.53 mm2; the number of distortion areas is 8 (more than the preset number of 3, the preset number is set according to the proportion of 10%: 30*10%=3); select 20 key points, calculate the coordinate deviation of the binocular camera and the depth camera; the number of abnormal key points with a Euclidean distance deviation of ≥0.3 mm is 5 (more than the preset number of 4, the preset number is set according to the proportion of 20%: 20*20%=4); because there are coordinate conflicts and distortions (8 distortion areas+5 abnormal key points) at the same time, it is determined that the shooting conditions need to be corrected, and the shooting angle needs to be increased; the shooting angle increase amount = the number of distortion areas x the angle coefficient (k=0.5, which can be adjusted according to the task accuracy); the number of angles to be increased = 8*0.5=4.
[0130] The intelligent factory voice and visual cooperative control system provided by the embodiment comprises:
[0131] A voice instruction recognition module is configured to recognize a voice instruction based on a preset wake-up word.
[0132] An image acquisition module is connected to the voice instruction recognition module and configured to acquire an image of a target object.
[0133] A feature point recognition module is connected to the voice instruction recognition module and the image acquisition module and configured to determine feature points of the target object and determine a grasping difficulty tendency type of the target object based on the feature points.
[0134] An analysis module is connected to the image acquisition module and the feature point recognition module and configured to determine a single-angle target object space three-dimensional coordinate parameter or a multi-angle target object space three-dimensional coordinate parameter as a calculation sub-parameter of a target pose three-dimensional coordinate of a robot arm based on the grasping difficulty tendency type of the target object, analyze whether there is a conflict of the target object space three-dimensional coordinate parameters at different angles, analyze whether there is a local space three-dimensional coordinate distortion phenomenon of the target object, and determine an interference coordinate to be excluded or a shooting condition to be corrected based on an analysis result.
[0135] A calculation module is connected to the image acquisition module, the feature point recognition module, and the analysis module and configured to calculate a grasping pose of the robot arm based on the determined calculation sub-parameter of the target pose three-dimensional coordinate of the robot arm.
[0136] Specifically, the specific structure of the voice instruction recognition module and the image acquisition module is not limited, and only needs to be able to realize the corresponding functions.
[0137] Specifically, the specific structure of the feature point recognition module, the analysis module and the calculation module is not limited, which can be composed of a logic component including a field programmable processor, a computer and a microprocessor in the computer.
[0138] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will fall within the protection scope of the present application.
[0139] The above is only the preferred embodiment of the present application, and is not used to limit the present application; for those skilled in the art, the present application can have various changes and variations, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for voice and vision collaborative control in a smart factory, characterized in that, include: Step S1: Wake up the device according to the preset wake-up word, recognize the voice command through the preset command word, and capture the image of the target object when the voice recognition is successful; Step S2: Determine the feature points of the target object based on the target object image, and determine the grasping difficulty tendency type according to the ratio of the number of feature points located in the feature region to the total number of feature points; Step S3: Based on the determined grasping difficulty tendency type and the shape complexity of the target object, determine the calculation parameters for the grasping posture of the target object, including: If the target object has a low grasping difficulty tendency and a low shape complexity, then the three-dimensional coordinate parameters of the target object in single-angle space are used as the calculation sub-parameters of the three-dimensional coordinates of the robot arm target posture. If the target object has a high grasping difficulty tendency or a high degree of shape complexity, then the three-dimensional coordinate parameters of the target object in space from multiple angles are used as the calculation sub-parameters of the three-dimensional coordinates of the target posture of the robotic arm. The analysis is conducted to determine whether there are conflicts in the three-dimensional coordinate parameters of the target object in space at different angles, and whether there is a distortion of the local three-dimensional coordinates of the target object. Based on the analysis results, the interference coordinates are determined to be eliminated or the shooting conditions are corrected. The analysis examines whether there are conflicts in the three-dimensional coordinate parameters of the target object at different angles, including: Transform the coordinates to the same reference frame. Data alignment is achieved using hand-eye calibration. Select the key points of the target object. Calculate the coordinates of the same key point at each angle. Calculate the Euclidean distance deviation. Identify key points whose Euclidean distance deviation is greater than or equal to the preset Euclidean distance deviation to obtain abnormal key points. Count the number of critical anomalies. If the number of abnormal key points is greater than or equal to the preset number, it is determined that there is a conflict in the three-dimensional coordinate parameters of the target object space under different angles. If the number of abnormal key points is less than the preset number, it is determined that there is no conflict in the three-dimensional coordinate parameters of the target object space under different angles; The analysis of whether there is local spatial three-dimensional coordinate distortion of the target object includes: The target object is divided into several local regions. Calculate the variance of the coordinates of a single local region. Calculate the average variance of the coordinates of each local region. The difference between the variance and the mean variance of the coordinates of a single local region is calculated to obtain a distortion characterization. The number of local regions whose distortion characteristics are greater than or equal to the preset distortion characteristics is counted to obtain the number of distorted regions. If the number of distorted regions is greater than or equal to the preset number of distorted regions, it is determined that there is a local spatial three-dimensional coordinate distortion phenomenon of the target object; If the number of distorted regions is less than the preset number of distorted regions, it is determined that there is no local spatial three-dimensional coordinate distortion phenomenon of the target object; The process of determining interference-eliminating coordinates or correcting shooting conditions based on analysis results includes: If there are conflicts in the three-dimensional coordinate parameters of the target object at different angles and there is a distortion of the three-dimensional coordinates of the target object in a local space, then the interfering coordinates are excluded. If there is only a conflict in the three-dimensional coordinate parameters of the target object at different angles or only a distortion in the three-dimensional coordinates of the target object in a local space, then the shooting conditions should be corrected. Step S4: Calculate the sub-parameters according to the determined grasping posture to determine the grasping pose of the robotic arm and grasp the target object.
2. The intelligent factory voice and vision collaborative control method according to claim 1, characterized in that, In step S2, the grasping difficulty tendency type is determined based on the ratio of the number of feature points located in the feature region to the total number of feature points, including: The ratio of the number of feature points located in the feature region to the total number of feature points is calculated to obtain a characterization of the crawling difficulty. If the grasping difficulty indicator is less than or equal to the preset grasping difficulty indicator, then the grasping difficulty tendency type of the target object is determined to be strong grasping difficulty tendency; If the grasping difficulty indicator is greater than the preset grasping difficulty indicator, then the grasping difficulty tendency type of the target object is determined to be weak grasping difficulty tendency.
3. The intelligent factory voice and vision collaborative control method according to claim 1, characterized in that, In step S3, the contour features of the target object are extracted based on the target object image, and the contour curvature is calculated to determine the shape complexity of the target object based on the contour curvature.
4. The intelligent factory voice and vision collaborative control method according to claim 1, characterized in that, The interference-eliminating coordinates include: Remove the abnormal key points and local areas where the distortion characterization is greater than or equal to the preset distortion characterization.
5. The intelligent factory voice and vision collaborative control method according to claim 4, characterized in that, The corrected shooting conditions include: Increasing the shooting angle increases the three-dimensional coordinate parameters of the target object, wherein the increase in shooting angle is positively correlated with the number of distorted regions.
6. A system applying the intelligent factory voice and vision collaborative control method according to any one of claims 1-5, characterized in that, include: The voice command recognition module is used to recognize wake-up commands based on preset wake-up words and to recognize voice commands based on preset command words. An image acquisition module, which is connected to the voice command recognition module, is used to acquire images of the target object; The feature point recognition module is connected to the voice command recognition module and the image acquisition module respectively, and is used to determine the feature points of the target object and determine the type of difficulty in grasping the target object based on the feature points; The analysis module, which is connected to the image acquisition module and the feature point recognition module respectively, is used to determine whether to use the three-dimensional coordinate parameters of the target object in a single angle or the three-dimensional coordinate parameters of the target object in a multi-angle angle as the calculation sub-parameters of the robot arm target posture three-dimensional coordinates based on the difficulty tendency type of the target object's grasping, analyze whether there is a conflict of the three-dimensional coordinate parameters of the target object in a different angle, analyze whether there is a local three-dimensional coordinate distortion phenomenon of the target object, and determine the coordinates to eliminate interference or the shooting conditions to be corrected based on the analysis results. The calculation module, which is connected to the image acquisition module, the feature point recognition module and the analysis module, is used to calculate the grasping pose of the robotic arm based on the calculation sub-parameters of the determined three-dimensional coordinates of the robotic arm target posture.
Citation Information
Patent Citations
Intelligent factory production line visual inspection system and method
CN119023577A
Unknown object six-degree-of-freedom grabbing method considering point cloud skeleton characteristics
CN116460845A
Mobile visual detection method and system and storage medium
CN118190953A