A humanoid robot target recognition and positioning method and system

Through technical means such as dual-view monocular visual positioning, directional gradient feature extraction, random sampling filtering algorithm and gait detection, the problem of low visual positioning accuracy in complex environments is solved, high-precision target recognition and positioning is achieved, and the stability and anti-interference ability of the system are enhanced.

CN118990480BActive Publication Date: 2025-05-09BEIJING ZHONGLIAN GUOCHENG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411143152.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2025-05-09
Estimated Expiration
2044-08-20

Smart Images

  • Figure CN118990480B_ABST
    Figure CN118990480B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for target recognition and positioning of a humanoid robot. The method comprises: using a dual-view monocular vision positioning method to calculate the three-dimensional position of a target to be identified; using image feature extraction based on directional gradients and a support vector machine classifier to preliminarily detect and classify the identified target; using a random sampling filtering algorithm to track the identified target; using a position algorithm based on gait detection to position the humanoid robot and obtain the motion trajectory of the humanoid robot; and correcting the motion trajectory of the humanoid robot through a sensor fusion algorithm. The system comprises a target positioning module, a target classification module, a target tracking module, a positioning module and a positioning correction module. The present invention can complete target recognition and positioning as well as the positioning of the humanoid robot itself, help the humanoid robot to understand and identify the environment, obtain the relative position of the humanoid robot and the target object, and facilitate subsequent work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of humanoid robots, and in particular to a target recognition and positioning method and system for a humanoid robot. Background Art

[0002] In the daily cognitive process of human beings, the visual system occupies a vital position. For robots, visual sensors also play an equally important role in the robot's perception of the external environment. With the development of robotics technology, how to make it more intelligent and sensitive is the core topic of robotic vision technology. Robotic vision technology analyzes and processes the information collected by visual sensors to achieve the recognition, tracking and positioning of target objects. However, relying solely on visual information has limitations. For example, in complex or adverse environmental conditions, the recognition effect of the visual system will be affected.

[0003] Existing target recognition methods mainly include the following categories: segmentation-based recognition methods, which use image segmentation technology to separate the target object from the background for subsequent processing; learning-based recognition methods, which use machine learning algorithms such as support vector machines (SVMs) and neural networks to train and classify target features; knowledge-based recognition methods, which combine expert knowledge bases to recognize and classify targets; model-based recognition methods, which match the features of the target object with the features in the actual scene by establishing a mathematical model of the target object; and information fusion-based recognition methods, which fuse information from multiple sensors to improve the accuracy and robustness of recognition. In complex environments, a single method often cannot meet the needs of accurate recognition, so recognition methods based on information fusion have become a trend.

[0004] In the field of mobile robots, positioning technology is the basis of robot autonomous navigation. According to the different sensors used, robot positioning technology can be divided into two types: traditional positioning and visual positioning. The traditional positioning method relies on a variety of sensors, such as electronic compasses, IMU units, GPS, ultrasonic sensors, etc. These sensors provide position and posture information. By integrating this information, the robot can achieve positioning and navigation in the environment. The visual positioning method mainly relies on the information provided by visual sensors (such as CMOS cameras, CCD cameras), extracts environmental features through image processing technology, and then realizes target recognition and positioning.

[0005] According to the different positioning methods, traditional positioning and visual positioning can be divided into incremental positioning and global positioning: Incremental positioning is based on the continuous update of the robot's current state, using the position and speed information of the previous step to infer the current position. This method is easy to implement, but the cumulative error is large; global positioning is based on known feature points in the environment and directly calculates the absolute position of the robot. This method has high accuracy but relies on known information in the environment.

[0006] Although the above positioning and recognition technologies have been widely studied and applied in robot applications, they still have the following shortcomings: limitations of visual information. In the case of insufficient lighting, occlusion, complex background, etc., recognition and positioning that rely solely on visual sensors are prone to failure, resulting in reduced or failed recognition accuracy; the complexity of sensor fusion. In the existing technology, although multi-sensor fusion methods have been adopted, how to effectively fuse information from different sensors is still a challenge, especially when facing the heterogeneity and real-time requirements of sensor data; limitations of positioning accuracy. Although the traditional incremental positioning method is simple and easy, the problem of cumulative error is prominent and it is difficult to meet the needs of precise positioning. Although global positioning has high accuracy, it is highly dependent on the environment. Summary of the invention

[0007] In view of this, the purpose of an embodiment of the present invention is to provide a target recognition and positioning method and system for a humanoid robot, which can complete target recognition and positioning as well as the humanoid robot's own positioning, help the humanoid robot to understand and identify the environment, obtain the relative position of the humanoid robot and the target object, and facilitate subsequent work.

[0008] The embodiment of the present invention is achieved as follows:

[0009] A target recognition and positioning method for a humanoid robot, comprising:

[0010] Use the dual-view monocular vision positioning method to calculate the three-dimensional position of the target to be identified.

[0011] Direction gradient-based image feature extraction and support vector machine classifier are used to perform preliminary detection and classification of identified targets.

[0012] The identified target is tracked using a random sampling filtering algorithm.

[0013] A position algorithm based on gait detection is used to locate the humanoid robot and obtain the motion trajectory of the humanoid robot.

[0014] The motion trajectory of the humanoid robot is corrected through a sensor fusion algorithm.

[0015] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of a humanoid robot, the use of a dual-view monocular vision positioning method to calculate the three-dimensional position of the target to be recognized includes:

[0016] The humanoid robot is equipped with two monocular cameras which are respectively installed at different heights, and the distance between the two monocular cameras is d.

[0017] The pixel coordinates of the target to be identified on the image plane of the first monocular camera are: , the pixel coordinates on the image plane of the second monocular camera are .

[0018] For the first monocular camera, establish the first plane pixel coordinates of the target to be identified The geometric relationship with the actual three-dimensional coordinates (X, Y, Z), ,in, is the focal length of the first monocular camera, and D is the depth information of the target to be identified.

[0019] For the second monocular camera, establish the second plane pixel coordinates of the target to be identified The geometric relationship with the actual three-dimensional coordinates (X, Y, Z), ,in, is the focal length of the second monocular camera.

[0020] Calculate the vertical disparity of the target to be identified in the two plane images of the monocular cameras .

[0021] Calculate the depth information of the target to be identified .

[0022] The three-dimensional position of the target to be identified is .

[0023] Its technical effect is: using two monocular cameras, calculating the target depth through parallax information to achieve real-time 3D reconstruction. The relative position of the two monocular cameras is a known baseline, and the 3D coordinates of the target can be directly calculated through simple geometric relationships. The algorithm is concise and computationally efficient, the required hardware settings are simple, the cost is low, and it is suitable for large-scale practical applications. Compared with traditional binocular stereo vision, only one baseline is required to obtain depth information, which simplifies the system design and is easy to implement. Since the relative position between the two cameras is fixed, the algorithm model is stable, the positioning error is small, and the accuracy of the system is guaranteed.

[0024] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of the humanoid robot, the use of image feature extraction based on directional gradient and support vector machine classifier to perform preliminary detection and classification of the recognition target includes:

[0025] The color image containing the recognition target is grayed and normalized to obtain an input image.

[0026] A directional gradient feature extraction model is established, the input image is used as the input of the directional gradient feature extraction model, and a directional gradient feature vector is obtained as output.

[0027] Images containing recognition targets are collected as samples, and corresponding category labels are respectively annotated for each of the samples to obtain a data set.

[0028] A support vector machine classifier is designed, the support vector machine classifier is trained using the data set, the directional gradient feature vector is used as the input of the trained support vector machine classifier, a category label is output, and a classification result of the identified target is obtained.

[0029] The technical effect is that the present invention can fully consider the image edge direction information through directional gradient feature extraction, effectively extract target features, significantly improve the accuracy of classification and recognition, and effectively optimize the classification effect by processing high-dimensional feature space through the kernel technique of SVM classifier. The organic combination of image features and classifiers enables the system to intelligently classify and recognize target shape features, with high recognition accuracy and strong credibility, providing reliable input for subsequent target tracking and positioning, which is concise, efficient, easy to use, and can be expanded to more categories of target recognition, and is suitable for robot vision systems with low computing power.

[0030] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of the humanoid robot, the step of establishing a directional gradient feature extraction model, taking the input image as the input of the directional gradient feature extraction model, and outputting a directional gradient feature vector comprises:

[0031] A directional gradient feature extraction model is established, wherein the directional gradient feature extraction model includes an input layer, a gradient calculation layer, a directional gradient histogram layer, a feature vector splicing layer and a normalization layer.

[0032] The input layer receives the input image after grayscale processing and normalization processing, wherein each pixel in the input image is represented as .

[0033] The gradient calculation layer calculates the horizontal gradient in the x direction for each pixel in the input image. , the vertical gradient in the y direction , the gradient amplitude representing the edge strength is calculated , represents the gradient direction of the edge direction .

[0034] The directional gradient histogram layer divides the input image into a number of N×N cell units, where N is the pixel value. In each cell unit, based on the calculated gradient direction, the gradient amplitude is added to the corresponding direction interval b. For each direction interval b, the gradient amplitude belonging to the direction is calculated and accumulated. , get the histogram vector of the gradient direction of the cell unit .

[0035] The feature vector concatenation layer concatenates the histogram vectors of the gradient direction of each adjacent cell unit in sequence to obtain a block feature vector v.

[0036] The normalization layer performs normalization processing on the spliced ​​block feature vectors. , get the directional gradient feature vector, where Is a positive constant.

[0037] The technical effect is that the present invention can directly extract edge information in the image, including edge strength and direction, through the gradient calculation layer and the directional gradient histogram layer, effectively capture the key structural information of the target, divide the image into several cell units, and construct a histogram of the gradient direction in each unit. By accumulating these histograms, it can effectively resist the influence of noise and illumination changes on the recognition effect. The directional gradient feature extraction model of the present invention provides reliable technical support for the target recognition and positioning of humanoid robots by efficiently and accurately extracting the edge and structural information of the image, combined with powerful classification capabilities, and at the same time has good real-time and robustness, and can achieve stable target recognition and tracking in complex environments.

[0038] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of a humanoid robot, the support vector machine classifier is designed, the support vector machine classifier is trained using the data set, the directional gradient feature vector is used as the input of the trained support vector machine classifier, and the category label is output, and the classification result of the recognized target is obtained, which includes:

[0039] Construct a multi-class support vector machine classifier and train one support vector machine classifier for each pair of categories, requiring a total of J(J−1) / 2 classifiers, where J is the number of categories of the recognition target.

[0040] Train each type of support vector machine classifier and adjust the loss function through the optimization algorithm , where F is the penalty term, e is the weight vector, q is the bias term, and M is the number of samples. The weight vector e and the bias term q are adjusted to minimize the loss function to a preset value.

[0041] The directional gradient feature vector extracted from the input image is used as input, and is input into each type of the trained support vector machine classifier, and each type of the support vector machine classifier outputs a category label, and the classification result of the recognition target of the input image is obtained. .

[0042] Its technical effect is: by introducing a multi-class classification strategy and combining the directional gradient eigenvector features, it can effectively identify and classify multiple targets. This method has higher accuracy when dealing with multi-category object recognition tasks and is particularly suitable for target classification tasks in complex scenarios.

[0043] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of a humanoid robot, the use of a random sampling filter algorithm to track the recognized target includes:

[0044] The initial position of the identified target is , an initial search window is defined at the initial position of the identified target, and the radius of the initial search window is .

[0045] For each frame in the input image set, calculate the distribution of the color histogram in the three-dimensional color space ,in, is a color value in the color space, is the pixel color value in the target area, is the indicator function, the color difference satisfies the condition The value of the indicator function is 1 when , otherwise it is 0. is the color difference threshold.

[0046] Update the center position of the identified target ,in, , , is the displacement of the identified target in the current frame, , , .

[0047] Update the radius of the search window ,in, The area or volume of the target in the current frame.

[0048] Update the position of the identified target in each frame of the input image set to obtain the three-dimensional trajectory of the identified target during the tracking process .

[0049] The technical effects are as follows: the random sampling filtering algorithm of the present invention can quickly calculate and update the target position in each frame of the input image, is suitable for high frame rate video stream processing, and ensures the real-time performance of target tracking; by continuously updating the radius and position of the search window, the algorithm can adapt to the scale change and displacement of the target, so that even if the target changes during the movement, it can continue to track stably; using the color histogram distribution in the three-dimensional color space as the target feature, it can effectively cope with the appearance change of the target and the interference of the complex background; the algorithm not only tracks the two-dimensional plane position of the target, but also can combine the position information in the three-dimensional space to output the three-dimensional trajectory of the target, providing accurate spatial position information for the navigation and operation of the humanoid robot in a complex environment. The random sampling filtering algorithm of the present invention is simple to calculate and has low resource consumption, is suitable for embedded systems or resource-constrained application scenarios, and reduces the dependence on hardware resources while ensuring the tracking accuracy.

[0050] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of a humanoid robot, the positioning of the humanoid robot using a position algorithm based on gait detection to obtain the motion trajectory of the humanoid robot includes:

[0051] The humanoid robot is walking, and the current number of steps is , wherein the joint angle of the left leg of the humanoid robot during walking is , the joint angle of the right leg during walking is , is an index function for gait detection. When a complete gait cycle is detected, the index function value for gait detection is 1.

[0052] Combined with the IMU angle data, the current three-dimensional coordinate position of the humanoid robot is obtained. ,in, is the angular change of the humanoid robot's trunk around the Z axis, and s is the step length of each gait cycle of the humanoid robot.

[0053] Its technical effects are: by combining random sampling filtering with three-dimensional position modeling, it overcomes the limitation of traditional gait detection position algorithm that can only track targets in a two-dimensional plane, so that the algorithm can be applied in more complex three-dimensional environments; by fusing multiple sensor information for trajectory correction, it effectively reduces the errors caused by gait detection and IMU detection, and enhances the robustness of the system in dynamic environments; the position algorithm based on gait detection is combined with visual information, so that the robot can not only autonomously locate, but also maintain accurate trajectory tracking in complex terrain.

[0054] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of the humanoid robot, the correction of the motion trajectory of the humanoid robot by the sensor fusion algorithm comprises:

[0055] According to the current position of the humanoid robot and the speed at the current time t , predicting the humanoid robot at the next moment location, , , ,in, is the time interval.

[0056] The sensor fusion measurement value is collected, and the X-axis position is , the Y-axis position is , the Z-axis position is .

[0057] For the predicted position of the humanoid robot at the next moment, update the covariance matrix ,in, is the state covariance matrix at the current moment, is the process noise covariance matrix.

[0058] Calculate Kalman gain ,in, is the measurement noise covariance matrix.

[0059] Correcting the predicted position of the humanoid robot, , , .

[0060] The technical effects are as follows: the present invention effectively reduces the influence of single sensor error by fusing the measurement values ​​of multiple sensors and combining with the Kalman filter to correct the motion trajectory in real time, thereby improving the positioning accuracy of the robot; by calculating and updating the covariance matrix and Kalman gain, it can continuously optimize error control in real-time processing, reduce system uncertainty, and improve the accuracy and efficiency of trajectory prediction and correction; it effectively integrates the information provided by different sensors (such as IMU, visual sensor, etc.), fully utilizes the advantages of multi-source data, improves the reliability and anti-interference ability of the system, and significantly improves the motion trajectory accuracy of the humanoid robot and the stability of the system through the integration, real-time prediction and correction of multi-sensor data.

[0061] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of a humanoid robot, the method for obtaining the sensor fusion measurement value includes:

[0062] Get the acceleration of the IMU measurement unit of the humanoid robot at time t , the speed of the humanoid robot is calculated to be , , , the position of the humanoid robot is calculated as , , , obtain the IMU measurement unit measurement value.

[0063] Get the distance measured by the ultrasonic sensor of the humanoid robot at time t ,in, is the propagation speed of sound waves, is the time from the emission to the reception of the sound wave. The ultrasonic sensor is located in front of the humanoid robot and is With fixed offset , the position of the humanoid robot is calculated as , , ,in, is the pitch angle of sound wave propagation, is the azimuth of sound wave propagation, and is the ultrasonic sensor measurement value.

[0064] The IMU measurement unit measurement value and the ultrasonic sensor measurement value are fused to obtain the sensor fusion measurement value. The X-axis position is , the Y-axis position is , the Z-axis position is ,in, and is the weight coefficient.

[0065] Its technical effects are as follows: IMU provides dynamic information of acceleration and velocity, which is suitable for fast motion tracking in a short period of time, while ultrasonic sensors provide distance measurements, which are suitable for accurate distance measurement of the environment in static or slow motion; by simultaneously utilizing the data of the IMU measurement unit and the ultrasonic sensor, multi-dimensional information such as acceleration, velocity, position, and distance are integrated. The measurement errors and noises of the IMU and ultrasonic sensors can be effectively suppressed through the fusion algorithm, reducing the impact of the error of a single sensor on the final positioning result, and being able to more accurately correct and predict the position of the robot, reduce positioning errors, and improve the accuracy of trajectory tracking.

[0066] A target recognition and positioning system for a humanoid robot, comprising:

[0067] The target positioning module is used to calculate the three-dimensional position of the target to be identified using a dual-view monocular vision positioning method.

[0068] The target classification module is used to perform preliminary detection and classification of the identified targets using image feature extraction based on directional gradient and support vector machine classifier.

[0069] The target tracking module is used to track the identified target using a random sampling filtering algorithm.

[0070] The positioning module is used to use a position algorithm based on gait detection to position the humanoid robot and obtain the motion trajectory of the humanoid robot.

[0071] The positioning correction module is used to correct the motion trajectory of the humanoid robot through a sensor fusion algorithm.

[0072] The beneficial effects of the embodiments of the present invention are:

[0073] The present invention uses a dual-view monocular vision method to calculate the three-dimensional position of the target. By setting two cameras at different heights and calculating their parallax, real-time three-dimensional reconstruction can be achieved. Compared with traditional binocular stereo vision, the hardware configuration is simpler and the cost is low, while maintaining high computing efficiency and accuracy. Through the known baseline distance and geometric relationship, the calculation process of the three-dimensional coordinates is simplified, the system complexity is reduced and the stability is improved.

[0074] The present invention adopts a directional gradient feature extraction model to effectively extract edge and structural information in images, and combines it with a support vector machine classifier to perform target detection and classification, making full use of the edge information of the image and improving the classification accuracy through the optimization of high-dimensional feature space. It can achieve accurate target recognition in complex environments and provide a reliable foundation for subsequent target tracking and positioning.

[0075] The present invention tracks the target through a random sampling filtering algorithm (such as particle filtering), can update the position of the target in real time in each frame of the input image, and adapt to the scale change and displacement of the target. It performs feature matching through the color histogram distribution in the three-dimensional color space, can effectively cope with the appearance change of the target and background interference, improves the robustness and accuracy of tracking, is suitable for high frame rate video stream processing, and has simple calculation and low resource consumption.

[0076] The present invention combines gait detection algorithm with IMU data to realize the positioning of humanoid robots in three-dimensional environments, overcoming the limitation that traditional gait detection can only track in two-dimensional planes. Combining IMU data enhances the robustness of the system in dynamic environments. By fusing gait detection with IMU angle data, the positioning accuracy and trajectory tracking capability of the robot in complex terrain are improved.

[0077] The present invention corrects the robot's motion trajectory through a sensor fusion algorithm (such as Kalman filtering), which can effectively reduce the error of a single sensor and improve positioning accuracy. It fuses the data of IMU and ultrasonic sensors, integrates acceleration, speed, position and distance information, and optimizes error control through real-time correction, thereby enhancing the system's reliability and anti-interference ability, and improving the system's positioning accuracy and stability.

[0078] The present invention uses advanced target recognition, tracking and positioning technologies to ultimately achieve high-precision target recognition and positioning, allowing the robot to better understand and interact with its surrounding environment, maintain stable and accurate behavior in real-time applications, support the robot's effective operation in a dynamic environment through real-time tracking capabilities, enhance the robot's autonomous navigation and operation capabilities in complex environments, and interact with targets in the environment more accurately, improving the user experience in scenarios that require high-precision operations and feedback. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0080] Figure 1 The present invention is a flowchart of the target recognition and positioning method of the humanoid robot. DETAILED DESCRIPTION

[0081] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0082] Please refer to Figure 1 The first embodiment of the present invention provides a target recognition and positioning method for a humanoid robot, comprising: using a dual-view monocular vision positioning method to calculate the three-dimensional position of a target to be identified; using image feature extraction based on directional gradients and a support vector machine classifier to perform preliminary detection and classification on the identified target; using a random sampling filtering algorithm to track the identified target; using a position algorithm based on gait detection to locate the humanoid robot and obtain the motion trajectory of the humanoid robot; and correcting the motion trajectory of the humanoid robot through a sensor fusion algorithm.

[0083] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of the humanoid robot, the use of the dual-view monocular vision positioning method to calculate the three-dimensional position of the target to be identified includes: equipping the humanoid robot with two monocular cameras, which are respectively installed at different heights, and the distance between the two monocular cameras is d; the pixel coordinates of the target to be identified on the image plane of the first monocular camera are , the pixel coordinates on the image plane of the second monocular camera are ; For the first monocular camera, establish the first plane pixel coordinates of the target to be identified The geometric relationship with the actual three-dimensional coordinates (X, Y, Z), ,in, is the focal length of the first monocular camera, D is the depth information of the target to be identified; for the second monocular camera, establish the second plane pixel coordinates of the target to be identified The geometric relationship with the actual three-dimensional coordinates (X, Y, Z), ,in, is the focal length of the second monocular camera; the vertical disparity of the target to be identified in the two plane images of the monocular cameras is calculated ; Calculate the depth information of the target to be identified ; The three-dimensional position of the target to be identified is .

[0084] Its technical effect is: using two monocular cameras, calculating the target depth through parallax information to achieve real-time 3D reconstruction. The relative position of the two monocular cameras is a known baseline, and the 3D coordinates of the target can be directly calculated through simple geometric relationships. The algorithm is concise and computationally efficient, the required hardware settings are simple, the cost is low, and it is suitable for large-scale practical applications. Compared with traditional binocular stereo vision, only one baseline is required to obtain depth information, which simplifies the system design and is easy to implement. Since the relative position between the two cameras is fixed, the algorithm model is stable, the positioning error is small, and the accuracy of the system is guaranteed.

[0085] In a preferred embodiment of the present invention, in the target recognition and positioning method of the above-mentioned humanoid robot, the use of image feature extraction based on directional gradient and support vector machine classifier to perform preliminary detection and classification of the recognition target includes: graying and normalizing the color image containing the recognition target to obtain an input image; establishing a directional gradient feature extraction model, using the input image as the input of the directional gradient feature extraction model, and outputting a directional gradient feature vector; collecting images containing the recognition target as samples, and labeling each of the samples with a corresponding category label to obtain a data set; designing a support vector machine classifier, using the data set to train the support vector machine classifier, using the directional gradient feature vector as the input of the trained support vector machine classifier, outputting a category label, and obtaining a classification result of the recognition target.

[0086] The technical effect is that the present invention can fully consider the image edge direction information through directional gradient feature extraction, effectively extract target features, significantly improve the accuracy of classification and recognition, and effectively optimize the classification effect by processing high-dimensional feature space through the kernel technique of SVM classifier. The organic combination of image features and classifiers enables the system to intelligently classify and recognize target shape features, with high recognition accuracy and strong credibility, providing reliable input for subsequent target tracking and positioning, which is concise, efficient, easy to use, and can be expanded to more categories of target recognition, and is suitable for robot vision systems with low computing power.

[0087] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot target recognition and positioning method, the establishment of a directional gradient feature extraction model, taking the input image as the input of the directional gradient feature extraction model, and outputting a directional gradient feature vector includes: establishing a directional gradient feature extraction model, the directional gradient feature extraction model includes an input layer, a gradient calculation layer, a directional gradient histogram layer, a feature vector splicing layer and a normalization layer; the input layer receives the input image after grayscale processing and normalization processing, wherein each pixel in the input image is represented as ; The gradient calculation layer calculates the horizontal gradient in the x direction for each pixel in the input image , the vertical gradient in the y direction , the gradient amplitude representing the edge strength is calculated , represents the gradient direction of the edge direction ; The directional gradient histogram layer divides the input image into a number of N×N cell units, where N is the pixel value. In each cell unit, based on the calculated gradient direction, the gradient amplitude is added to the corresponding direction interval b. Assuming that the direction is divided into 9 intervals, namely [0∘, 20∘], [20∘, 40∘], …, [160∘, 180∘], for each direction interval b, the gradient amplitude belonging to the direction is calculated and accumulated. , get the histogram vector of the gradient direction of the cell unit ; The feature vector concatenation layer concatenates the histogram vectors of the gradient direction of each adjacent cell unit in sequence to obtain a block feature vector v. For a cell unit of size 2×2, each cell unit includes 9 square intervals, and the block feature vector is 36-dimensional; the normalization layer normalizes the concatenated block feature vectors. , get the directional gradient feature vector, where Is a positive constant.

[0088] The technical effect is that the present invention can directly extract edge information in the image, including edge strength and direction, through the gradient calculation layer and the directional gradient histogram layer, effectively capture the key structural information of the target, divide the image into several cell units, and construct a histogram of the gradient direction in each unit. By accumulating these histograms, it can effectively resist the influence of noise and illumination changes on the recognition effect. The directional gradient feature extraction model of the present invention provides reliable technical support for the target recognition and positioning of humanoid robots by efficiently and accurately extracting the edge and structural information of the image, combined with powerful classification capabilities, and at the same time has good real-time and robustness, and can achieve stable target recognition and tracking in complex environments.

[0089] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of the humanoid robot, the design of the support vector machine classifier, the use of the data set to train the support vector machine classifier, the use of the directional gradient feature vector as the input of the trained support vector machine classifier, the output of the category label, and the classification result of the identified target include: constructing a multi-class support vector machine classifier, training a support vector machine classifier for each pair of categories, requiring a total of J(J−1) / 2 classifiers, where J is the number of categories of the identified target; training each class of the support vector machine classifier, and adjusting the loss function through an optimization algorithm , where F is a penalty term, e is a weight vector, q is a bias term, and M is the number of samples. The weight vector e and the bias term q are adjusted to minimize the loss function to a preset value; the directional gradient feature vector extracted from the input image is used as input, and is input into each type of the trained support vector machine classifier, and each type of the support vector machine classifier outputs a category label, and the classification result of the recognition target of the input image is obtained. .

[0090] Its technical effect is: by introducing a multi-class classification strategy and combining the directional gradient eigenvector features, it can effectively identify and classify multiple targets. This method has higher accuracy when dealing with multi-category object recognition tasks and is particularly suitable for target classification tasks in complex scenarios.

[0091] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of a humanoid robot, the use of a random sampling filter algorithm to track the identified target includes: the initial position of the identified target is , an initial search window is defined at the initial position of the identified target, and the radius of the initial search window is ; For each frame in the input image set, calculate the distribution of the color histogram in the three-dimensional color space ,in, is a color value in the color space, is the pixel color value in the target area, is the indicator function, the color difference satisfies the condition When the value of the indicator function is 1, it means that the color value of this pixel belongs to the corresponding color interval of the color histogram; otherwise, it is 0, which means that the color value of this pixel does not belong to this interval. is the color difference threshold; updates the center position of the identified target ,in, , , is the displacement of the identified target in the current frame, , , ; Update the radius of the search window ,in, is the area or volume of the target in the current frame; updates the position of the identified target in each frame of the input image set to obtain the three-dimensional trajectory of the identified target during the tracking process .

[0092] The technical effects are as follows: the random sampling filtering algorithm of the present invention can quickly calculate and update the target position in each frame of the input image, is suitable for high frame rate video stream processing, and ensures the real-time performance of target tracking; by continuously updating the radius and position of the search window, the algorithm can adapt to the scale change and displacement of the target, so that even if the target changes during the movement, it can continue to track stably; using the color histogram distribution in the three-dimensional color space as the target feature, it can effectively cope with the appearance change of the target and the interference of the complex background; the algorithm not only tracks the two-dimensional plane position of the target, but also can combine the position information in the three-dimensional space to output the three-dimensional trajectory of the target, providing accurate spatial position information for the navigation and operation of the humanoid robot in a complex environment. The random sampling filtering algorithm of the present invention is simple to calculate and has low resource consumption, is suitable for embedded systems or resource-constrained application scenarios, and reduces the dependence on hardware resources while ensuring the tracking accuracy.

[0093] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot target recognition and positioning method, the positioning of the humanoid robot using a position algorithm based on gait detection and obtaining the motion trajectory of the humanoid robot includes: during the walking process of the humanoid robot, the current number of walking steps is , wherein the joint angle of the left leg of the humanoid robot during walking is , the joint angle of the right leg during walking is , is the index function of gait detection. When a complete gait cycle is detected, the index function value of gait detection is 1. Combined with the IMU angle data, the current three-dimensional coordinate position of the humanoid robot is obtained. ,in, is the angular change of the humanoid robot's trunk around the Z axis, and s is the step length of each gait cycle of the humanoid robot.

[0094] Its technical effects are: by combining random sampling filtering with three-dimensional position modeling, it overcomes the limitation of traditional gait detection position algorithm that can only track targets in a two-dimensional plane, so that the algorithm can be applied in more complex three-dimensional environments; by fusing multiple sensor information for trajectory correction, it effectively reduces the errors caused by gait detection and IMU detection, and enhances the robustness of the system in dynamic environments; the position algorithm based on gait detection is combined with visual information, so that the robot can not only autonomously locate, but also maintain accurate trajectory tracking in complex terrain.

[0095] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot target recognition and positioning method, the correction of the motion trajectory of the humanoid robot by the sensor fusion algorithm comprises: and the speed at the current time t , predicting the humanoid robot at the next moment location, , , ,in, is the time interval; the sensor fusion measurement value is collected, and the X-axis position is , the Y-axis position is , the Z-axis position is ; For the predicted position of the humanoid robot at the next moment, update the covariance matrix ,in, is the state covariance matrix at the current moment, is the process noise covariance matrix; calculate the Kalman gain ,in, is the measurement noise covariance matrix; the predicted position of the humanoid robot is corrected, , , .

[0096] The technical effects are as follows: the present invention effectively reduces the influence of single sensor error by fusing the measurement values ​​of multiple sensors and combining with the Kalman filter to correct the motion trajectory in real time, thereby improving the positioning accuracy of the robot; by calculating and updating the covariance matrix and Kalman gain, it can continuously optimize error control in real-time processing, reduce system uncertainty, and improve the accuracy and efficiency of trajectory prediction and correction; it effectively integrates the information provided by different sensors (such as IMU, visual sensor, etc.), fully utilizes the advantages of multi-source data, improves the reliability and anti-interference ability of the system, and significantly improves the motion trajectory accuracy of the humanoid robot and the stability of the system through the integration, real-time prediction and correction of multi-sensor data.

[0097] In a preferred embodiment of the present invention, in the above-mentioned target recognition and positioning method of the humanoid robot, the method for obtaining the sensor fusion measurement value includes: obtaining the acceleration of the IMU measurement unit of the humanoid robot at time t , the speed of the humanoid robot is calculated to be , , , the position of the humanoid robot is calculated as , , , obtain the IMU measurement unit measurement value; obtain the distance measured by the ultrasonic sensor of the humanoid robot at time t ,in, is the propagation speed of sound waves, is the time from the emission to the reception of the sound wave. The ultrasonic sensor is located in front of the humanoid robot and is With fixed offset , the position of the humanoid robot is calculated as , , ,in, is the pitch angle of sound wave propagation, is the azimuth of sound wave propagation, and obtains the ultrasonic sensor measurement value; fuses the IMU measurement unit measurement value and the ultrasonic sensor measurement value to obtain the sensor fusion measurement value, and the X-axis position is , the Y-axis position is , the Z-axis position is ,in, and is the weight coefficient.

[0098] Its technical effects are as follows: IMU provides dynamic information of acceleration and velocity, which is suitable for fast motion tracking in a short period of time, while ultrasonic sensors provide distance measurements, which are suitable for accurate distance measurement of the environment in static or slow motion; by simultaneously utilizing the data of the IMU measurement unit and the ultrasonic sensor, multi-dimensional information such as acceleration, velocity, position, and distance are integrated. The measurement errors and noises of the IMU and ultrasonic sensors can be effectively suppressed through the fusion algorithm, reducing the impact of the error of a single sensor on the final positioning result, and being able to more accurately correct and predict the position of the robot, reduce positioning errors, and improve the accuracy of trajectory tracking.

[0099] A target recognition and positioning system for a humanoid robot, comprising: a target positioning module, used to calculate the three-dimensional position of a target to be recognized using a dual-view monocular vision positioning method; a target classification module, used to perform preliminary detection and classification of the recognized target using image feature extraction based on directional gradients and a support vector machine classifier; a target tracking module, used to track the recognized target using a random sampling filter algorithm; a positioning module, used to perform positioning of the humanoid robot using a position algorithm based on gait detection to obtain the motion trajectory of the humanoid robot; and a positioning correction module, used to correct the motion trajectory of the humanoid robot using a sensor fusion algorithm.

[0100] The second embodiment of the present invention provides a target recognition and positioning system for a humanoid robot, which includes: a target positioning module, which is used to use a dual-view monocular vision positioning method to calculate the three-dimensional position of a target to be identified; a target classification module, which is used to use image feature extraction based on directional gradients and a support vector machine classifier to perform preliminary detection and classification of the identified target; a target tracking module, which is used to track the identified target using a random sampling filtering algorithm; a positioning module, which is used to use a positioning algorithm based on gait detection to locate the humanoid robot and obtain the motion trajectory of the humanoid robot; a positioning correction module, which is used to correct the motion trajectory of the humanoid robot through a sensor fusion algorithm.

[0101] The computer program product of the humanoid robot target recognition and positioning method and device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method in the previous method embodiment. The specific implementation can be found in the method embodiment, which will not be repeated here.

[0102] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned humanoid robot target recognition and positioning method, thereby completing target recognition and positioning as well as the humanoid robot's own positioning, helping the humanoid robot to understand and recognize the environment, and obtain the relative position of the humanoid robot and the target object, so as to facilitate subsequent work.

[0103] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that can be executed by a processor. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0104] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A target recognition and positioning method for a humanoid robot, characterized in that: include: Use the dual-view monocular vision positioning method to calculate the three-dimensional position of the target to be identified; Use directional gradient-based image feature extraction and support vector machine classifier to perform preliminary detection and classification of identified targets; Tracking the identified target using a random sampling filtering algorithm; Using a position algorithm based on gait detection to locate the humanoid robot, and obtaining a motion trajectory of the humanoid robot; Correcting the motion trajectory of the humanoid robot through a sensor fusion algorithm; The use of image feature extraction based on directional gradient and support vector machine classifier to perform preliminary detection and classification of the identified target includes: Grayscale and normalize the color image containing the recognition target to obtain an input image; Establishing a directional gradient feature extraction model, taking the input image as the input of the directional gradient feature extraction model, and outputting a directional gradient feature vector; Collect images containing the recognition target as samples, and label each sample with a corresponding category label to obtain a data set; Designing a support vector machine classifier, using the data set to train the support vector machine classifier, taking the directional gradient feature vector as input of the trained support vector machine classifier, outputting a category label, and obtaining a classification result of the identified target; The step of establishing a directional gradient feature extraction model, taking the input image as the input of the directional gradient feature extraction model, and outputting a directional gradient feature vector comprises: Establishing a directional gradient feature extraction model, the directional gradient feature extraction model includes an input layer, a gradient calculation layer, a directional gradient histogram layer, a feature vector splicing layer and a normalization layer; The input layer receives the input image after grayscale processing and normalization processing, wherein each pixel in the input image is represented by I norm (i,j); The gradient calculation layer calculates the horizontal gradient G in the x direction for each pixel in the input image. x (i,j)=I norm (i,j+1)-I norm (i,j-1), vertical gradient G in the y direction y (i,j)=I norm (i+1,j)-I norm (i-1,j), calculate the gradient amplitude representing the edge strength The gradient direction indicating the edge direction The directional gradient histogram layer divides the input image into a number of N×N cell units, where N is the pixel value. In each cell unit, based on the calculated gradient direction, the gradient amplitude is added to the corresponding direction interval b. For each direction interval b, the gradient amplitude belonging to the direction is calculated and accumulated. Obtain the histogram vector H of the gradient direction of the cell unit = [H1, H2, ... H b ,...H m ]; The feature vector concatenation layer concatenates the histogram vectors of the gradient direction of each adjacent cell unit in sequence to obtain a block feature vector v; The normalization layer performs normalization processing on the spliced ​​block feature vectors. Get the directional gradient feature vector, where ∈ is a positive constant; The designing of the support vector machine classifier, using the data set to train the support vector machine classifier, using the directional gradient feature vector as the input of the trained support vector machine classifier, outputting the category label, and obtaining the classification result of the identified target include: Construct a multi-class support vector machine classifier, train one support vector machine classifier for each pair of categories, and need J(J-1) / 2 classifiers in total, where J is the number of categories of the recognition target; Train each type of support vector machine classifier and adjust the loss function through the optimization algorithm Wherein, F is a penalty term, e is a weight vector, q is a bias term, and M is the number of samples. The weight vector e and the bias term q are adjusted to minimize the loss function to a preset value; The directional gradient feature vector extracted from the input image is used as input, and is input into each type of the trained support vector machine classifier, and each type of the support vector machine classifier outputs a category label, so as to obtain a classification result y∈{1,2,…,J} of the recognition target of the input image; The use of a random sampling filtering algorithm to track the identified target includes: The initial position of the identified target is (X0, Y0, Z0), an initial search window is defined at the initial position of the identified target, and the radius of the initial search window is R0; For each frame in the input image set, calculate the distribution of the color histogram in the three-dimensional color space Among them, C i is a color value in the color space, C(x, y, z) is the pixel color value in the target area, I is the indicator function, and the color difference satisfies the condition |C i -C(x, y, z)|≤ΔC, the value of the indicator function is 1, otherwise it is 0, and ΔC is the threshold value of the color difference; Update the center position (X) of the identified target t =X t-1 +ΔX t ,Y t =Y t-1 +ΔY t ,Z t =Z t-1 +ΔZ t ), where ΔX t , ΔY t , ΔZ t is the displacement of the identified target in the current frame, Update the radius of the search window Among them, S t is the area or volume of the target in the current frame; Update the position of the identified target in each frame of the input image set to obtain the three-dimensional trajectory of the identified target during the tracking process Trajectory = {(X1, Y1, Z1), (X2, Y2, Z2), ..., (X t ,Y t ,Z t )}.

2. The target recognition and positioning method of a humanoid robot according to claim 1, characterized in that: The method of using a dual-view monocular vision positioning method to calculate the three-dimensional position of the target to be identified includes: The humanoid robot is equipped with two monocular cameras, which are respectively installed at different heights, and the distance between the two monocular cameras is d; The pixel coordinates of the target to be identified on the image plane of the first monocular camera are P1(x1, y1), and the pixel coordinates of the target to be identified on the image plane of the second monocular camera are P2(x2, y2); For the first monocular camera, a geometric relationship between the first plane pixel coordinates P1 (x1, y1) of the target to be identified and the actual three-dimensional coordinates (X, Y, Z) is established. Wherein, f1 is the focal length of the first monocular camera, and D is the depth information of the target to be identified; For the second monocular camera, a geometric relationship between the second plane pixel coordinates P2 (x2, y2) of the target to be identified and the actual three-dimensional coordinates (X, Y, Z) is established. Wherein, f2 is the focal length of the second monocular camera; Calculate the vertical disparity of the target to be identified in the two plane images of the monocular cameras Calculate the depth information of the target to be identified The three-dimensional position of the target to be identified is 3. The target recognition and positioning method of a humanoid robot according to claim 1, characterized in that: The method of using a position algorithm based on gait detection to locate the humanoid robot and obtain the motion trajectory of the humanoid robot comprises: The humanoid robot is walking, and the current number of steps is The joint angle of the left leg of the humanoid robot during walking is θ L (t), the joint angle of the right leg during walking is θ R (t), 1 is the index function of gait detection. When a complete gait cycle is detected, the index function value of gait detection is 1; Combined with the IMU angle data, the current three-dimensional coordinate position of the humanoid robot is obtained, X R (t) = X R (t-1)+s·n(t)·cos(φ(t)),Y R (t) = Y R (t-1)+s·n(t)·sin(φ(t)),Z R (t) = Z R (t-1), where φ(t) is the angular change of the humanoid robot's trunk around the Z axis, and s is the step length of each gait cycle of the humanoid robot.

4. The target recognition and positioning method of a humanoid robot according to claim 1, characterized in that: The correcting the motion trajectory of the humanoid robot by using a sensor fusion algorithm includes: According to the current position (X R (t),Y R (t),Z R (t)) and the velocity at the current time t (v x (t),v y (t),v z (t)), predict the position of the humanoid robot at the next time t+Δt, Where Δt is the time interval; The sensor fusion measurement value is collected, and the X-axis position is X s (t+Δt), Y axis position is Y s (t+Δt), Z axis position is Z s (t+Δt); For the predicted position of the humanoid robot at the next moment, update the covariance matrix P(t+Δt|t)=P(t)+Q(t), where P(t) is the state covariance matrix at the current moment, and Q(t) is the process noise covariance matrix; Calculate Kalman gain Where R(t) is the measurement noise covariance matrix; Correcting the predicted position of the humanoid robot, 5. The target recognition and positioning method of a humanoid robot according to claim 4, characterized in that: The method for obtaining the sensor fusion measurement value includes: Get the acceleration (a) of the IMU measurement unit of the humanoid robot at time t x (t),a y (t),a z (t)), the speed of the humanoid robot is calculated to be v x (t+Δt)=v x (t)+a x (t)·Δt,v y (t+Δt)=v y (t)+a y (t)·Δt,v z (t+Δt)=v z (t)+a z (t)·Δt, the position of the humanoid robot is calculated as Get the IMU measurement unit measurement value; Get the distance measured by the ultrasonic sensor of the humanoid robot at time t Among them, v s is the propagation speed of the sound wave, Δt is the time from the emission to the reception of the sound wave, and the ultrasonic sensor is located in front of the humanoid robot and is 1 / 400 mm away from the position of the humanoid robot (X R (t),Y R (t),Z R (t)) has a fixed offset (ΔX u ,ΔY u ,ΔZ u ), the position of the humanoid robot is calculated to be X US (t) = X R (t)+ΔX u +d(t)·cos(θ(t))·cos(φ(t)), Y US (t) = Y R (t)+ΔY u +d(t)·cos(θ(t))·sin(φ(t)),Z US (t) = Z R (t)+ΔZ u +d(t)·sin(θ(t)), where θ(t) is the pitch angle of the sound wave propagation, and φ(t) is the azimuth angle of the sound wave propagation, to obtain the ultrasonic sensor measurement value; The IMU measurement unit measurement value and the ultrasonic sensor measurement value are fused to obtain the sensor fusion measurement value, and the X-axis position is X s (t+Δt)=w1·X IMU (t+Δt)+w2·X US (t+Δt), Y axis position is Y s (t+Δt)=w1·Y IMU (t+Δt)+w2·Y US (t+Δt), Z axis position is Z s (t+Δt)=w1·Z IMU (t+Δt)+w2·Z US (t+Δt), where w1 and w2 are weight coefficients.

6. A target recognition and positioning system for a humanoid robot, characterized in that: include: The target positioning module is used to calculate the three-dimensional position of the target to be identified using a dual-view monocular vision positioning method; The target classification module is used to perform preliminary detection and classification of the identified targets using directional gradient-based image feature extraction and support vector machine classifier; A target tracking module, used to track the identified target using a random sampling filtering algorithm; A positioning module, used to use a position algorithm based on gait detection to locate the humanoid robot and obtain a motion trajectory of the humanoid robot; A positioning correction module, used to correct the motion trajectory of the humanoid robot through a sensor fusion algorithm; The operations performed by the target classification module include: Grayscale and normalize the color image containing the recognition target to obtain an input image; Establishing a directional gradient feature extraction model, taking the input image as the input of the directional gradient feature extraction model, and outputting a directional gradient feature vector; Collect images containing the recognition target as samples, and label each sample with a corresponding category label to obtain a data set; Designing a support vector machine classifier, using the data set to train the support vector machine classifier, taking the directional gradient feature vector as input of the trained support vector machine classifier, outputting a category label, and obtaining a classification result of the identified target; The step of establishing a directional gradient feature extraction model, taking the input image as the input of the directional gradient feature extraction model, and outputting a directional gradient feature vector comprises: Establishing a directional gradient feature extraction model, the directional gradient feature extraction model includes an input layer, a gradient calculation layer, a directional gradient histogram layer, a feature vector splicing layer and a normalization layer; The input layer receives the input image after grayscale processing and normalization processing, wherein each pixel in the input image is represented by I norm (i,j); The gradient calculation layer calculates the horizontal gradient G in the x direction for each pixel in the input image. x (i,j)=I norm (i,j+1)-I norm (i,j-1), vertical gradient G in the y direction y (i,j)=I norm (i+1,j)-I norm (i-1,j), calculate the gradient amplitude representing the edge strength The gradient direction indicating the edge direction The directional gradient histogram layer divides the input image into a number of N×N cell units, where N is the pixel value. In each cell unit, based on the calculated gradient direction, the gradient amplitude is added to the corresponding direction interval b. For each direction interval b, the gradient amplitude belonging to the direction is calculated and accumulated. Get the histogram vector H of the gradient direction of the cell unit = [H1, H2, ... H b ,...H m ]; The feature vector concatenation layer concatenates the histogram vectors of the gradient direction of each adjacent cell unit in sequence to obtain a block feature vector v; The normalization layer performs normalization processing on the spliced ​​block feature vectors. Get the directional gradient feature vector, where ∈ is a positive constant; The designing of the support vector machine classifier, using the data set to train the support vector machine classifier, using the directional gradient feature vector as the input of the trained support vector machine classifier, outputting the category label, and obtaining the classification result of the identified target include: Construct a multi-class support vector machine classifier, train one support vector machine classifier for each pair of categories, and need J(J-1) / 2 classifiers in total, where J is the number of categories of the recognition target; Train each type of support vector machine classifier and adjust the loss function through the optimization algorithm Wherein, F is a penalty term, e is a weight vector, q is a bias term, and M is the number of samples. The weight vector e and the bias term q are adjusted to minimize the loss function to a preset value; The directional gradient feature vector extracted from the input image is used as input, and is input into each type of the trained support vector machine classifier, and each type of the support vector machine classifier outputs a category label, so as to obtain a classification result y∈{1,2,...,J} of the recognition target of the input image; The operations performed by the target tracking module include: The initial position of the identified target is (X0, Y0, Z0), an initial search window is defined at the initial position of the identified target, and the radius of the initial search window is R0; For each frame in the input image set, calculate the distribution of the color histogram in the three-dimensional color space Among them, C i is a color value in the color space, C(x, y, z) is the pixel color value in the target area, I is the indicator function, and the color difference satisfies the condition |C i -C(x, y, z)|≤ΔC, the value of the indicator function is 1, otherwise it is 0, and ΔC is the threshold value of the color difference; Update the center position (X) of the identified target t =X t-1 +ΔX t ,Y t =Y t-1 +ΔY t ,Z t =Z t-1 +ΔZ t ), where ΔX t , ΔY t , ΔZ t is the displacement of the identified target in the current frame, Update the radius of the search window Among them, S t is the area or volume of the target in the current frame; Update the position of the identified target in each frame of the input image set to obtain the three-dimensional trajectory of the identified target during the tracking process Trajectory = {(X1, Y1, Z1), (X2, Y2, Z2), ..., (X t ,Y t ,Z t )}.

Citation Information

Patent Citations

  • Identification method of pointer circuit breaker based on patrol robot

    CN109344768A