Robot Arm Attitude Detection Method, Device, Equipment and Computer Storage Medium

By identifying joints and connection arms of the robotic arm pictures, and using machine learning models for posture combination and abnormal detection, the problems of high hardware costs and large computing resource utilization in the existing technology are solved, and efficient monitoring and safety improvement of multi-robot attitudes are achieved.

CN111798518BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010691783.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-17
Publication Date
2025-07-15
Estimated Expiration
2040-07-17

AI Technical Summary

Technical Problem

In the prior art, the attitude detection of the robotic arms relies on a 3D camera, resulting in high hardware costs and large computing resources, making it difficult to effectively monitor the attitude of multiple robotic arms, affecting the safety and efficiency of the production line.

Method used

By identifying joints and connection arms pictures of the robotic arm, the machine learning-trained model is used to extract joint and connection arms feature data, perform pose combination and abnormal detection, reducing hardware requirements and improving computing efficiency.

Benefits of technology

It realizes efficient monitoring of the posture of multiple robotic arms, reduces hardware costs and computing resource requirements, improves the safety and production efficiency of robotic arms work, and can detect and repair abnormalities in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111798518B_ABST
    Figure CN111798518B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and computer storage medium for detecting the posture of a robotic arm, which relates to the field of computer technology. In this method, a to-be-detected picture including at least one robotic arm is acquired, and joint feature data of each joint and connecting arm feature data of each connecting arm included in the at least one robotic arm in the to-be-detected picture are acquired; posture combination is performed according to the joint feature data of each joint and the connecting arm feature data of each connecting arm to determine the posture of the at least one robotic arm; posture anomaly detection is respectively performed on each robotic arm according to the posture of each robotic arm in the at least one robotic arm to obtain a posture detection result of each robotic arm, and the posture detection result indicates whether the posture of the robotic arm is abnormal, so as to realize the monitoring of the posture condition of the robotic arm, be able to discover in time when an anomaly occurs, and further improve the safety and effectiveness of the robotic arm's work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and particularly to the field of monitoring technology, and provides a method, a device, a device and a computer storage medium for detecting the posture of a robotic arm. Background Art

[0002] With the development of industrial automation, factory automation has become the mainstream. A robotic arm is a complex system with high precision and strong coupling. Due to its unique operation flexibility, it is widely used in the field of industrial assembly, undertakes highly refined work on the production line, and has become a key device in an automated factory. A simple robotic arm moves along a fixed trajectory and works in a relatively stable working condition, while a complex robotic arm operates in an unstable unstructured environment and calculates the moving trajectory in real time. If the robotic arm malfunctions, it will not only delay the progress of the production line but also may cause safety accidents. Especially for complex robotic arms, due to the instability of environmental factors, as well as the influence of sudden interventions and cumulative errors, it is more likely to malfunction.

[0003] Therefore, in order to ensure the safe and effective implementation of production work, it is necessary to monitor the working conditions of the robotic arm and then repair the abnormality in time when an abnormality occurs. Summary of the Invention

[0004] Embodiments of the present application provide a method, a device, a device and a computer storage medium for detecting the posture of a robotic arm, which are used to monitor the posture of the robotic arm, thereby improving the safety and effectiveness of the robotic arm work.

[0005] On the one hand, a method for detecting the posture of a robotic arm is provided, including:

[0006] Obtaining a to-be-detected picture including at least one robotic arm;

[0007] Obtaining joint feature data of each joint included in the at least one robotic arm and connecting arm feature data of each connecting arm in the to-be-detected picture;

[0008] Performing posture combination according to the joint feature data of each joint and the connecting arm feature data of each connecting arm to determine the posture of the at least one robotic arm;

[0009] Performing posture abnormality detection on each robotic arm respectively according to the posture of each robotic arm in the at least one robotic arm to obtain a posture detection result of each robotic arm, where the posture detection result indicates whether the posture of the robotic arm is abnormal.

[0010] On the one hand, a device for detecting the posture of a robotic arm is provided, including:

[0011] A picture acquisition unit, configured to obtain a to-be-detected picture including at least one robotic arm;

[0012] An identification unit for obtaining joint feature data of each joint included in the at least one robotic arm and connecting arm feature data of each connecting arm in the to-be-detected picture;

[0013] An attitude combination unit for performing attitude combination according to the joint feature data of each joint and the connecting arm feature data of each connecting arm to determine the attitude of the at least one robotic arm;

[0014] An attitude anomaly detection unit for respectively performing attitude anomaly detection on each robotic arm according to the attitude of each robotic arm in the at least one robotic arm to obtain an attitude detection result of each robotic arm, where the attitude detection result indicates whether the attitude of the robotic arm is abnormal.

[0015] Optionally, the identification unit is specifically configured to:

[0016] Perform joint identification on the to-be-detected picture to obtain the joint feature data; the joint feature data includes J joint feature maps corresponding to the joints of J parts, and each joint feature map is used to represent the probability that each pixel point in the to-be-detected picture is a joint pixel point;

[0017] Perform connecting arm identification on the to-be-detected picture to obtain the connecting arm feature data, where the connecting arm feature data includes C connecting arm feature maps corresponding to the connecting arms of C parts, and each connecting arm feature map is composed of vectors of each pixel point in the to-be-detected picture, and non-zero vectors in each vector field represent that the pixel points corresponding to the non-zero vectors are connecting arm pixel points.

[0018] Optionally, the identification unit is specifically configured to:

[0019] Use a trained joint identification model to perform joint identification on the to-be-detected picture to obtain the joint feature data;

[0020] Wherein, the joint identification model is trained by using a plurality of picture training samples, and each picture training sample is labeled with the regions where each joint is located in the picture training sample.

[0021] Optionally, the identification unit is specifically configured to:

[0022] Extract features from the to-be-detected picture to obtain initial joint features of the to-be-detected picture;

[0023] According to the initial joint features, respectively determine the probability that each pixel point on the to-be-detected picture is a joint pixel point of each part to obtain the joint feature data.

[0024] Optionally, the identification unit is specifically configured to:

[0025] Using the trained connecting arm recognition model, perform connecting arm recognition on the to-be-detected picture to obtain the connecting arm feature data;

[0026] Wherein, the connecting arm recognition model is trained using a plurality of picture training samples, and each picture training sample is labeled with the vectors of each connecting arm in the picture training sample.

[0027] Optionally, the recognition unit is specifically configured to:

[0028] Extract features from the to-be-detected picture to obtain the initial connecting arm features of the to-be-detected picture;

[0029] According to the initial connecting arm features, respectively determine the vectors of each pixel point on the to-be-detected picture to obtain the joint feature data.

[0030] Optionally, the pose combination unit is specifically configured to:

[0031] Perform pose combination according to the J joint feature maps and the C connecting arm feature maps to obtain the poses of the at least one robotic arm.

[0032] Optionally, the pose combination unit is specifically configured to:

[0033] According to the J joint feature maps, construct J joint sets, and each joint set is composed of joints of the same part of the at least one robotic arm;

[0034] Based on the J joint sets, construct a plurality of pose sets, and each pose set includes a plurality of poses formed by the J joints of the J joint sets, and any two joints among the J joints belong to different joint sets;

[0035] Determine the target pose set that is the pose of each robotic arm from the plurality of pose sets.

[0036] Optionally, the pose combination unit is specifically configured to:

[0037] According to the C connecting arm feature maps, determine the correlation degree between any two joints among the J joints that make up each pose, and the correlation degree is used to characterize the probability that two joints belong to the same robotic arm;

[0038] Based on the correlation degree, determine the target pose set that meets the preset conditions from the plurality of pose sets, wherein the preset conditions are that any two poses in the pose set do not include the same joints, and the sum of the correlation degrees of the pose set is the maximum value among the plurality of pose sets.

[0039] Optionally, the pose combination unit is specifically configured to:

[0040] Determine the correlation degree between any two joints according to the vectors of each pixel located between the any two joints and the vector formed by the any two joints.

[0041] Optionally, the posture anomaly detection unit is specifically configured to:

[0042] Determine the probability value that the posture of each robotic arm is a normal posture according to the coordinates of the joint points constituting the posture of each robotic arm and the trained posture distribution function; the posture distribution function is obtained through machine learning using multiple posture training samples, and each posture sample is labeled with a representation of normal or abnormal posture;

[0043] When it is determined that the probability value is less than a preset probability threshold, determine that the posture of each robotic arm is abnormal.

[0044] On the one hand, a robotic arm posture monitoring system is provided, including:

[0045] At least one image acquisition device, and each image acquisition device is used to acquire an image or video stream including at least one robotic arm;

[0046] A robotic arm posture detection device, which is used to adopt the steps of any of the above methods to obtain the posture detection results of each robotic arm in the video frame of the image or the video stream, and trigger an alarm when there is a posture detection result indicating that the robotic arm posture is abnormal.

[0047] On the one hand, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above methods are implemented.

[0048] On the one hand, a computer storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the steps of any of the above methods are implemented.

[0049] On the one hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any of the above methods.

[0050] In the embodiments of the present application, by identifying joints and connecting arms in a picture containing at least one robotic arm, and performing pose combination based on joint feature data and connecting arm feature data, the pose of each robotic arm in the picture is determined, and then whether the pose of each robotic arm is abnormal is determined. In this way, the embodiments of the present application can determine the poses of multiple robotic arms based on pictures of the robotic arms, with higher efficiency in determining the poses of the robotic arms. Moreover, by further determining whether the robotic arm is abnormal based on the identified pose of the robotic arm, maintenance can be carried out in a timely manner when an abnormality occurs, thereby improving the safety and effectiveness of the operation of the robotic arm and effectively ensuring the safe and effective implementation of production work by the robotic arm. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0052] Figure 1 An example diagram of a robotic arm provided by an embodiment of the present application;

[0053] Figure 2 Another example diagram of a robotic arm provided by an embodiment of the present application;

[0054] Figure 3 Still another example diagram of a robotic arm provided by an embodiment of the present application;

[0055] Figure 4 A schematic diagram of a robotic arm pose monitoring system provided by an embodiment of the present application;

[0056] Figure 5 A schematic diagram of the training process of a joint recognition model provided by an embodiment of the present application;

[0057] Figure 6 A schematic diagram of a labeling of a picture training sample provided by an embodiment of the present application;

[0058] Figure 7 Another schematic diagram of a labeling of a picture training sample provided by an embodiment of the present application;

[0059] Figure 8 A schematic diagram of a joint feature map obtained by converting a labeling map provided by an embodiment of the present application;

[0060] Figure 9 A schematic diagram of a structure of a joint recognition model provided by an embodiment of the present application;

[0061] Figure 10Schematic diagram of the training process of the connecting arm recognition model provided by the embodiment of the present application;

[0062] Figure 11 Schematic diagram of the annotation of the picture training sample provided by the embodiment of the present application;

[0063] Figure 12 Schematic diagram of a structure of the connecting arm recognition model provided by the embodiment of the present application;

[0064] Figure 13 Schematic diagram of the process of the robotic arm attitude detection method provided by the embodiment of the present application;

[0065] Figure 14 Schematic diagram of the joint set provided by the embodiment of the present application;

[0066] Figure 15 Schematic diagram of the bipartite graph based on the joint set provided by the embodiment of the present application;

[0067] Figure 16 Schematic diagram of the attitude of the combined robotic arm provided by the embodiment of the present application;

[0068] Figure 17 Schematic diagram of a structure of the robotic arm attitude detection device provided by the embodiment of the present application;

[0069] Figure 18 Schematic diagram of a structure of the computer device provided by the embodiment of the present application. Detailed implementation manners

[0070] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the scope of protection of the present application. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0071] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application are first explained here:

[0072] Robotic arm: A mechanical device composed of multiple joints and connecting arms, such as Figure 1 , Figure 2 and Figure 3As shown, it is an example diagram of several robotic arms. Among them, the position of the white dashed box indicates the location of the joints of the robotic arm, and the white dashed line indicates the direction of the connecting arm of the robotic arm. For example Figures 1 to 3 as shown in Figures 1 to 3 , a posture of a robotic arm can be obtained through the combination of joints and connecting arms.

[0073] Convolutional Neural Networks (CNN): Used in the field of machine learning, it is a type of feed-forward neural network with a deep structure that contains convolutional calculations. Convolutional neural networks can learn grid-like topology features, such as pixels, with a relatively small amount of computation, have stable effects, and have no additional feature engineering requirements for data. The most important part of CNN is the convolutional layer. For a device, an image is essentially stored in the form of a pixel matrix. Therefore, the processing of an image is essentially based on this pixel matrix. In the convolutional layer, the pixel matrix is convolved with a convolutional kernel of a preset step size and a preset size.

[0074] Feature map: It is extracted through the convolutional layer of the above-mentioned convolutional neural network. In essence, it is also a pixel matrix. Each element in the pixel matrix can be regarded as a pixel point on the feature map, and the value at the position of this pixel point is the feature value of a region or a pixel point in the original image.

[0075] Part Affinity Field (PAF): In the pose detection algorithm based on PAF, PAF is a set of one or more two-dimensional (2D) vectors. Each set of 2D vectors encodes the position and direction of a connecting arm and can be used to measure the affinity between two joints. The role of PAF is reflected in the pose combination stage.

[0076] Vector field: In the embodiments of the present application, the vector field is composed of vectors at the positions of each pixel point where the connecting arm is located. The vector at each pixel point position is the unit vector of the connecting arm where it is located.

[0077] To ensure the safe and effective implementation of production work, it is necessary to monitor the working conditions of the robotic arm, such as detecting the posture of the robotic arm, and then repairing the abnormality in time when an abnormality occurs. In the prior art, the posture detection of the robotic arm generally relies on the 3D video images captured by a 3D camera for posture estimation. However, this solution relying on a 3D camera can only monitor the posture of one robotic arm at the same time. In an automated factory, there is often a production line composed of multiple robotic arms. Based on this solution, a large number of 3D cameras are required, resulting in high hardware costs. Moreover, a large number of video images need to be subjected to posture estimation, which requires a large amount of computing resources. Therefore, there is an urgent need for a robotic arm posture detection solution with low cost and high efficiency.

[0078] In view of this, the embodiments of the present application provide a robotic arm posture detection method. In this method, by identifying joints and connecting arms in a picture containing at least one robotic arm, and combining postures according to joint feature data and connecting arm feature data, the postures of each robotic arm in the picture are determined, and then whether the postures of each robotic arm are abnormal is determined. It can be seen that the technical solution of the embodiments of the present application can determine the postures of multiple robotic arms simultaneously according to the pictures of the robotic arms. In this way, fewer image acquisition devices such as cameras are required to monitor the robotic arms, reducing the hardware cost. In addition, when determining the postures of the same number of robotic arms, the required computing resources are correspondingly less, the processing efficiency is higher, and the computing resources are more saved. Moreover, by further determining whether the robotic arm is abnormal based on the identified robotic arm posture, the abnormality can be maintained in time when an abnormality occurs, thereby improving the safety and effectiveness of the robotic arm work and effectively ensuring the safe and effective implementation of production work by the robotic arm.

[0079] In the embodiments of the present application, a large number of manually labeled training samples are used to train a model by machine learning, so as to identify joints and connecting arms in the picture through the trained model, improve the recognition rate of robotic arm joints and connecting arms, and provide a feature data basis for posture recognition.

[0080] After introducing the design concept of the embodiments of the present application, the following briefly introduces the application scenarios applicable to the technical solutions of the embodiments of the present application. It should be noted that the application scenarios introduced below are only for illustrating the embodiments of the present application rather than limiting. In the specific implementation process, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.

[0081] The solution provided by the embodiments of the present application can be applied to most scenarios that require robotic arm posture monitoring, and is particularly suitable for the monitoring of automated production lines composed of multiple robotic arms, such as Figure 4 As shown, it is a schematic diagram of a robotic arm posture monitoring system provided by the embodiments of the present application. The system includes an image acquisition device 40 and a robotic arm posture detection device 41.

[0082] The image acquisition device 40 is used to acquire images or video streams including robotic arms, such as a camera. The robotic arm can be, for example, a robotic arm working on an assembly line, or a robotic arm included in a mobile device, such as a robotic arm included in a mobile robot, etc.

[0083] The robotic arm pose detection device 41 is a computer device with certain processing capabilities, such as a personal computer (PC), a laptop, or a server, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms, but is not limited thereto.

[0084] The robotic arm pose detection device 41 includes one or more processors 411, a memory 412, and an I / O interface 413 for interacting with other devices, etc. In addition, the robotic arm pose detection device 41 can also be configured with a database 414, and the database 414 can be used to store data such as model data and received video streams involved in the solutions provided in the embodiments of the present application. Among them, program instructions for the robotic arm pose detection method provided in the embodiments of the present application can be stored in the memory 412 of the robotic arm pose detection device 41. When these program instructions are executed by the processor 411, they can be used to implement the steps of the robotic arm pose detection method provided in the embodiments of the present application to determine whether the poses of the robotic arms in the picture are abnormal.

[0085] In the specific implementation process, the image acquisition device 40 can be set at a position where the robotic arm can be photographed, and acquire images or video streams of the robotic arm. The acquired images or video streams can be stored in the server, and the robotic arm pose detection device 41 obtains the images or video streams from the server. Of course, they can also be directly sent to the robotic arm pose detection device 41 for pose detection. When acquiring images, the image acquisition device 40 can acquire images at regular intervals, or acquire images when it detects the presence of the robotic arm according to the actual situation.

[0086] For the video stream collected by the image acquisition device 40, the robotic arm attitude detection device 41 can intercept video frames therefrom as the pictures to be detected, and perform attitude recognition and attitude anomaly detection on the robotic arm in the pictures to be detected, so as to determine whether the attitude of the robotic arm is abnormal. When it is determined that the attitude of the robotic arm is abnormal, an alarm will be triggered. For example, a warning message will be sent to the device of the relevant personnel, or an instruction will be sent to the alarm device related to the robotic arm to control the alarm device to give an alarm, so that the relevant personnel can know in time that the robotic arm has a fault and repair it in time, or a warning will be sent to the management device of the robotic arm. The management device can determine the type of the fault of the robotic arm according to the operation log of the robotic arm. If the fault can be repaired by software, a repair program can be sent to the robotic arm to make the robotic arm repair itself.

[0087] In practical applications, in order to improve the accuracy of attitude determination, for the same robotic arm, multiple image acquisition devices can be respectively set to capture images of these robotic arms, and then the attitudes of the respective robotic arms can be recognized from these images respectively, and then the accuracy of each attitude result can be determined by synthesizing multiple attitude recognition results. In addition, the attitude can also be estimated by sensors arranged on the robotic arm. The sensors can be, for example, gyroscopes, distance sensors, etc. Then, the attitudes of the respective robotic arms can be comprehensively determined by combining the image recognition method and the sensor recognition method, further improving the accuracy of attitude recognition.

[0088] The image acquisition device 40 and the robotic arm attitude detection device 41 can be directly or indirectly communicatively connected through one or more networks 42. The network 42 can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and the embodiments of the present invention do not limit this.

[0089] Of course, the method provided by the embodiments of the present application is not limited to Figure 4 the application scenarios shown, and can also be used in other possible application scenarios, and the embodiments of the present application do not limit this. For Figure 4 the functions that can be realized by each device in the application scenarios shown will be described together in the subsequent method embodiments, and will not be elaborated here too much. Next, the technologies involved in the embodiments of the present application will be briefly introduced.

[0090] In an alternative embodiment, the embodiments of the present application can adopt an entity device combined with artificial intelligence (AI) technology to implement the robotic arm attitude detection process. In another alternative embodiment, the robotic arm attitude detection process can also be implemented by combining cloud technology with AI technology.

[0091] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important supporting technology. The background services of network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system background, which can only be achieved through cloud computing. Specifically, in addition to being able to execute program processes through physical computing resources and use physical storage resources to achieve data storage, the embodiments of the present application can also use the computing resources provided by the cloud to perform robotic arm pose detection, and the data involved in the pose detection process can all be stored through the storage resources provided by the cloud.

[0092] AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing (NLP) technology, and machine learning / deep learning. The technical solutions provided by the embodiments of the present application mainly involve technologies such as machine learning / deep learning in artificial intelligence.

[0093] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Specifically, in the embodiments of the present application, the attitude detection of the robotic arm can be performed through a model obtained by machine learning.

[0094] In the embodiments of the present application, both the joint recognition and the connecting arm recognition in the attitude recognition process can be implemented through a pre-trained model. Therefore, before introducing the robotic arm attitude detection method of the embodiments of the present application, the process of model training will be introduced first.

[0095] As Figure 5 shown, it is a schematic diagram of the training process of the joint recognition model.

[0096] Step 501: Obtain picture training samples.

[0097] Before the joint recognition model is specifically put into use, it is necessary to perform pre-training on the joint recognition model to obtain a mature joint recognition model. Among them, the joint recognition model can be trained through multiple picture training samples, and each picture training sample marks the regions where each joint is located in the picture training sample. As Figure 6 and Figure 7 shown, it is a schematic diagram of the annotation of the picture training sample. Since each robotic arm can include multiple joints, all joints can be marked on one picture, but different joints use different annotations. Of course, to classify the joints, that is, the joints in different parts belong to one category, each joint can be marked separately, and one type of joint is marked on one picture. As Figure 6 shown, it marks the region of joint 1 at the chuck part, and Figure 7 shown, it marks the region of joint 2 at the upper joint part. A picture training sample can include multiple annotation pictures. For a picture training sample, if the robotic arm in it includes several joints, then the picture training sample can include the corresponding number of annotation pictures. For example, when the robotic arm includes 3 joints, the number of annotation pictures can also be 3.

[0098] After the annotation of each picture training sample is completed and before training, preprocessing of each picture training sample can also be performed. Specifically, the preprocessing can convert the annotation pictures of each picture training sample into joint feature pictures. As Figure 8As shown in the figure, it is a schematic diagram of the joint feature map obtained by converting the annotation map. Each small square can represent a pixel point, and the value in the small box represents the probability that the pixel point at this position is a joint pixel point. Among them, the probability value of the small square corresponding to the annotation area is 1, while the probability of the remaining areas except the annotation area is 0. Of course, the gradient descent method can also be used, that is, as the distance from the center position of the annotation area increases, the probability value gradually decreases.

[0099] Step 502: Obtain the joint feature maps of each picture training sample through the joint recognition model.

[0100] After preprocessing the samples of each picture training sample, they can be used for the training of the joint recognition model. As Figure 9 shown, it is a schematic structural diagram of the joint recognition model provided by an embodiment of the present application. The joint recognition model can be a model based on the CNN network, including multiple convolutional layers, an output layer, and a loss layer. Of course, this structural diagram is only one possible structure and not the only one. Under possible circumstances, other possible structures can also be used, and the embodiments of the present application do not limit this.

[0101] Among them, before inputting the picture training sample into the joint recognition model, image preprocessing can also be performed on each picture training sample. Image preprocessing can be to use the CNN network to perform certain feature extraction on the picture training sample to obtain one or more feature maps of the picture training sample, and then input the feature map into the joint recognition model. Of course, the step of image preprocessing can be selected whether to implement in actual applications.

[0102] In the first training, the joint recognition model is an initial model. The convolutional layer of the joint recognition model is used to extract features from the input picture training sample layer by layer to obtain the initial features of the picture training sample. Specifically, the features obtained by the previous convolutional layer can be used as the input of the next convolutional layer, so as to obtain the features of the picture training sample layer by layer. The initial features output by the last convolutional layer are input into the output layer. The output layer is used to determine the probability that each pixel point on the picture training sample is a joint pixel point based on the initial features and output the joint feature map. Corresponding to the annotation map, the number of joint feature maps can also be the same as the joint type, that is, when the picture training sample includes joints of J parts, the number of obtained joint feature maps is also J. Each joint feature map represents the probability that each pixel point belongs to the joint pixel points of this type of joint.

[0103] Step 503: Determine whether the loss value is less than the preset loss value.

[0104] During the training process, the loss layer can be used to calculate the loss value between the joint feature map output by the joint recognition model and the corresponding joint feature map of the annotation map, so as to determine whether the loss value is less than the preset loss value. When it is determined that the loss value is less than the preset loss value, the model training ends. If it is determined that the loss value is not less than the preset loss value, the model needs to be continuously trained. The loss layer can use a loss function to calculate the loss value. For example, the loss function can be an absolute value loss function, a squared loss function, a logarithmic loss function, or a cross entropy error function, etc. Of course, it can also be other possible loss functions, and the embodiments of the present application do not limit this.

[0105] Among them, in addition to using the loss value to measure whether the model needs to be continuously trained, the accuracy of the current model can also be statistically calculated, and it can be determined whether the model needs to be continuously trained according to the accuracy.

[0106] Step 504: If the determination result in step 503 is no, the model parameters are adjusted according to the loss value.

[0107] Specifically, the adjustment value of the model parameters can be determined according to the loss value. After adjusting the model parameters based on the adjustment value, the adjusted model can be used to continue to obtain the joint feature maps of each picture training sample, that is, enter a new round of loop of steps 502 to 504 until it is determined that the loss value is less than the preset loss value, and the process ends.

[0108] Next, the training process of the connecting arm recognition model will be introduced. As Figure 10 shown, it is a schematic diagram of the training process of the connecting arm recognition model.

[0109] Step 1001: Obtain picture training samples.

[0110] Before the connecting arm recognition model is specifically put into use, it is necessary to perform pre-training on the connecting arm recognition model to obtain a mature connecting arm recognition model. Among them, the connecting arm recognition model can be trained through multiple picture training samples, and each picture training sample annotates the regions where each connecting arm is located in the picture training sample. Since each robotic arm can include multiple connecting arms, all the connecting arms can be marked on one picture, but different connecting arms use different markings. Of course, in order to classify the connecting arms, that is, the connecting arms in different parts belong to one category, each connecting arm can be marked separately, and one type of connecting arm is marked on one picture, such as Figure 11As shown in the figure, it is a schematic diagram of the annotation of the picture training sample, which annotates the vector of the connecting arm of one of the parts. This vector is essentially the vector from one joint connected by the connecting arm to another joint. Such a vector can be called the Part Affinity Field (PAF). It stores the position and direction information in the support area of the connecting arm. The part affinity field is a two-dimensional vector field for each connecting arm, and it can be defined as follows: for each pixel in the area belonging to the connecting arm, the two-dimensional vector encodes the direction from one joint connected by the connecting arm to another joint. Each type of connecting arm has a corresponding vector field to connect their corresponding joints.

[0111] A picture training sample can include multiple annotation diagrams. For a picture training sample, if the robotic arm in it has a certain number of connecting arms, then the picture training sample can include the corresponding number of annotation diagrams. For example, when the robotic arm has 2 connecting arms, the number of annotation diagrams can also be 2.

[0112] After the annotation of each picture training sample is completed and before training, preprocessing of each picture training sample can also be performed. Specifically, the preprocessing can convert the annotation diagrams of each picture training sample into connecting arm feature maps. For a picture training sample with C robotic arms, C connecting arm feature maps can be obtained. Each connecting arm feature map can be a matrix of w*h*2, where w is the width of the picture training sample, h is the height of the picture training sample, and 2 refers to the number of channels. In the connecting arm feature map, the value at each pixel point represents the vector of that pixel point, and the non-zero vectors in each vector field indicate that the pixel points corresponding to the non-zero vectors are the pixel points of the connecting arm.

[0113] Specifically, each connecting arm feature map can be represented by L i i ranges from 1 to C. For each L i , a 2D vector is generated for each pixel point on the picture. The calculation method of the vector is as follows: if the pixel point p is on the c-th type of connecting arm of the k-th robotic arm, then L c,k (p) is the unit vector of this connecting arm, otherwise it is 0. Since there may be an overlap of robotic arms, when the pixel point p is in the overlapping area of the connecting arms, then L c (p) is the mean value of the unit vectors of the connecting arms in the overlapping part. Then there is the following relationship:

[0114]

[0115]

[0116] Step 1002: Obtain the connecting arm feature maps of each picture training sample through the connecting arm recognition model.

[0117] After preprocessing the sample of each picture training sample, it can be used for training the connecting arm recognition model. As Figure 12 shown, a schematic structural diagram of the connecting arm recognition model provided by an embodiment of the present application. The connecting arm recognition model can be a model based on a CNN network, including multiple convolutional layers, an output layer, and a loss layer. Of course, this schematic structural diagram is only one possible structure and not the only one. Under possible circumstances, other possible structures can also be adopted, and the embodiments of the present application do not limit this.

[0118] Among them, before inputting the picture training sample into the connecting arm recognition model, image preprocessing can also be performed on each picture training sample. The image preprocessing can be to perform certain feature extraction on the picture training sample using a CNN network to obtain one or more feature maps of the picture training sample, and then input the feature maps into the connecting arm recognition model. Of course, the step of image preprocessing can be selected whether to be implemented in actual applications.

[0119] During the first training, the connecting arm recognition model is an initial model. The convolutional layer of the connecting arm recognition model is used to extract features from the input picture training sample layer by layer to obtain the initial features of the picture training sample. Specifically, the features obtained by the previous convolutional layer can be used as the input of the next convolutional layer, so as to obtain the features of the picture training sample layer by layer. The initial features output by the last convolutional layer are input into the output layer. The output layer is used to determine the probability that each pixel point on the picture training sample is a connecting arm pixel point based on the initial features and output a connecting arm feature map. Corresponding to the annotation map, the number of connecting arm feature maps can also be the same as the type of connecting arm. That is, when the picture training sample includes connecting arms of C parts, the number of obtained connecting arm feature maps is also C. The value of each pixel point on each connecting arm feature map represents the vector of the pixel point.

[0120] Step 1003: Determine whether the loss value is less than a preset loss value.

[0121] During the training process, the loss layer can be used to calculate the loss value between the connecting arm feature map output by the connecting arm recognition model and the connecting arm feature map corresponding to the annotation map, so as to determine whether the loss value is less than the preset loss value. When it is determined that the loss value is less than the preset loss value, the model training ends. If it is determined that the loss value is not less than the preset loss value, the model needs to be continuously trained.

[0122] Among them, in addition to using the loss value to measure whether the model needs to be continuously trained, the accuracy of the current model can also be statistically analyzed, and it can be determined whether the model needs to be continuously trained according to the accuracy.

[0123] Step 1004: If the determination result of step 1003 is no, adjust the model parameters according to the loss value.

[0124] Specifically, the adjustment value of the model parameters can be determined according to the loss value. After adjusting the model parameters based on the adjustment value, the adjusted model can be used to continue to obtain the connecting arm feature maps of each picture training sample, that is, enter a new round of loop of step 1002 to step 1004 until it is determined that the loss value is less than the preset loss value, and the process ends.

[0125] In the specific application process, since there is a certain connection between the joints and the connecting arms, the joint recognition model and the connecting arm recognition model can be trained simultaneously. For example, after obtaining the joint feature maps and the connecting arm feature maps through the joint recognition model and the connecting arm recognition model respectively, the loss layer can be used to calculate the sum of the loss values of the joint recognition model and the connecting arm recognition model, and then determine whether to end the training according to the sum of the loss values.

[0126] In the embodiment of the present application, after the model training is completed, the model can be applied to the actual pose detection process. Please refer to Figure 13 , which is a schematic flowchart of the robotic arm pose detection method provided by the embodiment of the present application. This method can be executed by Figure 4 the robotic arm pose detection device 41 in, and the process of this method is introduced as follows.

[0127] Step 1301: Obtain a to-be-detected picture including at least one robotic arm.

[0128] In the embodiment of the present application, the to-be-detected picture can be a picture directly collected by an image acquisition device, or a video frame intercepted from a video stream collected by the image acquisition device. Among them, if it is intercepted from a video stream, it can be intercepted regularly, and the regular duration can be set according to the actual situation. In addition, it can also be intercepted randomly. Considering that when monitoring a movable device such as a mobile robot, there may be a picture that does not include a robotic arm. Therefore, for the intercepted video frame, it is also necessary to further confirm whether there is a robotic arm in the video frame. If not, the video frame is discarded. If so, the subsequent process is performed on the video frame.

[0129] In practical applications, the robotic arm pose detection process for each to-be-detected picture is the same. Therefore, the robotic arm pose detection process is introduced below taking one to-be-detected picture as an example.

[0130] Step 1302: Obtain the joint feature data of each joint included in at least one robotic arm and the connecting arm feature data of each connecting arm in the to-be-detected picture.

[0131] In the embodiments of the present application, after obtaining the picture to be detected, the picture to be detected can be preprocessed. The image preprocessing can be to use a CNN network to extract certain features from the picture to be detected to obtain one or more feature maps of the picture to be detected, and subsequent processing can be performed based on these one or more feature maps. Of course, the steps of image preprocessing can be selected whether to be implemented in actual applications.

[0132] Considering that the posture of the robotic arm is mainly composed of each joint and each connecting arm included in the robotic arm, therefore, in order to obtain the posture of each robotic arm, the joint feature data of each joint and the connecting arm feature data of each connecting arm in the picture to be detected can be obtained, so as to combine the postures to obtain the postures of each robotic arm.

[0133] Specifically, the joint feature data or the connecting arm feature data can be data representing the positions of each joint or each connecting arm in the picture to be detected. For example, through a border detection algorithm, the borders of the regions where each joint and each connecting arm are located can be marked in the picture to be detected, and then the adjacent joints and connecting arms can be combined according to the positions of the borders, so as to obtain the postures of each robotic arm.

[0134] Specifically, obtaining the joint feature data of each joint included in at least one robotic arm in the picture to be detected can also be to perform joint recognition on the picture to be detected to obtain the joint feature data. Among them, performing joint recognition on the picture to be detected can be to use a trained joint recognition model, that is, a joint recognition model trained through the Figure 5 process shown, to perform joint recognition on the picture to be detected to obtain the joint feature data.

[0135] Among them, the joint feature data can include J joint feature maps (S1, S2,..., S J ) corresponding to the joints of J parts. Each joint feature map S j is a matrix of w*h, the same size as the picture to be detected. Each joint feature map S j is used to represent the probability that each pixel point in the picture to be detected is a joint pixel point. The value of j ranges from 1 to J, and J is a positive integer.

[0136] When using the joint recognition model to perform joint recognition on the picture to be detected, as Figure 9 shown in the structure of the joint recognition model shown, the convolutional layer included in the joint recognition model can be used to extract features from the picture to be detected layer by layer to obtain the initial joint features of the picture to be detected, and then according to the initial joint features, the probability that each pixel point on the picture to be detected is a joint pixel point of each part is determined respectively to obtain the joint feature data. For the determination of the probability of each pixel point, classification can be performed using a classification algorithm to determine the probability that each pixel point belongs to each type of joint.

[0137] Specifically, the connection arm feature data of each connection arm included in at least one robotic arm in the image to be detected can also be obtained by performing connection arm recognition on the image to be detected to obtain the connection arm feature data. Among them, performing connection arm recognition on the image to be detected can be using a trained connection arm recognition model, that is, the connection arm recognition model trained through the Figure 10 process shown, to perform connection arm recognition on the image to be detected to obtain the connection arm feature data.

[0138] Among them, the connection arm feature data can include C connection arm feature maps (L1, L2,..., L c ) corresponding to C parts of the connection arm. Each connection arm feature map L i is a matrix of w*h*2. Each connection arm feature map is composed of vectors of each pixel point in the image to be detected. Each non-zero vector in each connection arm feature map L i represents that the pixel point corresponding to the non-zero vector is a connection arm pixel point. The value of i ranges from 1 to C, and C is a positive integer.

[0139] When using the connection arm recognition model to perform connection arm recognition on the image to be detected, as Figure 12 shown in the structure of the connection arm recognition model, the convolutional layers included in the connection arm recognition model can be used to extract features from the image to be detected layer by layer to obtain the initial connection arm features of the image to be detected, and then according to the initial connection arm features, vectors of each pixel point on the image to be detected are determined respectively to obtain the connection arm feature data.

[0140] Step 1303: Perform pose combination according to the joint feature data of each joint and the connection arm feature data of each connection arm to determine the pose of at least one robotic arm.

[0141] In the embodiments of the present application, through the process of step 1302, the obtained joint feature data can be J joint feature maps (S1, S2,..., S J ), and the obtained connection arm feature data can be C connection arm feature maps, (L1, L2,..., L c ). Then, pose combination can be performed according to the J joint feature maps and the C connection arm feature maps to obtain the poses of each robotic arm in the image to be detected.

[0142] The process of pose combination is essentially a process of using the J joint feature maps and the C connection arm feature maps as a basis to find the joints belonging to the same robotic arm. Once the joints of each part of each robotic arm are found, then according to the combination of each joint, or according to the combination of each joint and the connection arm, the poses of each robotic arm can be obtained.

[0143] Based on J joint feature maps, J joint sets can be constructed. Each joint set is composed of joints of the same part of at least one robotic arm. Based on the J joint sets, multiple pose sets can be constructed. Each pose set includes multiple poses formed by the J joints of the J joint sets. Any two joints among the J joints belong to different joint sets. For example, Figure 14 As shown, when there are 2 robotic arms in the image to be detected, namely robotic arm 1 and robotic arm 2, and each robotic arm includes 3 joints, namely joint A, joint B, and joint C, then 3 joint sets S A =(A1, A2), S B =(B1, B2) and S C =(C1, C2) can be constructed. These 3 joint sets can construct multiple pose sets. For example, (A1 - B1 - C1, A2 - B2 - C2), (A1 - B2 - C1, A2 - B1 - C2), and (A1 - B1 - C1, A2 - B1 - C2), etc. Then the process of determining the pose is to determine the target pose set that is truly the pose of each robotic arm from multiple pose sets.

[0144] Obviously, compared with the real situation, it is impossible for two robotic arms to share a joint. Therefore, all joints that make up any two poses are different, which is a constraint condition for determining the target pose set. In addition, there is a connecting arm between two adjacent joints that are truly the pose of the robotic arm. On the image, it is manifested that the pixel points between two adjacent joints are connecting arm pixel points. Therefore, when determining the target pose set, the correlation degree between any two joints among the J joints that make up each pose can also be determined according to the connecting arm feature map. The correlation degree can represent the probability that two joints belong to the same robotic arm. Then, based on the correlation degree, the target pose set that meets the preset conditions can be determined from multiple pose sets. The preset conditions are that any two poses in the pose set do not include the same joint, and the sum of the correlation degrees of the pose set is the maximum value among multiple pose sets. Among them, the correlation degree between any two joints can be determined according to the vectors of each pixel located between any two joints and the vector formed by any two joints.

[0145] Specifically, it is defined that A joint feature map is a feature map of joints of the same part. In this feature map, it includes the features of the same joint of multiple robotic arms. represents the k-th predicted position of the j-th joint. Ideally, a predicted position can represent a joint. Of course, in actual prediction, there may be interference in the model prediction result, that is, non-joint pixel points are recognized as joint pixel points. At the same time, it is defined that is used to represent whether two joints belong to the same robotic arm. For example, when it is 0, it means Do not belong to the same robotic arm, for example When it is 1, it means Belong to the same robotic arm, where j1 and j2 represent joints of two different parts, and k1 and k2 represent different joint prediction points in the joint feature map, which can be understood as joints of different robotic arms.

[0146] In order to obtain The correlation degree of belonging to the same robotic arm, it is necessary to use the connecting arm feature map, and the calculation formula is as follows:

[0147]

[0148] p(u) = (1 - u)d j1 + ud j2

[0149] Among them, E refers to L c Along the prediction point d j1 And d j2 The correlation degree between two points calculated by linear integral along the line segment between two points of d j1 And d j2 Indicates the expected value that two pixel points d j1 And d j2 Belong to the same robotic arm, and p(u) represents the pixel point between two pixel points d

[0150] For the J joint sets constructed from J joint feature maps, that is D j1 , D j2 , … and other J joint sets, the problem of finding the connected joints belonging to the same robotic arm can be transformed into a bipartite graph problem. Each point in the bipartite graph is a joint in the joint set, and the weight of each edge is E calculated by the above formula. The goal of the bipartite graph problem is to find an edge set of the bipartite graph where no two edges share a point and the sum of the weights is the largest. As Figure 15 Shown, there are joint sets 1 to 3. The dots with different grayscales represent joints in different joint sets. The connecting edges between the dots represent the relationship between two joints, and the weight of the connecting edge is E calculated by the above formula. Therefore, the goal is to find an edge set where no two edges share a dot and the sum of the weights of the connecting edges in this edge set is the largest among all edge sets. The obtained edge set is the target pose set. This problem can be solved using the maximum matching solution algorithm, such as the Hungarian algorithm (The Hungarian algorithm). Of course, other possible maximum matching solution algorithms can also be used to implement it, and the embodiments of this application do not limit this.

[0151] In specific implementation, for the J joint sets constructed from J joint feature maps, that is … and other J sets of joints. It is also possible to separately determine the joints belonging to the same robotic arm in two adjacent sets of joints, and then perform data integration to determine all the joints belonging to the same robotic arm. For example, for Figure 14 it is possible to separately find the joints belonging to the same robotic arm among joint A and joint B, and the joints belonging to the same robotic arm among joint B and joint C.

[0152] As Figure 16 shown, it is a schematic diagram of the posture of the combined robotic arm. Among them, through the combination of each joint and the connecting arm, the posture of each robotic arm can be obtained.

[0153] Step 1304: Respectively perform posture anomaly detection on each robotic arm according to the postures of the robotic arms in at least one robotic arm to obtain the posture detection results of each robotic arm.

[0154] In the embodiments of the present application, after obtaining the postures of each robotic arm, it is possible to perform posture anomaly detection on each robotic arm according to the postures of the robotic arms in at least one robotic arm to obtain the posture detection results of each robotic arm. Among them, the posture detection result indicates whether the posture of the robotic arm is abnormal. When the posture detection result indicates an abnormal posture, an alarm can be issued to repair the abnormality in a timely manner.

[0155] Specifically, an already trained anomaly detection model can be used to perform posture anomaly detection on each robotic arm. Among them, a posture distribution function can be learned by using a machine learning method with multiple posture training samples. Each posture sample is labeled as representing a normal or abnormal posture. According to the coordinates of the joint points constituting the posture of the robotic arm and the above-mentioned posture distribution function, the probability value of each robotic arm having a normal posture can be determined. Furthermore, the probability value is compared with a preset probability threshold to determine whether the posture of the robotic arm is abnormal. When it is determined that the probability value is less than the preset probability threshold, it can be determined that the posture of each robotic arm is abnormal. Otherwise, when the probability value is greater than or equal to the preset probability threshold, it can be determined that the posture of each robotic arm is normal.

[0156] Among them, the anomaly detection model can be, for example, a Gaussian distribution anomaly detection model. Taking the Gaussian distribution anomaly detection model as an example, this model is obtained through machine learning with a large number of posture training samples. Each posture sample can be a sample with a normal posture. Through model training, a Gaussian distribution function of the normal posture can be obtained.

[0157] Among them, the set of normal posture samples is represented as follows:

[0158] P = {p1, p2,... p n}

[0159] p1 = (x 11 , y 11 ), … (x 1j , y 1j )

[0160] …

[0161] p n = (x n1 , y n1 ), … (x nj , y nj )

[0162] Among them, P is the set of normal pose samples, p1, p2, ... p n represent each normal pose sample, where j is the number of joint types, (x ni , y ni ) represents the coordinates of the i-th joint included in the robotic arm in the n-th pose sample, and the value range of i is 1 to J.

[0163] Through the above pose sample set, its Gaussian distribution function can be calculated:

[0164]

[0165]

[0166] μ is the mean parameter of the Gaussian distribution function, and ∑ is the variance parameter of the Gaussian distribution function.

[0167] After obtaining the Gaussian distribution function, it can be used to predict whether the current pose of the robotic arm is abnormal. Specifically, the probability that the coordinates of each joint of the robotic arm are in a normal pose can be obtained through the following formula:

[0168]

[0169] Among them, P(p) represents the probability value that the pose of the robotic arm is normal. Furthermore, it can be determined whether the probability value is less than a preset probability threshold. For example, it can be set to 90% or 80%. When the calculated probability value is less than 90% or 80%, it can be determined that the pose of the robotic arm is abnormal. Of course, the setting of the threshold can be adjusted according to actual needs or based on experience.

[0170] Of course, abnormal pose samples can also be added to the pose training samples. Among them, normal pose samples and abnormal pose samples can be collected according to a preset ratio. For example, the number of normal pose samples is greater than 10 6 pieces, and the number of abnormal pose samples is greater than 10 3 pieces. Of course, the quantity can be adjusted according to actual needs.

[0171] Specifically, the postures of the robotic arms can be respectively determined whether they are abnormal by matching the postures of the robotic arms with the preset postures in the preset posture database. For a robotic arm, if the robotic arm can match a preset posture close to its own, it is determined that the posture of the robotic arm is normal, otherwise it is determined that the posture of the robotic arm is abnormal.

[0172] In summary, the robotic arm posture detection method provided by the embodiments of the present application uses a bottom-up algorithm to first extract joints and connecting arms, and then estimate postures, which can simultaneously estimate the postures of multiple robotic arms in the same captured image, greatly improving the efficiency. Moreover, when training the model, a large number of robotic arm picture samples are used for training, and the model can recognize various robotic arms, with strong versatility. Correspondingly, when configuring the image acquisition device, a smaller number of devices can be configured, which can meet the monitoring of the working conditions of the robotic arms while saving more hardware resources and costs.

[0173] Please refer to Figure 17 , based on the same inventive concept, the embodiments of the present application also provide a robotic arm posture detection device 170, including:

[0174] An image acquisition unit 1701, configured to acquire a to-be-detected image including at least one robotic arm;

[0175] An identification unit 1702, configured to acquire joint feature data of each joint and connecting arm feature data of each connecting arm included in at least one robotic arm in the to-be-detected image;

[0176] A posture combination unit 1703, configured to perform posture combination according to the joint feature data of each joint and the connecting arm feature data of each connecting arm to determine the postures of at least one robotic arm;

[0177] A posture abnormality detection unit 1704, configured to respectively perform posture abnormality detection on each robotic arm according to the postures of the robotic arms in at least one robotic arm to obtain a posture detection result of each robotic arm, where the posture detection result indicates whether the posture of the robotic arm is abnormal.

[0178] Optionally, the identification unit 1702 is specifically configured to:

[0179] Perform joint recognition on the to-be-detected image to obtain joint feature data; the joint feature data includes J joint feature maps corresponding to the joints of J parts, and each joint feature map is used to characterize the probability that each pixel point in the to-be-detected image is a joint pixel point;

[0180] Perform connecting arm recognition on the image to be detected to obtain connecting arm feature data, where the connecting arm feature data includes C connecting arm feature maps corresponding to C parts, and each connecting arm feature map is composed of vectors of each pixel point in the image to be detected. Each non-zero vector in each vector field indicates that the pixel point corresponding to the non-zero vector is a connecting arm pixel point.

[0181] Optionally, the recognition unit 1702 is specifically configured to:

[0182] Use the trained joint recognition model to perform joint recognition on the image to be detected to obtain joint feature data;

[0183] Among them, the joint recognition model is trained using multiple image training samples, and each image training sample labels the regions where each joint is located in the image training sample.

[0184] Optionally, the recognition unit 1702 is specifically configured to:

[0185] Extract features from the image to be detected to obtain the initial joint features of the image to be detected;

[0186] According to the initial joint features, respectively determine the probability that each pixel point on the image to be detected is a joint pixel point of each part, so as to obtain joint feature data.

[0187] Optionally, the recognition unit 1702 is specifically configured to:

[0188] Use the trained connecting arm recognition model to perform connecting arm recognition on the image to be detected to obtain connecting arm feature data;

[0189] Among them, the connecting arm recognition model is trained using multiple image training samples, and each image training sample labels the vectors of each connecting arm in the image training sample.

[0190] Optionally, the recognition unit 1702 is specifically configured to:

[0191] Extract features from the image to be detected to obtain the initial connecting arm features of the image to be detected;

[0192] According to the initial connecting arm features, respectively determine the vectors of each pixel point on the image to be detected, so as to obtain joint feature data.

[0193] Optionally, the pose combination unit 1703 is specifically configured to:

[0194] Perform pose combination according to J joint feature maps and C connecting arm feature maps to obtain the poses of at least one robotic arm.

[0195] Optionally, the pose combination unit 1703 is specifically configured to:

[0196] Construct J joint sets according to J joint feature maps, where each joint set is composed of joints of the same part of at least one robotic arm;

[0197] Construct multiple pose sets based on the J joint sets, where each pose set includes multiple poses formed by the J joints of the J joint sets, and any two joints among the J joints belong to different joint sets;

[0198] Determine the target pose set that is the pose of each robotic arm from the multiple pose sets.

[0199] Optionally, the pose combination unit 1703 is specifically configured to:

[0200] Determine the correlation degree between any two joints among the J joints that make up each pose according to C connecting arm feature maps, where the correlation degree is used to represent the probability that the two joints belong to the same robotic arm;

[0201] Based on the correlation degree, determine the target pose set that meets the preset conditions from the multiple pose sets, where the preset condition is that any two poses in the pose set do not include the same joints, and the sum of the correlation degrees of the pose set is the maximum value among the multiple pose sets.

[0202] Optionally, the pose combination unit 1703 is specifically configured to:

[0203] Determine the correlation degree between any two joints according to the vectors of each pixel located between any two joints and the vector formed by any two joints.

[0204] Optionally, the pose anomaly detection unit 1704 is specifically configured to:

[0205] Determine the probability value that the pose of each robotic arm is a normal pose according to the coordinates of the joint points that make up the pose of each robotic arm and the trained pose distribution function; the pose distribution function is obtained through machine learning using multiple pose training samples, and each pose sample is labeled as representing a normal or abnormal pose;

[0206] When it is determined that the probability value is less than the preset probability threshold, determine that the pose of each robotic arm is abnormal.

[0207] This device can be used to execute Figures 5 to 16 the method shown in the embodiments shown, therefore, for the functions that each functional module of this device can achieve, etc., reference can be made to Figures 5 to 16 the description of the embodiments shown, and details are not repeated here.

[0208] Please refer to Figure 18, Based on the same inventive concept, an embodiment of the present application also provides a computer device 180, which may include a memory 1801 and a processor 1802.

[0209] The memory 1801 is used to store a computer program executed by the processor 1802. The memory 1801 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the computer device, etc. The processor 1802 may be a central processing unit (CPU), or a digital processing unit, etc. In the embodiment of the present application, the specific connection medium between the above-mentioned memory 1801 and the processor 1802 is not limited. In the embodiment of the present application Figure 18 it is shown that the memory 1801 and the processor 1802 are connected through a bus 1803. The bus 1803 is represented by a thick line in Figure 18 For the connection manners between other components, only schematic illustrations are provided and are not limited thereto. The bus 1803 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 18 only a thick line is used to represent it in

[0210] but it does not mean that there is only one bus or one type of bus. The memory 1801 may be a volatile memory, such as a random-access memory (RAM); the memory 1801 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or the memory 1801 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1801 may be a combination of the above memories.

[0211] The processor 1802 is used to execute the method executed by the device in the embodiment as Figures 5 to 16 shown.

[0212] In some possible embodiments, aspects of the methods provided in this application may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the methods according to various exemplary embodiments of this application described above in this specification. For example, the computer device may execute the method performed by the device in the embodiment shown as Figures 5 to 16 shown in the embodiments.

[0213] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0214] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of this application.

[0215] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A method for detecting the posture of a robotic arm, characterized in that The method includes: Obtaining a to-be-detected picture including at least one robotic arm; Performing joint recognition on the to-be-detected picture to obtain joint feature data of each joint included in the at least one robotic arm in the to-be-detected picture; the joint feature data includes J joint feature maps corresponding to joints of J parts, and each joint feature map is used to represent the probability that each pixel point in the to-be-detected picture is a joint pixel point; Performing connecting arm recognition on the to-be-detected picture to obtain connecting arm feature data of each connecting arm included in the at least one robotic arm, the connecting arm feature data includes C connecting arm feature maps corresponding to connecting arms of C parts, and each connecting arm feature map is composed of vectors of each pixel point in the to-be-detected picture, and non-zero vectors in each vector field represent that the pixel points corresponding to the non-zero vectors are connecting arm pixel points; Performing pose combination according to the joint feature data of each joint and the connecting arm feature data of each connecting arm to determine the pose of the at least one robotic arm; Respectively performing pose anomaly detection on each robotic arm according to the pose of each robotic arm in the at least one robotic arm to obtain a pose detection result of each robotic arm, and the pose detection result indicates whether the pose of the robotic arm is abnormal.

2. The method according to claim 1, wherein Performing joint recognition on the to-be-detected picture to obtain joint feature data of each joint included in the at least one robotic arm in the to-be-detected picture, including: Using a trained joint recognition model to perform joint recognition on the to-be-detected picture to obtain the joint feature data; Wherein, the joint recognition model is trained by using a plurality of picture training samples, and each picture training sample is labeled with the area where each joint is located in the picture training sample.

3. The method according to claim 2, wherein The using the trained joint recognition model to perform joint recognition on the to-be-detected picture to obtain the joint feature data includes: Performing feature extraction on the to-be-detected picture to obtain initial joint features of the to-be-detected picture; According to the initial joint features, respectively determining the probability that each pixel point on the to-be-detected picture is a joint pixel point of joints of each part to obtain the joint feature data.

4. The method according to claim 1, characterized in that, Performing connecting arm recognition on the to-be-detected picture to obtain connecting arm feature data of each connecting arm included in the at least one robotic arm, including: Using a trained connecting arm recognition model to perform connecting arm recognition on the to-be-detected picture to obtain the connecting arm feature data; Wherein, the connecting arm recognition model is trained by using a plurality of picture training samples, and each picture training sample is labeled with the vectors of each connecting arm in the picture training sample.

5. The method according to claim 4, wherein The using the trained connecting arm recognition model to perform connecting arm recognition on the to-be-detected picture to obtain the connecting arm feature data includes: Performing feature extraction on the to-be-detected picture to obtain initial connecting arm features of the to-be-detected picture; According to the initial connecting arm features, respectively determining the vectors of each pixel point on the to-be-detected picture to obtain the joint feature data.

6. The method according to any one of claims 1 to 5, characterized in that The performing pose combination according to the joint feature data of each joint and the connecting arm feature data of each connecting arm to determine the pose of the at least one robotic arm includes: Performing pose combination based on the J joint feature maps and the C connecting arm feature maps to obtain the pose of the at least one robotic arm.

7. The method according to claim 6, wherein Performing pose combination based on the J joint feature maps and the C connecting arm feature maps to obtain the pose of the at least one robotic arm, including: Based on the J joint feature maps, constructing J joint sets, each joint set being composed of joints of the same part of the at least one robotic arm; Based on the J joint sets, constructing a plurality of pose sets, each pose set including a plurality of poses constituted by the J joints of the J joint sets, where any two joints among the J joints belong to different joint sets; Determining a target pose set that is the pose of each of the robotic arms from the plurality of pose sets.

8. The method according to claim 7, characterized in that, Determining a target pose set that is the pose of each of the robotic arms from the plurality of pose sets, including: According to the C connecting arm feature maps, determining the correlation degree between any two joints among the J joints constituting each pose, where the correlation degree is used to characterize the probability that two joints belong to the same robotic arm; Based on the correlation degree, determining the target pose set that meets the preset conditions from the plurality of pose sets, where the preset conditions are that any two poses in the pose set do not include the same joint, and the sum of the correlation degrees of the pose set is the maximum value among the plurality of pose sets.

9. The method according to claim 8, wherein According to the C connecting arm feature maps, determining the correlation degree between any two joints among the J joints constituting each pose, including: Determining the correlation degree between any two joints according to the vectors of each pixel located between the any two joints and the vector constituted by the any two joints.

10. The method according to claim 1, characterized in that, For each robotic arm, performing pose anomaly detection according to the pose of each robotic arm to obtain the pose detection result of each robotic arm, including: According to the coordinates of the joint points constituting the pose of each robotic arm and the trained pose distribution function, determining the probability value that the pose of each robotic arm is a normal pose; the pose distribution function is obtained through machine learning using a plurality of pose training samples, and each pose sample is labeled as representing a normal or abnormal pose; When it is determined that the probability value is less than the preset probability threshold, determining that the pose of each robotic arm is abnormal.

11. A robotic arm attitude detection device, characterized in that, Including: An image acquisition unit for acquiring a to-be-detected image including at least one robotic arm; An identification unit for performing joint identification on the to-be-detected image to obtain the joint feature data of each joint included in the at least one robotic arm in the to-be-detected image; The joint feature data includes J joint feature maps corresponding to the joints of J parts, and each joint feature map is used to characterize the probability that each pixel point in the to-be-detected picture is a joint pixel point; and, by performing connecting arm recognition on the to-be-detected picture, the connecting arm feature data of each connecting arm included in the at least one robotic arm is obtained, and the connecting arm feature data includes C connecting arm feature maps corresponding to the connecting arms of C parts, and each connecting arm feature map is composed of vectors of each pixel point in the to-be-detected picture, and non-zero vectors in each vector field indicate that the pixel points corresponding to the non-zero vectors are connecting arm pixel points; A pose combination unit, configured to perform pose combination according to the joint feature data of each joint and the connecting arm feature data of each connecting arm to determine the pose of the at least one robotic arm; A pose anomaly detection unit, configured to perform pose anomaly detection on each robotic arm respectively according to the poses of the robotic arms in the at least one robotic arm to obtain the pose detection results of the robotic arms, and the pose detection results indicate whether the poses of the robotic arms are abnormal.

12. A robotic arm attitude monitoring system, characterized in that, Comprising: At least one image acquisition device, and each image acquisition device is configured to acquire an image or a video stream including at least one robotic arm; A robotic arm pose detection device, configured to use the method according to any one of claims 1 to 10 to obtain the pose detection results of the robotic arms in the video frames of the image or the video stream, and trigger an alarm when there is a pose detection result indicating that the robotic arm pose is abnormal.

13. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

14. A computer storage medium, on which computer program instructions are stored, wherein When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Device and method for detecting working states of sensing equipment on mechanical arm, mechanical arm and medical robot

    CN109109018A

  • Auxiliary photographing equipment used for movement disorder symptom analysis, control method and device

    CN110561399A