Method, device and electronic device for robotic arm interaction based on visual recognition

Through a visual recognition-based method, the image acquisition data set is used to control the movement of the robotic arm in real time, which solves the problem of lack of interactivity in the robotic arm dance, and achieves real-time interaction and ornamental improvement with the external environment.

CN119635651BActive Publication Date: 2025-07-08BEIJING APAILANG CREATIVITY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411966136.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-07-08
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In the prior art, the dance performance of the robotic arm lacks interactivity and cannot adjust the movements in real time according to changes in the external environment, resulting in poor viewing experience.

Method used

Through the image acquisition data set, the target robotic arm action mapping indication data is determined, the robotic arm performs actions in real time, and the visual recognition technology is used to improve interactivity.

Benefits of technology

Real-time interaction between the robotic arm and the external environment is realized, the viewing experience and interactivity are improved, and the movements can be adjusted according to different target objects, enhancing the flexibility and artistry of the performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119635651B_ABST
    Figure CN119635651B_ABST
Patent Text Reader

Abstract

The present invention provides a robotic arm interaction method, device, and electronic device based on visual recognition. The method includes: obtaining an image acquisition data set, where the image acquisition data set includes m image acquisition data, and one image acquisition data includes images collected by each of n cameras at one image acquisition moment, and both m and n are positive integers; determining target robotic arm action mapping indication data corresponding to a target object based on the image acquisition data set, where the target object is the object with the largest object area indication value among at least one object, and the object area indication value of one object is used to indicate the size of the area occupied by the corresponding object; determining a target robotic arm action based on the target robotic arm action mapping indication data, and controlling the robotic arm to execute the target robotic arm action. Embodiments of the present invention can improve the interactivity of the robotic arm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and robot control, and in particular, to a robotic arm interaction method, device, and electronic device based on visual recognition. Background Art

[0002] With the continuous development of robot technology, robots are gradually moving from traditional industrial application fields to fields closer to people's lives, such as services and entertainment. Currently, robot dancing, as a novel performance form, has attracted more and more attention; among them, the robotic arm, as a common basic form of a robot, is used to realize robot dancing. However, in related technologies, the implementation of robot dancing mainly relies on pre-programming to control the robot to complete a series of fixed actions, and it is unable to adjust the actions in real time according to changes in the external environment, lacking interactivity and making it difficult to bring a better viewing experience to the audience. Based on this, there is currently no good solution to how to improve the interactivity of the robotic arm. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a robotic arm interaction method, device, and electronic device based on visual recognition to solve problems such as the lack of interactivity of the robotic arm involved in related technologies; that is, embodiments of the present invention can determine the target robotic arm action in real time through an image acquisition data set and control the robotic arm to execute the target robotic arm action, so as to adjust the actions of the robotic arm according to changes in the external environment, thereby effectively improving the interactivity of the robotic arm and bringing a better viewing experience to the audience.

[0004] According to one aspect of the embodiments of the present invention, a robotic arm interaction method based on visual recognition is provided, and the method includes:

[0005] Obtain an image acquisition data set, where the image acquisition data set includes m image acquisition data, and one image acquisition data includes images collected by each of n cameras at one image acquisition moment, and both m and n are positive integers;

[0006] Based on the image acquisition data set, determine target robotic arm action mapping indication data corresponding to a target object; where the target object is the object with the largest object area indication value among at least one object, and the object area indication value of an object is used to indicate the size of the area occupied by the corresponding object;

[0007] Based on the target robotic arm action mapping indication data, determine a target robotic arm action, and control the robotic arm to execute the target robotic arm action.

[0008] According to another aspect of the embodiments of the present invention, a robotic arm interaction device based on visual recognition is provided, and the device includes:

[0009] An image acquisition module, configured to acquire an image acquisition data set, where the image acquisition data set includes m image acquisition data, and one image acquisition data includes images acquired by each of n cameras at one image acquisition moment, and both m and n are positive integers;

[0010] An image analysis and processing module, configured to determine target robotic arm action mapping indication data corresponding to a target object based on the image acquisition data set; where the target object is the object with the largest object area indication value among at least one object, and the object area indication value of one object is used to indicate the size of the area occupied by the corresponding object;

[0011] An action generation module, configured to determine a target robotic arm action based on the target robotic arm action mapping indication data;

[0012] A robotic arm control module, configured to control the robotic arm to execute the target robotic arm action.

[0013] According to another aspect of the embodiments of the present invention, an electronic device is provided, where the electronic device includes a processor and a memory storing a program, and where the program includes instructions that, when executed by the processor, cause the processor to execute the method mentioned above.

[0014] According to another aspect of the embodiments of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to cause a computer to execute the method mentioned above.

[0015] After the image acquisition data set is acquired in the embodiments of the present invention, based on the image acquisition data set, target robotic arm action mapping indication data corresponding to a target object can be determined. The image acquisition data set includes m image acquisition data, and one image acquisition data includes images acquired by each of n cameras at one image acquisition moment, and both m and n are positive integers; where the target object is the object with the largest object area indication value among at least one object, and the object area indication value of one object is used to indicate the size of the area occupied by the corresponding object. Based on this, a target robotic arm action can be determined based on the target robotic arm action mapping indication data, and the robotic arm can be controlled to execute the target robotic arm action. It can be seen that the embodiments of the present invention can determine the target robotic arm action in real time through the image acquisition data set, and control the robotic arm to execute the target robotic arm action, so as to adjust the action of the robotic arm according to the changes in the external environment, that is, a robotic arm interaction method based on visual recognition can be realized, thereby effectively improving the interactivity of the robotic arm and bringing a better viewing experience to the audience. Description of the Drawings

[0016] In the following description of exemplary embodiments with reference to the drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:

[0017] Figure 1 Shows a schematic flowchart of a robotic arm interaction method based on visual recognition according to an exemplary embodiment of the present invention;

[0018] Figure 2 Shows a schematic flowchart of another robotic arm interaction method based on visual recognition according to an exemplary embodiment of the present invention;

[0019] Figure 3 Shows a schematic block diagram of a robotic arm interaction device based on visual recognition according to an exemplary embodiment of the present invention;

[0020] Figure 4 Shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present invention. Detailed implementation manners

[0021] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0022] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0023] As used herein, the term "comprising" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0024] It should be noted that the modifications of "one" and "plural" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0026] It should be noted that the execution subject of the robotic arm interaction method based on visual recognition provided in the embodiments of the present invention can be one or more electronic devices, and the embodiments of the present invention do not limit this; among them, the electronic device can be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices, and at least one terminal and at least one server are included in the multiple electronic devices, the robotic arm interaction method based on visual recognition provided in the embodiments of the present invention can be jointly executed by the terminal and the server. Correspondingly, the terminals mentioned here can include, but are not limited to: robots, tablet computers, laptop computers, desktop computers, and so on. The servers mentioned here can be independent physical servers, or server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, and so on. Exemplarily, when the execution subject of the robotic arm interaction method based on visual recognition is one electronic device (such as a robot), all steps of the robotic arm interaction method based on visual recognition can be completed by this electronic device; when the execution subject of the robotic arm interaction method based on visual recognition is multiple electronic devices (such as including a robot and an electronic device other than the robot), the various electronic devices can interact. For example, certain data determined by an electronic device other than the electronic device (such as a robot) used to control the robotic arm to execute actions can be transmitted to the electronic device used to control the robotic arm to execute actions, so that the electronic device used to control the robotic arm to execute actions can execute subsequent operation steps until the robotic arm is controlled to execute corresponding actions, and so on; the embodiments of the present invention do not limit this.

[0027] Based on the above description, the embodiments of the present invention propose a robotic arm interaction method based on visual recognition, and this robotic arm interaction method based on visual recognition can be executed by one or more of the above-mentioned electronic devices. For the sake of elaboration, hereinafter, it will be described by taking one electronic device (such as a robot) executing this robotic arm interaction method based on visual recognition as an example; as Figure 1 shown, this robotic arm interaction method based on visual recognition may include the following steps S101 - S103:

[0028] S101. Obtain an image acquisition data set, where the image acquisition data set includes m pieces of image acquisition data, and one piece of image acquisition data includes images acquired by each of the n cameras at one image acquisition moment. Both m and n are positive integers.

[0029] Among them, the above n cameras can be used for image acquisition; optionally, the electronic device may include n cameras, that is, n cameras can be installed on the electronic device. At this time, the electronic device may include an image acquisition module to acquire images through the image acquisition module; or, the electronic device can establish a connection with the n cameras (such as a wired or wireless connection, etc.), then the n cameras can transmit the acquired images to the electronic device so that the electronic device can obtain the images acquired by the n cameras, and so on; the embodiments of the present invention do not limit this. It should be noted that the time among the n cameras is synchronized, so that in the subsequent image processing process, the images acquired by each camera at the same acquisition moment can be determined; optionally, the embodiments of the present invention can be synchronized by hardware. In this case, it is preferable to use a hardware synchronization device to ensure that multiple cameras acquire images at the same moment. For example, through an external trigger signal, the acquisition clocks of all cameras are synchronized so that they can start and end image acquisition at the same time; or, it can be synchronized by software. In this case, a time stamp can be added to the image data acquired by each camera to record the image acquisition time, and then in the subsequent processing process, the images are aligned according to the time stamps, and so on; the embodiments of the present invention do not limit this.

[0030] Optionally, the acquisition method of the image acquisition data set may include but is not limited to at least one of the following:

[0031] The first acquisition method: Each piece of image acquisition data in the m pieces of image acquisition data (that is, the image acquisition data at m image acquisition moments) is the image acquisition data newly acquired by the n cameras, that is, newly acquired image acquisition data; in this case, the electronic device can add the images acquired by the n cameras at each of the m image acquisition moments to the image acquisition data set to obtain the image acquisition data set. At this time, the image acquisition data in the image acquisition data set are all newly acquired image acquisition data.

[0032] The second acquisition method: The m image acquisition data includes at least one real-time acquired image acquisition data and at least one historical image acquisition data; in this case, the electronic device can acquire the images acquired by the n cameras at each real-time image acquisition moment among at least one real-time image acquisition moment (i.e., newly acquired images), and determine at least one historical image acquisition data that is closest to the current system time from the historical data (which may include the image acquisition data acquired by the n cameras before), so as to add at least one historical image acquisition data and at least one real-time acquired image data to the image acquisition data set respectively, so as to realize the acquisition of the image acquisition data set, and so on.

[0033] When the value of n is greater than 1, the electronic device needs to perform spatial calibration and data fusion, so as to convert the coordinates in each camera coordinate system into the integrated world coordinates and image coordinates. Optionally, the internal parameters of each camera can be calibrated separately to obtain its internal parameters. For example, taking the i-th camera among the n cameras as an example, the internal parameters may include the focal length (f i ), the optical center coordinates and the pixel size etc., which can be completed by the traditional checkerboard calibration method; using these internal parameters, the image coordinates can be converted into the coordinates in the camera coordinate system to prepare for subsequent triangulation. Optionally, the electronic device can also determine the relative position relationship between the n cameras, that is, the external parameters; through precise measurement tools and specific calibration processes, the rotation matrix (R i ) and the translation vector (T i ) of each camera relative to a common reference coordinate system (such as the world coordinate system) can be obtained. For example, this step can be implemented using an extended form of Zhang Zhengyou calibration method or a multi-camera calibration method based on a special calibration object to ensure accurate relative position information, so as to accurately estimate the position of the human face in the three-dimensional space through triangulation subsequently.

[0034] Optionally, for the camera C i , the internal parameter matrix K i can be defined as shown in Formula 1.1:

[0035]

[0036] Based on this, assuming a point D(X w , Y w , Z w ) in the world coordinate system, its coordinates D i (X i , Y i , Z i , Z i ) in the camera C i coordinate system can be calculated by Formula 1.2:

[0037]

[0038] Then, formula 1.3 can be adopted to convert the coordinates in the camera coordinate system into image coordinates (u i , v i , i ) through the intrinsic parameter matrix K:

[0039]

[0040] S102. Based on the image acquisition dataset, determine the target robotic arm motion mapping indication data corresponding to the target object; wherein, the target object is the object with the largest object region indication value among at least one object, and the object region indication value of an object is used to indicate the size of the region occupied by the corresponding object.

[0041] Optionally, an object can be a person or an item, and the embodiments of the present invention do not limit this. Optionally, when an object is a person, the object region indication value of the object can be the percentage of the body region of the object in the entire image (such as the total number of pixels in the body region of the object divided by the total number of pixels in the image where it is located), or the percentage of all regions of the object in the entire image. When the sizes of the images collected by each camera are the same (i.e., the total number of pixels in the images is the same), it can also be the total number of pixels in the body region of the object, etc.; the embodiments of the present invention do not limit this. Optionally, when an object is an item, the object region indication value of the object can be the percentage of the object region (i.e., the region where the object is located) in the entire image (i.e., the percentage of the object region in the corresponding image, such as the total number of pixels in the object region divided by the total number of pixels in the image where the object is located). When the sizes of the images collected by each camera are the same, it can also be the total number of pixels in the object region, etc.; the embodiments of the present invention do not limit this. Optionally, when the object region indication value of an object is the percentage of the corresponding region of the corresponding object in the image, the object region indication value of an object can also be referred to as the visual area percentage of the corresponding object.

[0042] Based on this, the object region indication value of an object is positively correlated with the size of the region occupied by the corresponding object (i.e., the region occupied by the corresponding object in the corresponding image); that is to say, the larger the object region indication value of an object, the larger the region occupied by the object, and the smaller the object region indication value of an object, the smaller the region occupied by the object.

[0043] Correspondingly, the electronic device can determine the object with the largest occupied area from at least one object, and use the object with the largest occupied area as the target object; that is to say, the electronic device can select the object with the largest object area indication value from at least one object as the target object. Optionally, the at least one object may include all objects in each image of the image acquisition dataset; or, the at least one object may include objects that exist in the image acquisition data at consecutive multiple image acquisition times in the image acquisition dataset; or, the at least one object may include all objects in the image acquisition data at any one of the m image acquisition times, and so on; the embodiments of the present invention do not limit this. Optionally, the number of target objects may be one or more, and the embodiments of the present invention do not limit this; for example, the electronic device may also use all objects in the image acquisition data at each image acquisition time as at least one object to determine the target object respectively; in this case, the target object at each image acquisition time may be determined respectively, and so on. Among them, the specific implementation manner of determining the target object can be seen as shown below, and the embodiments of the present invention will not elaborate here.

[0044] S103. Based on the target robotic arm motion mapping indication data, determine the target robotic arm motion, and control the robotic arm to execute the target robotic arm motion.

[0045] Optionally, the number of target robotic arm motions may be one or more, and the embodiments of the present invention do not limit this; among them, one target robotic arm motion may be determined based on one target robotic arm motion mapping indication data. Then, when the target robotic arm motion mapping indication data changes, that is, when the number of target robotic arm motion mapping indication data is multiple, multiple target robotic arm motions can be determined, that is, the target robotic arm motion corresponding to each target robotic arm motion mapping indication data can be determined respectively, and so on.

[0046] In an embodiment of the present invention, after obtaining an image acquisition data set, based on the image acquisition data set, target robotic arm action mapping indication data corresponding to a target object can be determined. The image acquisition data set includes m image acquisition data. One image acquisition data includes images acquired by each of the n cameras at one image acquisition moment. Both m and n are positive integers. Among them, the target object is the object with the largest object region indication value among at least one object. The object region indication value of an object is used to indicate the size of the region occupied by the corresponding object. Based on this, the target robotic arm action can be determined based on the target robotic arm action mapping indication data, and the robotic arm can be controlled to execute the target robotic arm action. It can be seen that the embodiment of the present invention can determine the target robotic arm action in real time through the image acquisition data set, and control the robotic arm to execute the target robotic arm action, so as to adjust the action of the robotic arm according to the change of the external environment, that is, a robotic arm interaction method based on visual recognition can be realized, thereby effectively improving the interactivity of the robotic arm and bringing a better viewing experience to the audience.

[0047] Based on the above description, another robotic arm interaction method based on visual recognition is proposed in an embodiment of the present invention. Correspondingly, this robotic arm interaction method based on visual recognition can be executed by one or more of the above-mentioned electronic devices. For the sake of convenience of explanation, in the following, it is taken as an example that an electronic device (such as a robotic arm) executes this robotic arm interaction method based on visual recognition; please refer to Figure 2 , this robotic arm interaction method based on visual recognition may include the following steps S201 - S205:

[0048] S201, obtain an image acquisition data set, where the image acquisition data set includes m image acquisition data. One image acquisition data includes images acquired by each of the n cameras at one image acquisition moment.

[0049] S202, perform human detection on each image in the image acquisition data set respectively to obtain human detection data for each image. The human detection data for one image includes the human segmentation mask image of the corresponding image and the human key point indication data for each person in the corresponding image. The human segmentation mask image of one image includes the segmentation masks for each person in the corresponding image.

[0050] Optionally, the electronic device can use algorithms such as PoseNet or MoveNet (a technical solution for pose detection) to perform human detection on each image in the image acquisition data set respectively. The embodiment of the present invention does not limit this. Among them, the PoseNet algorithm can not only generate a specific segmentation mask of a person (distinguishing different parts of the human body, that is, the masks of different parts of the human body are different), which is used to calculate the object region indication value, such as the proportion of the number of pixels of the human body region in the image; but also can identify the key points of the human body for human pose recognition.

[0051] Optionally, the human key-point indication data of a person may include the position information of the human key points of the corresponding person, so that the changes in the coordinates of these key points can be analyzed to obtain the action information of the person, such as the inclination angle of the body, the movement direction and amplitude of the limbs (i.e., the angle between the arm and the torso), etc.

[0052] S203. Based on the human segmentation mask images of each image, determine the target object; wherein, the target object is the object with the largest object area indication value among at least one object, and the object area indication value of an object is used to indicate the size of the area occupied by the corresponding object.

[0053] Among them, the human segmentation mask image of an image may include the segmentation masks of each person in the corresponding image; based on this, for any person in the image acquisition dataset (i.e., any person among all the persons indicated by the image acquisition dataset), the object area indication value of any person can be determined based on the segmentation mask of any person, and a person can be an object. Exemplarily, when the object area indication value of a person is the percentage of the image occupied by the body area of the corresponding person, the number of pixels in the body area of any person can be determined based on the segmentation mask of any person, and then the ratio between the number of pixels in the body area of any person and the number of pixels in the image where any person is located can be used as the object area indication value of any person, and so on.

[0054] Optionally, the electronic device may also perform item detection on each image in the image acquisition dataset to obtain the item detection data of each image, and the item detection data of an image includes the item segmentation mask image of the corresponding image. Among them, the item segmentation mask image of an image may include the segmentation masks of each item in the corresponding image; correspondingly, for any item in the image acquisition dataset (i.e., any item among all the items indicated by the image acquisition dataset), the object area indication value of any item can be determined based on the segmentation mask of any item, and an item (also referred to as an object) can be an object; Exemplarily, when the object area indication value of an item is the percentage of the image occupied by the corresponding item area, the number of pixels in the item area of any item can be determined based on the segmentation mask of any item, and then the ratio between the number of pixels in the item area of any item and the number of pixels in the image where any item is located can be used as the object area indication value of any item, and so on. Optionally, when performing item detection on any image, the item semantic information of each item in any image can also be obtained (which can be used to indicate the category of the item, such as a cat, a dog, etc.).

[0055] Optionally, the electronic device may adopt an object recognition deep learning model to perform object detection on each image respectively, so as to identify the objects in the image and generate a segmentation mask of the corresponding objects. Optionally, the object recognition deep learning model may be any pre-trained object recognition deep learning model, and the embodiments of the present invention do not limit this; for example, the object recognition deep learning model may be a SAM (Segment Anything Model) model (an image segmentation task model), and so on.

[0056] Based on this, when determining the target object based on the human segmentation mask images of each image, the target object can be determined based on the human segmentation mask images and object segmentation mask images of each image, and the target object is a person or an object. Optionally, for any object in any image in the image acquisition dataset (that is, any object indicated by any image), when any object is a person, the segmentation mask of any object can be determined from the human segmentation mask image of the image where any object is located, and when any object is an object, the segmentation mask of any object can be determined from the object segmentation mask image of the image where any object is located. Furthermore, the object region indication value of any object can be determined based on the segmentation mask of any object.

[0057] In one implementation, when determining the target object based on the human segmentation mask images and object segmentation mask images of each image, the electronic device can respectively determine the object region indication value of each object in each image based on the human segmentation mask images and object segmentation mask images of each image; then, based on the object region indication value of each object in each image, the pending object with the largest object region indication value at each image acquisition moment among the m image acquisition moments can be determined respectively (the pending object with the largest object region indication value at one image acquisition moment can be the object with the largest object region indication value among all the objects indicated by the images acquired by each camera at the corresponding image acquisition moment), so as to determine the target object from the m pending objects by using the voting method. At this time, the target object can be the pending object with the largest number of the same pending objects among the m pending objects. Optionally, when both objects are objects, whether they are the same object can be judged by the object semantic information of the two objects, and when both objects are people, whether they are the same person can be judged based on the similarity of the face regions of the two objects, and so on; the embodiments of the present invention do not limit this.

[0058] In another implementation, for any one of the m image acquisition moments, the electronic device may first determine a fusion object set at any one image acquisition moment from all the objects in the images acquired by each camera at any one image acquisition moment. One fusion object in the fusion object set may correspond to one or more objects among all the objects in the images acquired by each camera at any one image acquisition moment. That is to say, the same object (i.e., the same object) among all the objects in the images acquired by each camera at any one image acquisition moment can be used as one fusion object. For example, it can be determined whether it is the same object by item semantic information or face area, etc. Then, for any one fusion object in the fusion object set, the object area indication value of any one fusion object can be determined according to the weights of the cameras that acquired the images of each object corresponding to any one fusion object and the object area indication values of each object corresponding to any one fusion object. In this case, methods such as weighted average method or Bayesian inference can be used to determine the object area indication value of any one fusion object. Further, after obtaining the fusion object sets at each image acquisition moment, the fusion object with the largest object area indication value can be determined from the fusion object sets at each image acquisition moment as the pending object at the corresponding image acquisition moment, so as to determine m pending objects, and then the target object can be determined from the m pending objects by voting method.

[0059] In yet another implementation, the electronic device may also use the pending objects at each image acquisition moment as the target objects at the corresponding image acquisition moments respectively, that is, they can be used as one target object respectively. At this time, the number of target objects can be m; or the pending object at the m-th image acquisition moment can be used as the target object. At this time, the target object can be determined in the environment at the nearest moment, etc.; the embodiments of the present invention do not limit this.

[0060] Correspondingly, when the target object is a person, the electronic device may trigger and execute the following to determine the target robotic arm motion mapping indication data corresponding to the target object based on the person key point indication data of the target object; when the target object is an item, the item semantic information of the target object can be determined, and the item semantic information of the target object can be added to the target robotic arm motion mapping indication data corresponding to the target object. One item semantic information corresponds to one robotic arm motion.

[0061] Optionally, when the number of target objects is multiple, the target robotic arm motion mapping indication data corresponding to each target object among the multiple target objects can be determined respectively; optionally, the target robotic arm motion mapping indication data corresponding to different target objects can be the same or different, and the embodiments of the present invention do not limit this. For the convenience of description, one target object will be used as an example for subsequent description.

[0062] S204. Determine the target robotic arm motion mapping indication data corresponding to the target object based on the human key point indication data of the target object.

[0063] In an embodiment of the present invention, the electronic device can determine the human body posture motion of the target object based on the human key point indication data of the target object, and add the human body posture motion identifier of the human body posture motion of the target object to the target robotic arm motion mapping indication data corresponding to the target object for determining the target robotic arm motion. Optionally, the number of the human key point indication data of the target object can be one or more; optionally, the electronic device can determine the human key point indication data of all objects that are the same object as the target object from the human key point indication data of each person in each image (persons with a similarity greater than a preset similarity threshold in the face regions of different images can be the same person) to obtain the human key point indication data of the target object at each image acquisition moment among one or more image acquisition moments of the target object (i.e., multiple human key point indication data of the target object); where one human key point indication data can be used to indicate one human body posture. Optionally, one or more human key point indication data (i.e., human key point coordinates) of the target object can be represented as a sequence of human key point indication data of the target object {P j}, where j represents the serial number of the jth human body posture. Optionally, the electronic device can determine the human body posture motion of the target object based on the change situation of multiple human key point indication data of the target object, such as waving horizontally or vertically, etc.; or, a specific human body posture in the human body postures indicated by one or more human key point indication data of the target object can be used as the human body posture motion of the target object, etc.; the embodiments of the present invention do not limit this. Optionally, the specific human body posture can be any human body posture in the human body posture set; optionally, the human body posture set can be set according to experience or actual requirements, and the embodiments of the present invention do not limit this. Based on this, the human body posture motion of the target object can be static or dynamic, and the embodiments of the present invention do not limit this.

[0064] Optionally, the electronic device may also determine the action guidance data of the target object based on the character key point indication data of the target object, and add the action guidance data to the target robotic arm action mapping indication data; wherein the action guidance data includes at least one of the following: the character's arm amplitude (i.e., the angle between the arm and the body trunk), the character's action speed, and the character's action frequency, and the action guidance data supports the determination of robotic arm action parameter data of the robotic arm, so that the robotic arm performs the target robotic arm action according to the robotic arm action parameter data. It should be noted that the embodiments of the present invention do not limit the specific method for determining the action guidance data; for example, based on the character key point indication data of the target object, the angle between the upper arm and the torso of the target object can be determined as the character arm amplitude of the target object, or the angle between the lower arm and the torso of the target object can be determined, so as to perform weighted summation on the angle between the upper arm and the torso and the angle between the lower arm and the torso to obtain the character arm amplitude of the target object, or when the character arm amplitudes indicated by each character key point indication data of the target object are different, the maximum value of the character arm amplitudes can be used as the character arm amplitude of the target object, or the character arm amplitudes indicated by each character key point indication data of the target object can be respectively used as the character arm amplitude of the target object at the corresponding image acquisition time. For example, based on the target object's key point indication data, the maximum position change distance and position change time (i.e., the time it takes to generate the maximum position change distance) of the target key point with the largest position information change can be determined, and the ratio between the maximum position change distance and the position change time of the target key point can be used as the character's action speed, or the ratio between the maximum position change distance and the corresponding position change time of the designated key point can be used as the character's action speed, etc.; for example, the maximum value of the number of action changes can be used as the target object's character action frequency (i.e., action change frequency), such as when swinging left and swinging right alternately, the change from swinging left to swinging right or from swinging right to swinging left can be one action change, etc. Optionally, the designated key point can be set according to experience or according to actual needs, and the embodiment of the present invention is not limited to this.

[0065] S205 , determining a target robotic arm motion based on the target robotic arm motion mapping indication data, and controlling the robotic arm to execute the target robotic arm motion.

[0066] In an embodiment of the present invention, when the target object is a person, the target robotic arm motion mapping indication data includes a human body posture motion identifier; when the target object is an item, the target robotic arm motion mapping indication data includes item semantic information. Based on this, when determining the target robotic arm motion based on the target robotic arm motion mapping indication data, when the target robotic arm motion mapping indication data includes a human body posture motion identifier, based on the human body posture motion identifier in the target robotic arm motion mapping indication data, a target robotic arm motion identifier that matches the human body posture motion identifier in the target robotic arm motion mapping indication data is determined from a first set of robotic arm motions. The target robotic arm motion identifier is used to indicate the target robotic arm motion to achieve the determination of the target robotic arm motion. Among them, a human body posture motion identifier can be used to indicate a human body posture motion; optionally, a human body posture motion identifier can be any human body posture motion identifier in a set of human body posture motion identifiers; optionally, a human body posture motion identifier can be the name of a human body posture motion or the number of a human body posture motion, etc., and the embodiments of the present invention do not limit this. Correspondingly, a robotic arm motion identifier can be used to indicate a robotic arm motion; optionally, a robotic arm motion identifier can be the name of a robotic arm motion or the number of a robotic arm motion, etc., and the embodiments of the present invention do not limit this. Optionally, the set of human body posture motion identifiers can be set according to experience or actual needs, and the embodiments of the present invention do not limit this.

[0067] Optionally, the first set of robotic arm motions may include robotic arm motion identifiers (i.e., first robotic arm motion identifiers) of each of at least one first robotic arm motion; optionally, the first set of robotic arm motions can be pre-compiled according to experience or actual needs, and the embodiments of the present invention do not limit this. Optionally, there may be a mapping relationship between the set of human body posture motion identifiers and the first set of robotic arm motions, that is, any human body posture motion identifier in the set of human body posture motion identifiers can match a first robotic arm motion identifier in the first set of robotic arm motions, so that a target robotic arm motion identifier that matches the human body posture motion identifier in the target robotic arm motion mapping indication data can be determined from the first set of robotic arm motions.

[0068] Correspondingly, when the target robotic arm motion mapping indication data includes an item semantic information, based on the item semantic information in the target robotic arm motion mapping indication data, determine a target robotic arm motion identifier that matches the item semantic information in the target robotic arm motion mapping indication data from the second robotic arm motion set. Optionally, the second robotic arm motion set may include the robotic arm motion identifiers (i.e., the second robotic arm motion identifiers) of each of at least one second robotic arm motion; Optionally, the second robotic arm motion set may be pre-compiled according to experience or actual requirements, and the embodiments of the present invention do not limit this. Optionally, each item semantic information in the item semantic information set may correspond to (i.e., match) a second robotic arm motion identifier in the second robotic arm motion set, that is to say, there is a mapping relationship between the item semantic information set and the second robotic arm motion set; Optionally, the item semantic information set may be set according to experience or actual requirements, and the embodiments of the present invention do not limit this; wherein, the item semantic information in the target robotic arm motion mapping indication data (i.e., the item semantic information of the target object) may be any item semantic information in the item semantic information set. Optionally, the second robotic arm motion set may be the same as the first robotic arm motion set (in this case, a robotic arm motion set may have a mapping relationship with the item semantic information set and the human body posture motion identifier set respectively), or may be different from the first robotic arm motion set, and the embodiments of the present invention do not limit this. Exemplarily, when the item semantic information is a cat or a toy cat, the robotic arm may make gentle motions imitating a cat; when the item semantic information is a robot or a toy robot, the robotic arm may make jerky motions imitating a robot, and so on.

[0069] Optionally, when the target object is a person, the target robotic arm motion mapping indication data may include motion guidance data for the target object, and the motion guidance data includes at least one of the following: the amplitude of the person's arm, the speed of the person's motion, and the frequency of the person's motion; in this case, the electronic device may also determine the pending robotic arm motion parameter data of the robotic arm based on the motion guidance data, and the pending robotic arm motion parameter data includes at least one of the following: the pending amplitude of the robotic arm motion, the pending frequency of the robotic arm motion, and the pending speed of the robotic arm motion; where the pending amplitude of the robotic arm motion is determined based on the amplitude of the person's arm, the pending frequency of the robotic arm motion is determined based on the frequency of the person's motion, and the pending speed of the robotic arm motion is determined based on the speed of the person's motion; then correspondingly, the robotic arm motion parameter data of the robotic arm may be determined based on the pending robotic arm motion parameter data. Based on this, when controlling the robotic arm to execute the target robotic arm motion, the robotic arm may be controlled to execute the target robotic arm motion according to the robotic arm motion parameter data; that is to say, after determining the robotic arm motion parameter data, the robotic arm may be controlled to execute the target robotic arm motion according to the robotic arm motion parameter data, such as making the speed, frequency, and amplitude of the target robotic arm motion the same as the corresponding data indicated by the robotic arm motion parameter data. Among them, the target robotic arm motion may also be referred to as the robotic arm motion content.

[0070] Optionally, when determining the pending robotic arm motion parameter data of the robotic arm based on the motion guidance data, the electronic device may determine the pending amplitude of the robotic arm motion based on the amplitude of the person's arm and the original amplitude of the robotic arm motion, such as the pending amplitude of the robotic arm motion B θ = B0×(θ / 90), where B0 represents the original amplitude of the robotic arm motion (i.e., the original amplitude of the robotic arm), θ represents the amplitude of the person's arm, and the greater the amplitude of the person's arm, the greater the pending amplitude of the robotic arm motion; and / or, the pending frequency of the robotic arm motion may be determined based on the frequency of the person's motion and the original frequency of the robotic arm motion, such as the pending frequency of the robotic arm motion F k = F0×(F pk / F p0 ), where F0 represents the original frequency of the robotic arm motion, F pk represents the frequency of the person's motion, F p0 represents the standard frequency of the person waving the arm, and the higher the frequency of the person's motion, the higher the pending frequency of the robotic arm motion; and / or, the pending speed of the robotic arm motion may be determined based on the speed of the person's motion and the original moving speed of the robotic arm, such as the pending speed of the robotic arm motion V k = V0×(V pk / V p0 ), where V0 represents the original moving speed of the robotic arm, V pk represents the speed of the person's motion, V p0It represents the standard movement speed of a person. The faster the person's action speed, the faster the to-be-determined robotic arm's action speed, and so on. Optionally, the standard frequency of the person waving their arm, the standard movement speed of the person, etc. can all be set according to experience or actual requirements, and the embodiments of the present invention do not limit this.

[0071] In one implementation, when determining the robotic arm's action parameter data based on the to-be-determined robotic arm's action parameter data, the electronic device can perform face detection on each image in the image acquisition dataset to obtain the face detection results of each image. In a specific implementation, the electronic device can use any pre-trained face detection model (such as a face detection model based on a convolutional neural network, etc. This face detection model has been trained with a large amount of face image data and can identify the face region in the image) to perform face detection on each image. Optionally, the face detection result can include face position information, such as represented by rectangular box coordinates, etc. In another specific implementation, for any image in the image acquisition dataset, the electronic device can use the Haar feature-based face detection algorithm to slide a window on any image to extract the Haar features of any image, and use the trained AdaBoost classifier to determine whether the window area is a face region to record the position of the face, such as represented by image coordinates, such as the upper left corner coordinates and the lower right corner coordinates and size information (width w i and height h i ), and at the same time extract the Haar feature vector of the face for subsequent feature matching. i represents the serial number of the i-th camera C i . Suppose the Haar feature vector of the face region detected from the camera C i is (d is the dimension of the feature vector).

[0072] Then, based on the face detection results of each image, the person count result can be determined, and based on the person count result, the robotic arm's action adjustment data can be determined for the robotic arm. The robotic arm's action adjustment data includes adjusting the robotic arm's action amplitude and / or adjusting the robotic arm's action frequency. Optionally, when determining the person count result based on the face detection results of each image, the electronic device can, based on the face detection results of each image, respectively determine the number of faces at each image acquisition moment, and perform a mean operation on the number of faces at each image acquisition moment to obtain the person count result; or, the number of faces at the m-th image acquisition moment can be used as the person count result; or, the number of faces at each image acquisition moment can be used as a person count result respectively, and so on; the embodiments of the present invention do not limit this.

[0073] In a specific implementation, for any one of the m image acquisition moments, the electronic device can determine the position and size information of the face regions in the images captured by each camera at any one image acquisition moment based on the face detection results of the images captured by each camera at any one image acquisition moment, and perform a preliminary screening to remove obviously mismatched cases; for example, the positions are too far apart (that is, if the position distance between any two face regions is greater than a preset distance threshold, it can be considered that they are not the same face), the size ratio difference is extremely large (such as if the size ratio between any two face regions is greater than a preset size ratio threshold, etc., it can be considered that they are not the same face), etc.; optionally, both the preset distance threshold and the preset size ratio threshold can be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this. Then, for any remaining pair of face regions (that is, any pair of face regions in the pairs of all face regions in the images captured by each camera at any one image acquisition moment except for the obviously mismatched pairs of face regions, the two face regions in a pair of face regions are the face regions in the images captured by different cameras), the similarity of any pair of face regions (such as the similarity of Haar feature vectors, using the cosine similarity function to measure) can be calculated, so as to determine whether the two faces indicated by any pair of face regions are the same face based on the similarity. Among them, camera C i and C j The cosine similarity calculation of the feature vectors between the face regions in the captured images is shown in Formula 2.1:

[0074]

[0075] Among them, S in Formula 2.1 represents the cosine similarity. Optionally, when the similarity (such as cosine similarity) is greater than a preset similarity threshold, it can be further verified whether it is the same face through triangulation. Optionally, let the optical centers of camera C i and C j be O i and O j respectively, and the image planes be π i and π j respectively. For the possibly matching face regions detected in the images of the two cameras (whose image coordinates are known), according to the internal and external parameters of the cameras, the three-dimensional coordinates (X w , Y w , Z w ) of the face in the world coordinate system are calculated through the principle of similar triangles. Optionally, the specific calculation method can be to establish a similar triangle relationship composed of the optical center, the pixel points on the image plane, and the spatial points, and use the known internal and external parameters to solve the three-dimensional coordinates. Let the optical center of camera C i be the coordinate in the world coordinate system as Camera Cj The optical center coordinates of For the face region pixel coordinates (u i detected in the image collected by the camera C i , v i )(such as including all pixel point coordinates in the corresponding face region) and the pixel coordinates (u j of the possibly matching face region corresponding in the image collected by the camera C j , v j ), according to the intrinsic matrix K i and K j and the extrinsic relationship, the system of equations shown in Formula 2.2 can be listed:

[0076]

[0077] Based on this, the three-dimensional coordinates of the face region in the world coordinate system can be obtained by solving Formula 2.2; if the three-dimensional coordinates obtained by triangulation are within a reasonable error range (consistent with other determined face positions or known position relationships in the scene), it is determined that these two face regions belong to the same face for counting statistics to avoid double counting; for example, when the three-dimensional coordinates of any two face regions in the world coordinate system are within the specified spatial range, it can be determined that any two face regions belong to the same face, and when the three-dimensional coordinates of any two face regions in the world coordinate system are not within the specified spatial range, it can be determined that any two face regions do not belong to the same face, and so on. Optionally, the specified spatial range can be set according to experience or according to actual needs, and the embodiments of the present invention do not limit this.

[0078] Optionally, in other embodiments, two face regions in a pair of face regions with a similarity greater than a preset similarity threshold can also be determined to be the same face, and two face regions in a pair of face regions with a similarity less than or equal to the preset similarity threshold can be determined to be different faces, and so on. Optionally, the preset similarity threshold can be set according to experience or actual needs, and the embodiments of the present invention do not limit this.

[0079] Based on this, the electronic device can sequentially determine whether any two faces in the images collected by each camera at any image acquisition moment are the same (i.e., whether they are the same face) based on the face detection results of the images collected by each camera at any image acquisition moment, so as to determine the number of different faces in the images collected by each camera at any image acquisition moment, and use the number of different faces as the number of faces at any image acquisition moment.

[0080] In another specific implementation, the electronic device can also determine the positions (i.e., positioning) and postures (such as pitch angles, etc.) of each camera in space, for example, by means of a calibration method; and based on the internal parameters of each camera (such as focal length, pixel size, etc.) and external parameters (such as positioning and posture, etc.), the overlapping area of the fields of view of n cameras can be determined. For example, the field of view range of each camera can be projected into a three-dimensional space to determine the overlapping part between them; Exemplarily, for two cameras with different perspectives, the field of view cones in space can be obtained by calculating their optical axis directions and field of view angles, and the overlapping area is the intersecting part of the two field of view cones; Based on this, based on the face detection results of the images collected by each camera at any image acquisition moment, it can be successively determined whether any two faces in the overlapping area of the images collected by each camera at any image acquisition moment are the same, so as to determine the number of mutually different faces in the images collected by each camera at any image acquisition moment, and use the number of mutually different faces as the number of faces at any image acquisition moment, and so on; The embodiments of the present invention do not limit this.

[0081] Optionally, when determining the robotic arm motion adjustment data of the robotic arm based on the people counting result, the electronic device can use the people counting result to determine the amplitude adjustment factor, and based on the amplitude adjustment factor and the original motion amplitude of the robotic arm, determine the adjusted motion amplitude of the robotic arm. For example, the amplitude adjustment factor is B amp (H), where H represents the people counting result, and B amp (H)=1 + 0.1H, then the adjusted motion amplitude of the robotic arm B = B0×B amp (H); and / or, the people counting result and the frequency adjustment factor can be used to determine the adjusted motion frequency of the robotic arm. For example, the frequency adjustment factor is F freq (H), F freq (H)=1 + 0.05H, then the adjusted motion frequency of the robotic arm F = F0×F freq (H), and so on.

[0082] Further, based on the to-be-determined robotic arm motion parameter data and the robotic arm motion adjustment data, the robotic arm motion parameter data of the robotic arm can be determined. Optionally, the electronic device can add the to-be-determined robotic arm motion speed to the robotic arm motion parameter data; and / or, the to-be-determined robotic arm motion amplitude and the adjusted robotic arm motion amplitude can be weighted and summed to obtain an amplitude integration result. When the amplitude integration result is less than or equal to the preset amplitude threshold, the amplitude integration result can be added to the robotic arm motion parameter data. When the amplitude integration result is greater than the preset amplitude threshold, the preset amplitude threshold can be added to the robotic arm motion parameter data; and / or, the to-be-determined robotic arm motion frequency and the adjusted robotic arm motion frequency can be weighted and summed to obtain a frequency integration result. When the frequency integration result is less than or equal to the preset frequency threshold, the frequency integration result can be added to the robotic arm motion parameter data. When the frequency integration result is greater than the preset frequency threshold, the preset frequency threshold can be added to the robotic arm motion parameter data to realize the determination of the robotic arm motion parameter data of the robotic arm, and so on.

[0083] In another implementation, the electronic device can also use the to-be-determined robotic arm motion parameter data as the robotic arm motion parameter data of the robotic arm, and so on.

[0084] Optionally, the electronic device can generate a robotic arm control instruction based on the target robotic arm motion and / or the robotic arm motion parameter data, and send the robotic arm control instruction to the robotic arm to control the robotic arm to execute the target robotic arm motion. Optionally, in the embodiments of the present invention, the robotic arm motion parameter vector can be set as A (used to indicate the target robotic arm motion, etc.), and the robotic arm motion instruction vector can be set as I (i.e., the robotic arm control instruction). Then, the conversion model from the robotic arm motion parameter data to the robotic arm motion instruction can be expressed as: I = G(A); where G() is a conversion function defined according to the robotic arm control instruction format.

[0085] Optionally, the electronic device can continuously iterate through steps S201 - S205 to obtain the image acquisition data set in real time, so as to update the target robotic arm motion in real time, and so on. Based on this, in the embodiments of the present invention, according to the feedback information after the robotic arm executes the action (such as the action completion situation (such as the completion percentage, etc.), the change situation of visual information (such as the change amount of the number of people and / or the change of the target object, etc.)), the mapping function L() (used to determine the target robotic arm motion) and the conversion function G() can be adjusted to improve the robustness and accuracy of the algorithm; optionally, let the feedback information be l b , the adjusted mapping function be L′(), and the adjusted conversion function be G′(). Then, the adjustment process can be expressed as: L′ = L + ΔL(l b ), G′ = G + ΔG(l b) for representing an interactive feedback adjustment model, which can indicate that the embodiments of the present invention can adjust the actions of the robotic arm according to the real-time environmental conditions; where, l b Represents feedback information, which may include but is not limited to at least one of the following: l comp (Action completion status, completion status) and l vis (Visual information variation, visual variation), etc., and the embodiments of the present invention do not limit this; ΔL and ΔG represent adjustment functions for correcting the mapping function and the conversion function according to the feedback information.

[0086] After obtaining the image acquisition data set, the embodiments of the present invention can perform person detection on each image in the image acquisition data set to obtain the person detection data of each image. The person detection data of an image includes the person segmentation mask image of the corresponding image and the person key point indication data of each person in the corresponding image. The person segmentation mask image of an image includes the segmentation mask of each person in the corresponding image; and based on the person segmentation mask images of each image, the target object is determined. Correspondingly, based on the person key point indication data of the target object, the target robotic arm action mapping indication data corresponding to the target object can be determined. Based on this, the target robotic arm action can be determined based on the target robotic arm action mapping indication data, and the robotic arm can be controlled to execute the target robotic arm action. It can be seen that the embodiments of the present invention can achieve multi-dimensional visual information fusion interaction. That is to say, the embodiments of the present invention comprehensively consider multi-dimensional visual information such as the number of faces, key person postures, and item semantics, and fuse this information for the generation and control of the target robotic arm action, achieving a richer and more intelligent interaction effect between the robotic arm and the surrounding environment (such as an interactive dancing effect), breaking through the limitation of the single action setting of traditional robotic arms. Therefore, the robotic arm interaction method based on visual recognition can also be called the robotic arm interaction dance method or the robotic arm interactive dancing method based on visual recognition, etc.; and, the embodiments of the present invention can implement a dynamic switching mechanism for the target object, that is, it can perform dynamic switching when different target objects appear, enabling the robotic arm to respond in real time to the action changes of different people, enhancing the flexibility and real-time nature of the interaction; in addition, the embodiments of the present invention can also implement action customization based on the semantics of key items, that is, the actions of the robotic arm can be customized according to the semantic characteristics of the items, making the actions of the robotic arm better match the items in the environment, and improving the artistry and ornamental value of the performance.

[0087] Based on the description of the related embodiments of the above-mentioned robotic arm interaction method based on visual recognition, the embodiments of the present invention also propose a robotic arm interaction device based on visual recognition. The robotic arm interaction device based on visual recognition can be a computer program (including program code) running in an electronic device; such as Figure 3As shown, the robotic arm interaction device based on visual recognition may include an image acquisition module 301, an image analysis and processing module 302, an action generation module 303, and a robotic arm control module 304. The robotic arm interaction device based on visual recognition can execute Figure 1 or Figure 2 the robotic arm interaction method based on visual recognition shown, that is, the robotic arm interaction device based on visual recognition can run the above units:

[0088] An image acquisition module 301, configured to acquire an image acquisition data set, where the image acquisition data set includes m image acquisition data, and one image acquisition data includes images acquired by each of n cameras at one image acquisition moment, and both m and n are positive integers;

[0089] An image analysis and processing module 302, configured to determine target robotic arm action mapping indication data corresponding to a target object based on the image acquisition data set; where the target object is the object with the largest object area indication value among at least one object, and the object area indication value of an object is used to indicate the size of the area occupied by the corresponding object;

[0090] An action generation module 303, configured to determine a target robotic arm action based on the target robotic arm action mapping indication data;

[0091] A robotic arm control module 304, configured to control the robotic arm to execute the target robotic arm action.

[0092] In one implementation manner, when the image analysis and processing module 302 determines the target robotic arm action mapping indication data corresponding to the target object based on the image acquisition data set, it may specifically be configured to:

[0093] Perform human detection on each image in the image acquisition data set respectively to obtain human detection data of each image. The human detection data of one image includes the human segmentation mask image of the corresponding image and the human key point indication data of each person in the corresponding image. The human segmentation mask image of one image includes the segmentation masks of each person in the corresponding image;

[0094] Determine the target object based on the human segmentation mask images of the respective images;

[0095] Determine the target robotic arm action mapping indication data corresponding to the target object based on the human key point indication data of the target object.

[0096] In another implementation manner, when the image analysis and processing module 302 determines the target robotic arm action mapping indication data corresponding to the target object based on the human key point indication data of the target object, it may specifically be configured to:

[0097] Based on the human key point indication data of the target object, determine the human body posture action of the target object, and add the human body posture action identifier of the human body posture action of the target object to the target robotic arm action mapping indication data corresponding to the target object for determining the target robotic arm action; and / or,

[0098] Based on the human key point indication data of the target object, determine the action guidance data of the target object, and add the action guidance data to the target robotic arm action mapping indication data; wherein, the action guidance data includes at least one of the following: the amplitude of the human arm, the speed of the human action, and the frequency of the human action, and the action guidance data supports determining the robotic arm action parameter data of the robotic arm so that the robotic arm executes the target robotic arm action according to the robotic arm action parameter data.

[0099] In another implementation manner, the image analysis and processing module 302 can also be used for:

[0100] Perform item detection on each image in the image acquisition dataset respectively to obtain the item detection data of each image, and the item detection data of one image includes the item segmentation mask image of the corresponding image;

[0101] When the image analysis and processing module 302 determines the target object based on the human segmentation mask image of each image, it can be specifically used for:

[0102] Based on the human segmentation mask image and the item segmentation mask image of each image, determine the target object, and the target object is a person or an item;

[0103] The image analysis and processing module 302 can also be used for:

[0104] When the target object is a person, trigger the execution of determining the target robotic arm action mapping indication data corresponding to the target object based on the human key point indication data of the target object;

[0105] When the target object is an item, determine the item semantic information of the target object, and add the item semantic information of the target object to the target robotic arm action mapping indication data corresponding to the target object, and one item semantic information corresponds to one robotic arm action.

[0106] In another implementation manner, when the target object is a person, the target robotic arm action mapping indication data includes a human body posture action identifier; when the target object is an item, the target robotic arm action mapping indication data includes item semantic information; when the action generation module 303 determines the target robotic arm action based on the target robotic arm action mapping indication data, it can be specifically used for:

[0107] When the target robotic arm motion mapping indication data includes a human body posture motion identifier, based on the human body posture motion identifier in the target robotic arm motion mapping indication data, determine a target robotic arm motion identifier that matches the human body posture motion identifier in the target robotic arm motion mapping indication data from a first set of robotic arm motions. The target robotic arm motion identifier is used to indicate the target robotic arm motion to achieve the determination of the target robotic arm motion;

[0108] When the target robotic arm motion mapping indication data includes an item semantic information, based on the item semantic information in the target robotic arm motion mapping indication data, determine a target robotic arm motion identifier that matches the item semantic information in the target robotic arm motion mapping indication data from a second set of robotic arm motions.

[0109] In another implementation, when the target object is a person, the target robotic arm motion mapping indication data includes motion guidance data of the target object. The motion guidance data includes at least one of the following: the amplitude of the person's arm, the speed of the person's motion, and the frequency of the person's motion. The motion generation module 303 can also be used to:

[0110] Based on the motion guidance data, determine the pending robotic arm motion parameter data of the robotic arm. The pending robotic arm motion parameter data includes at least one of the following: the pending amplitude of the robotic arm motion, the pending frequency of the robotic arm motion, and the pending speed of the robotic arm motion. Among them, the pending amplitude of the robotic arm motion is determined based on the amplitude of the person's arm, the pending frequency of the robotic arm motion is determined based on the frequency of the person's motion, and the pending speed of the robotic arm motion is determined based on the speed of the person's motion;

[0111] Based on the pending robotic arm motion parameter data, determine the robotic arm motion parameter data of the robotic arm;

[0112] When the robotic arm control module 304 controls the robotic arm to execute the target robotic arm motion, it can specifically be used to:

[0113] Control the robotic arm to execute the target robotic arm motion according to the robotic arm motion parameter data.

[0114] In another implementation, when the motion generation module 303 determines the robotic arm motion parameter data of the robotic arm based on the pending robotic arm motion parameter data, it can specifically be used to:

[0115] Perform face detection on each image in the image acquisition dataset to obtain the face detection results of each image;

[0116] Based on the face detection results of the respective images, determine the people counting result, and based on the people counting result, determine the robotic arm motion adjustment data of the robotic arm; the robotic arm motion adjustment data includes the adjusted robotic arm motion amplitude and / or the adjusted robotic arm motion frequency of the robotic arm;

[0117] Based on the to-be-determined robotic arm motion parameter data and the robotic arm motion adjustment data, determine the robotic arm motion parameter data of the robotic arm.

[0118] According to an embodiment of the present invention, Figure 3 Each unit in the robotic arm interaction device based on visual recognition shown can be respectively or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, any robotic arm interaction device based on visual recognition can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0119] According to another embodiment of the present invention, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding methods shown in Figure 1 or Figure 2 on a general electronic device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a robotic arm interaction device based on visual recognition as shown in Figure 3 and to implement the robotic arm interaction method based on visual recognition of the embodiments of the present invention. The computer program can be recorded on a computer storage medium, for example, and loaded into the above-mentioned electronic device through the computer storage medium and run therein.

[0120] Based on the descriptions of the above method embodiments and device embodiments, an exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiments of the present invention.

[0121] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiment of the present invention.

[0122] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiment of the present invention.

[0123] Referring Figure 4 , a block diagram of an electronic device 400 that can be a server or a client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0124] As Figure 4 shown, the electronic device 400 includes a computing unit 401 that can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0125] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device capable of inputting information into the electronic device 400. The input unit 406 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 407 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include, but is not limited to, a magnetic disk and an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0126] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above. For example, in some embodiments, the robotic arm interaction method based on visual recognition can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to execute the robotic arm interaction method based on visual recognition by any other suitable means (e.g., by means of firmware).

[0127] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0128] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0129] As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0130] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0131] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0132] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other.

[0133] Moreover, it should be understood that the above - disclosed are only preferred embodiments of the present invention, and of course cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made in accordance with the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A robotic arm interaction method based on visual recognition, characterized in that Including: Obtain an image acquisition data set, where the image acquisition data set includes m image acquisition data. One image acquisition data includes images collected by each of n cameras at one image acquisition moment. Both m and n are positive integers; Based on the image acquisition data set, determine target robotic arm action mapping indication data corresponding to a target object; where the target object is the object with the largest object area indication value among at least one object, and the object area indication value of one object is used to indicate the size of the area occupied by the corresponding object; when the target object is a person, the target robotic arm action mapping indication data includes a human body pose action identifier; when the target object is an item, the target robotic arm action mapping indication data includes item semantic information; Based on the target robotic arm action mapping indication data, determine a target robotic arm action, including: when the target robotic arm action mapping indication data includes a human body pose action identifier, based on the human body pose action identifier in the target robotic arm action mapping indication data, determine a target robotic arm action identifier that matches the human body pose action identifier in the target robotic arm action mapping indication data from a first set of robotic arm action identifiers. The target robotic arm action identifier is used to indicate the target robotic arm action to implement the determination of the target robotic arm action; when the target robotic arm action mapping indication data includes item semantic information, based on the item semantic information in the target robotic arm action mapping indication data, determine a target robotic arm action identifier that matches the item semantic information in the target robotic arm action mapping indication data from a second set of robotic arm action identifiers; Control the robotic arm to execute the target robotic arm action.

2. The method according to claim 1, wherein The determining the target robotic arm action mapping indication data corresponding to the target object based on the image acquisition data set includes: Perform person detection on each image in the image acquisition data set to obtain person detection data for each image. The person detection data for one image includes a person segmentation mask image for the corresponding image and person key point indication data for each person in the corresponding image. A person segmentation mask image for one image includes a segmentation mask for each person in the corresponding image; Based on the person segmentation mask images of the respective images, determine the target object; Based on the person key point indication data of the target object, determine the target robotic arm action mapping indication data corresponding to the target object.

3. The method according to claim 2, wherein The determining the target robotic arm action mapping indication data corresponding to the target object based on the person key point indication data of the target object includes: Based on the person key point indication data of the target object, determine the human body pose action of the target object, and add the human body pose action identifier of the human body pose action of the target object to the target robotic arm action mapping indication data corresponding to the target object for determining the target robotic arm action; and / or, Based on the human key point indication data of the target object, determine the action guidance data of the target object, and add the action guidance data to the target robotic arm action mapping indication data; wherein, the action guidance data includes at least one of the following: human arm amplitude, human action speed, and human action frequency, and the action guidance data supports determining the robotic arm action parameter data of the robotic arm so that the robotic arm executes the target robotic arm action according to the robotic arm action parameter data.

4. The method according to claim 2, wherein The method further includes: Perform object detection on each image in the image acquisition dataset respectively to obtain the object detection data of each image, and the object detection data of one image includes the object segmentation mask image of the corresponding image. The determining the target object based on the human segmentation mask images of the respective images includes: Determine the target object based on the human segmentation mask images and object segmentation mask images of the respective images, and the target object is a human or an object. The method further includes: When the target object is a human, trigger the execution of determining the target robotic arm action mapping indication data corresponding to the target object based on the human key point indication data of the target object. When the target object is an object, determine the object semantic information of the target object, and add the object semantic information of the target object to the target robotic arm action mapping indication data corresponding to the target object, and one object semantic information corresponds to one robotic arm action.

5. The method according to any one of claims 1-4, characterized in that, When the target object is a human, the target robotic arm action mapping indication data includes the action guidance data of the target object, and the action guidance data includes at least one of the following: human arm amplitude, human action speed, and human action frequency; the method further includes: Based on the action guidance data, determine the pending robotic arm action parameter data of the robotic arm, and the pending robotic arm action parameter data includes at least one of the following: pending robotic arm action amplitude, pending robotic arm action frequency, and pending robotic arm action speed; wherein, the pending robotic arm action amplitude is determined based on the human arm amplitude, the pending robotic arm action frequency is determined based on the human action frequency, and the pending robotic arm action speed is determined based on the human action speed. Based on the pending robotic arm action parameter data, determine the robotic arm action parameter data of the robotic arm. The controlling the robotic arm to execute the target robotic arm action includes: Control the robotic arm to execute the target robotic arm action according to the robotic arm action parameter data.

6. The method according to claim 5, wherein The determining the robotic arm action parameter data of the robotic arm based on the pending robotic arm action parameter data includes: Perform face detection on each image in the image acquisition dataset respectively to obtain the face detection results of the respective images. Based on the face detection results of the respective images, determine the population statistics result, and based on the population statistics result, determine the robotic arm motion adjustment data for the robotic arm; the robotic arm motion adjustment data includes the adjusted robotic arm motion amplitude and / or the adjusted robotic arm motion frequency of the robotic arm; Based on the to-be-determined robotic arm motion parameter data and the robotic arm motion adjustment data, determine the robotic arm motion parameter data of the robotic arm.

7. A robotic arm interaction device based on visual recognition, characterized in that, The device includes: An image acquisition module, configured to acquire an image acquisition data set, the image acquisition data set including m image acquisition data, and one image acquisition data including images acquired by each of the n cameras at one image acquisition moment, where both m and n are positive integers; An image analysis and processing module, configured to determine, based on the image acquisition data set, target robotic arm motion mapping indication data corresponding to a target object; where the target object is the object with the largest object area indication value among at least one object, and the object area indication value of an object is used to indicate the size of the area occupied by the corresponding object; when the target object is a person, the target robotic arm motion mapping indication data includes a human body posture motion identifier; when the target object is an item, the target robotic arm motion mapping indication data includes item semantic information; A motion generation module, configured to determine a target robotic arm motion based on the target robotic arm motion mapping indication data, including: when the target robotic arm motion mapping indication data includes a human body posture motion identifier, based on the human body posture motion identifier in the target robotic arm motion mapping indication data, determine a target robotic arm motion identifier that matches the human body posture motion identifier in the target robotic arm motion mapping indication data from a first robotic arm motion set, and the target robotic arm motion identifier is used to indicate the target robotic arm motion to implement the determination of the target robotic arm motion; when the target robotic arm motion mapping indication data includes an item semantic information, based on the item semantic information in the target robotic arm motion mapping indication data, determine a target robotic arm motion identifier that matches the item semantic information in the target robotic arm motion mapping indication data from a second robotic arm motion set; A robotic arm control module, configured to control the robotic arm to execute the target robotic arm motion.

8. An electronic device, characterized in that, It includes: A processor; And a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Robot control method and device, electronic equipment and computer readable storage medium

    CN109397286A