Mechanical arm control method and device, computer device and storage medium
By acquiring real-world images and task information of the robotic arm, and using a fixation point solution model and potential field diagram to generate motion parameters, the problem of poor accuracy in the robotic arm's motion trajectory was solved, thus improving the task success rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUNNY OPTICAL ZHEJIANG RES INST CO LTD
- Filing Date
- 2024-11-21
- Publication Date
- 2026-05-22
AI Technical Summary
In existing technologies, the robotic arm has poor accuracy in determining the motion trajectory after acquiring environmental images, resulting in a low task success rate.
By acquiring real-world images and task information collected by the robotic arm, the gaze point information is determined using a gaze point solution model, an attention distribution potential field map is generated, and motion parameters are determined by combining the robotic arm grasping model to control the movement of the robotic arm.
This improved the accuracy of the robotic arm's motion trajectory and increased the success rate of tasks.
Smart Images

Figure CN122071114A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automatic control technology, and in particular to a robotic arm control method, device, computer equipment, and storage medium. Background Technology
[0002] A robotic arm is a widely used automated mechanical device, increasingly applied in fields such as industrial manufacturing, medical treatment, and entertainment services. A robotic arm consists of multiple joints (axes), each capable of independent movement. Through the coordinated movements of these joints, the robotic arm can perform various tasks. In use, the robotic arm first acquires the task to be performed, and then moves along a predetermined trajectory and speed to execute the task.
[0003] In current technologies, after acquiring a task, the robotic arm needs to determine its motion trajectory based on the acquired environmental image, allowing it to execute the task accordingly. However, the motion trajectory determined solely from the environmental image lacks accuracy, resulting in a low success rate for the robotic arm in performing the task. Summary of the Invention
[0004] Therefore, it is necessary to provide a robotic arm control method, device, computer equipment, and storage medium to address the aforementioned technical problems.
[0005] In a first aspect, this application provides a robotic arm control method, the method comprising: acquiring a real-world image captured by the robotic arm and task information to be performed by the robotic arm; inputting the real-world image and the task information into a gaze point solving model to determine gaze point information corresponding to the real-world image; determining an attention distribution potential field map based on the gaze point information and the real-world image; inputting the attention distribution potential field map and the real-world image into a robotic arm grasping model to obtain motion parameters of the robotic arm, and controlling the movement of the robotic arm based on the motion parameters.
[0006] In one embodiment, before acquiring the environmental image collected by the robotic arm and the task information to be performed by the robotic arm, the method further includes: constructing a first training dataset; the first training dataset includes: multiple training task information, a time series corresponding to each training task information, and training acquisition data corresponding to each time point in the time series; the training acquisition data includes: training environment image, first training gaze point position, training potential field map, and robotic arm training parameters; training a first neural network model based on the training task information in the first training dataset, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training gaze point position corresponding to each time point to obtain the gaze point solving model; training a second neural network model based on the training environment image corresponding to each time point in the first training dataset, the training potential field map, and the robotic arm training parameters to obtain the robotic arm grasping model.
[0007] In one embodiment, constructing the first training dataset includes: acquiring the training task information corresponding to the user controlling the robotic arm to perform a grasping task, the time series corresponding to the training task information, and the training data collected at each time point in the time series; and constructing the first training dataset based on the multiple training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series.
[0008] In one embodiment, acquiring the training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series when the user controls the robotic arm to perform a grasping task includes: acquiring the training task information, the time series corresponding to the training task information, the training environment image, the eye image, and the robotic arm training parameters at each time point in the time series when the user controls the robotic arm to perform a grasping task; inputting the eye image at each time point into the gaze point recognition model to determine the position of the first training gaze point at each time point in the time series; and determining the training potential field map at each time point in the time series based on the position of the first training gaze point and the training environment image.
[0009] In one embodiment, before inputting the eye image corresponding to each time point into the gaze point recognition model to determine the first training gaze point position corresponding to each time point in the time series, the method further includes: constructing a second training dataset; the second training dataset includes: training eye images and second training gaze point positions corresponding to the training eye images; and training a third neural network model based on the second training dataset to obtain the gaze point recognition model.
[0010] In one embodiment, constructing the second training dataset includes: controlling a display device to display a marker point image; the marker point image includes multiple marker points; and recording the corresponding training eye image and the position of the second training gaze point corresponding to the training eye image when the user gazes at each marker point.
[0011] In one embodiment, the third neural network model is a deep convolutional neural network model; the third neural network model includes: multiple convolutional layers, a first fully connected layer, and a second fully connected layer; the multiple convolutional layers are connected sequentially, and the first fully connected layer and the second fully connected layer are respectively connected to the last convolutional layer; the first fully connected layer is used to output the horizontal coordinate; the second fully connected layer is used to output the vertical coordinate.
[0012] In one embodiment, determining the training potential field map corresponding to each time point in the time series based on the first training fixation point position and the training environment image includes: obtaining the target position coordinates in the training environment image; the target position coordinates are any position coordinates other than the first training fixation point position; determining the relative distance based on the target position coordinates and the first training fixation point position; determining the potential field value of the target position coordinates based on the relative distance; and determining the training potential field map corresponding to the training environment image based on the potential field value of each position coordinate in the training environment image.
[0013] In one embodiment, the first neural network model includes a long short-term memory network; the first neural network includes an input layer, a hidden layer, an output layer, and an optimization network; the input layer, hidden layer, and output layer are connected sequentially; the optimization network is connected to the input layer, hidden layer, and output layer respectively; the input layer is used to partition the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training fixation point position corresponding to each time point according to the task information and the time series; the hidden layer is used to generate long-term memory features and short-term memory features based on the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training fixation point position corresponding to each time point; the output layer is used to output the training result based on the long-term memory features and short-term memory features; the optimization network includes a loss function for optimizing the parameters of the hidden layer.
[0014] In one embodiment, the second neural network model includes a long short-term memory network.
[0015] Secondly, this application also provides a robotic arm control device, the device comprising: an acquisition module for acquiring a real-world image captured by the robotic arm and task information to be performed by the robotic arm; a gaze point calculation module for inputting the real-world image and the task information into a gaze point solving model to determine the gaze point information corresponding to the real-world image; a potential field map calculation module for determining an attention distribution potential field map based on the gaze point information and the real-world image; and a control module for inputting the attention distribution potential field map and the real-world image into a robotic arm grasping model to obtain the motion parameters of the robotic arm, and controlling the movement of the robotic arm based on the motion parameters.
[0016] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement any of the robotic arm control methods described in the first aspect.
[0017] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements any one of the robotic arm control methods described in the first aspect.
[0018] The aforementioned robotic arm control method, device, computer equipment, and storage medium acquire real-world environmental images and task information collected by the robotic arm. The real-world environmental information and task information are input into a gaze point solving model to determine the corresponding gaze point information in the real-world environmental images. Then, based on the gaze point information and the real-world environmental images, an attention distribution potential field map is determined. Finally, the attention distribution potential field map and the real-world environmental images are input into a robotic arm grasping model to obtain the robotic arm's motion parameters, and the robotic arm's movement is controlled based on these motion parameters. By determining gaze point information from real-world environmental images, generating an attention distribution potential field map based on the gaze point information, and finally combining the attention distribution potential field map and the real-world environmental images to determine the robotic arm's motion parameters, the accuracy of the robotic arm's motion trajectory for the task to be performed is improved, further increasing the success rate of the robotic arm in completing the task. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating a robotic arm control method in one embodiment;
[0020] Figure 2 This is a flowchart illustrating a model training method in one embodiment;
[0021] Figure 3 This is a flowchart illustrating a method for obtaining training data in one embodiment;
[0022] Figure 4 This is a training environment image in one embodiment;
[0023] Figure 5 An eye image from one embodiment;
[0024] Figure 6 This is a flowchart illustrating a method for calculating the fixation point position in one embodiment;
[0025] Figure 7 This is a marker point image from one embodiment;
[0026] Figure 8 This is a training potential field diagram in one embodiment;
[0027] Figure 9 This is a flowchart illustrating the robotic arm control process in one embodiment;
[0028] Figure 10 This is a structural block diagram of a robotic arm control device in one embodiment;
[0029] Figure 11 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0031] A robotic arm is an automated device that simulates the movements of a human arm. It typically consists of multiple joints (axes), each capable of independent movement. Through programming and control, robotic arms can perform various complex tasks, such as grasping, handling, and assembly. A robotic arm generally includes: a base, joints, links, an end effector, and sensors. The base is used to fix the robotic arm and is usually mounted on the ground, a workbench, or a mobile platform. Joints are the moving parts of the robotic arm; each joint can perform rotational or linear motion. Links connect the various joints. The end effector adapts to different tasks, such as grippers or suction cups. Sensors can be placed at any position on the robotic arm; for example, image sensors can collect environmental information while the robotic arm is running. In related technologies, when performing tasks, robotic arms rely solely on image sensors to collect environmental information and control the arm based solely on this information. This results in poor accuracy of the robotic arm's movement trajectory during task completion, further leading to a low success rate in task execution.
[0032] In one embodiment, such as Figure 1 As shown, a robotic arm control method is provided, including the following steps:
[0033] Step 101: Obtain real-world images captured by the robotic arm and information on the tasks to be performed by the robotic arm.
[0034] When using a robotic arm to perform a task, the system first acquires images of the real-world environment captured by the robotic arm, as well as task information corresponding to the task to be performed. An image sensor is located at the end effector of the robotic arm, capturing images of the current external environment, i.e., the real-world environment. The task information is the specific task to be performed, such as "pick up the water glass on the table" or "take the tape from the first drawer and put it on the table."
[0035] Step 102: Input the real environment image and task information into the gaze point solution model to determine the corresponding gaze point information in the real environment image.
[0036] After acquiring the real-world environment image and task information, the image and task information are input into the gaze point determination model to determine the corresponding gaze point information in the real-world environment image. During the robotic arm's task execution, the robotic arm continuously acquires real-world environment images and continuously inputs these images and task information into the gaze point determination model, thus obtaining the gaze point information corresponding to the real-world environment image acquired at each time point. The gaze point information refers to the location of the gaze point in the real-world environment image. The gaze point determination model simulates the gaze point of a real user when completing the corresponding task. This gaze point not only maps the user's focus of attention on the target object in the task but also provides richer execution logic for the robotic arm during grasping tasks. The gaze point determination model is obtained after training a neural network model.
[0037] Step 103: Determine the attention distribution potential field map based on the gaze point information and the real environment image.
[0038] After obtaining the gaze point information corresponding to the real environment image, the attention distribution potential field map is determined based on the gaze point information and the real environment image. Here, the gaze point information refers to the location of the gaze point in the real environment image. First, each coordinate information in the real environment image is obtained. Based on each coordinate information and the gaze point information, the distance information between each coordinate information and the gaze point information in the real environment image is calculated. Then, based on this distance information, the potential field value corresponding to each coordinate information in the real environment image is calculated. Finally, the attention distribution potential field map is generated based on the potential field value corresponding to each coordinate information. For example, the gaze point information is denoted as (x0, y0), and each coordinate information in the real environment image is denoted as (xi, yi). The potential field value V corresponding to each coordinate information can be expressed as:
[0039]
[0040] Where k represents the potential field value coefficient, which can be set according to actual usage requirements. The potential field value corresponding to the gaze point information does not need to be calculated, or can be set to a large constant.
[0041] Step 104: Input the attention distribution potential field map and the real environment image into the robotic arm grasping model to obtain the motion parameters of the robotic arm, and control the movement of the robotic arm based on the motion parameters.
[0042] After determining the attention distribution potential field map, the map and the real environment image are input into the robotic arm grasping model to obtain the robotic arm's motion parameters. During the robotic arm's task execution, it continuously acquires real environment images. At each time point, the acquired real environment image and task information are input into the gaze point solving model to obtain the gaze point information corresponding to that time point. Further, the attention distribution potential field map corresponding to that time point is obtained, and this map, along with the real environment image, is input into the robotic arm grasping model to obtain the robotic arm's motion parameters for that time point. It can be understood that during task execution, the above calculations are performed for each time point to determine the corresponding robotic arm motion parameters, thereby completing the task to be performed by the robotic arm. The robotic arm's motion parameters include the rotation angle or linear movement distance of the robotic arm joints. Taking a robotic arm with six joints as an example, the motion parameters include the rotation angle or linear movement distance of each joint. After determining the robotic arm's motion parameters, its movement can be controlled based on these parameters.
[0043] The above embodiments acquire real-world images captured by the robotic arm and information about the task to be performed by the robotic arm. The real-world information and task information are input into a gaze point determination model to determine the corresponding gaze points in the real-world images. Then, based on the gaze point information and the real-world images, an attention distribution potential field map is determined. Finally, the attention distribution potential field map and the real-world images are input into a robotic arm grasping model to obtain the robotic arm's motion parameters, and the robotic arm's movement is controlled based on these parameters. By determining gaze point information from real-world images, generating an attention distribution potential field map based on the gaze point information, and finally combining the attention distribution potential field map and the real-world images to determine the robotic arm's motion parameters, the accuracy of the robotic arm's motion trajectory for the task to be performed is improved, further increasing the success rate of the robotic arm in completing the task.
[0044] In one embodiment, such as Figure 2 As shown, a model training method is provided, including the following steps:
[0045] Step 201: Construct the first training dataset.
[0046] The first training dataset includes: information on multiple training tasks, the time series corresponding to each training task, and the training data collected at each time point in the time series. The training data collected includes: training environment images, the location of the first training gaze point, the training potential field map, and the robotic arm training parameters. Details are as follows:
[0047]
[0048]
[0049] When constructing the first training dataset, we first obtain the training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series when the user controls the robotic arm to perform the grasping task.
[0050] The training task information includes any task that requires the robotic arm to perform, such as "picking up the water glass on the table" or "taking out the tape from the first drawer and placing it on the table." The time series represents each point in time when the user controls the robotic arm to perform the grasping task. The training data includes: the training environment image, the position of the first training gaze point, the training potential field map, and the robotic arm training parameters corresponding to each point in time when the user controls the robotic arm to perform the grasping task. For example, the time required for the user to control the robotic arm to perform the grasping task is T1-Tn. Training data is recorded at each time point. At T1, the collected training environment image is Real_Image1, the eye image is Eye_Image1, the position of the first training gaze point is (x1, y1), the training potential field map is V_Image1, and the robotic arm training parameters are {a11, a21, a31, a41, a51, a61}. The robotic arm training parameters represent the rotation angle or linear motion distance of each of the six joints of the robotic arm.
[0051] After obtaining the above data, the first training dataset is constructed based on multiple training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series.
[0052] Step 202: Train the first neural network model based on the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the position of the first training gaze point corresponding to each time point in the first training dataset to obtain the gaze point solution model.
[0053] After constructing the first training dataset, the first neural network model is trained based on the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training gaze point position corresponding to each time point, to obtain the gaze point solution model. The first neural network model can be a Long Short-Term Memory network.
[0054] Step 203: Based on the training environment image, training potential field diagram and robotic arm training parameters corresponding to each time point in the first training dataset, train the second neural network model to obtain the robotic arm grasping model.
[0055] After constructing the first training dataset, the second neural network model is trained based on the training environment image, training potential field diagram, and robotic arm training parameters at each time point to obtain the robotic arm grasping model. The second neural network model can be a long short-term memory network.
[0056] This embodiment constructs a first training dataset and trains a first neural network model and a second neural network model using the first training dataset to obtain a gaze point solution model and a robotic arm grasping model. Furthermore, the first and second neural network models utilize long short-term memory networks, which can fully utilize the temporal factors present in the first training dataset, thereby improving the accuracy of the model output results and further enhancing the accuracy of the robotic arm's motion trajectory for the task at hand.
[0057] In one embodiment, such as Figure 3 As shown, a method for obtaining training data is provided, including the following steps:
[0058] Step 301: Obtain the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point in the time series, the eye image corresponding to each time point in the time series, and the robotic arm training parameters corresponding to each time point in the time series when the user controls the robotic arm to perform the grasping task.
[0059] The user wears a wearable device and holds a remote control for the robotic arm; this wearable device can be VR glasses. The end effector of the robotic arm is equipped with an image sensor, such as a 3D camera. The image sensor captures real-time images of the training environment and transmits them wirelessly to the wearable device. The wearable device displays these images on its display component, allowing the user to observe the training environment from the robotic arm's perspective. Simultaneously, the wearable device includes an eye image acquisition device, such as a camera, to capture real-time images of the user's eyes. These eye images must include both eyes. The rotation angle or linear movement distance of each joint of the robotic arm can be captured in real-time. When acquiring training data, the user controls the robotic arm's movement using the remote control, based on the training task information to be performed. During the control of the robotic arm's movement, the user records the training environment image, eye image, and robotic arm training parameters at a preset time frequency for each point in time. For example, when the robotic arm needs to perform the task of "grabbing a water glass from a table," the training task information is recorded. The user controls the robotic arm's movement through a handheld remote control device, enabling the robotic arm to grab the water glass from the table. During the task execution, training environment images, eye images, and robotic arm training parameters are collected at a preset time frequency, i.e., by the image sensor set at the end of the robotic arm. Figure 4 The training environment image shown; the wearable device is equipped with an eye diagram acquisition device to collect images such as... Figure 5 The image shown is of an eye; the robotic arm acquires training parameters in real time during the movement process and generates the following table:
[0060]
[0061] Wherein, the time series T1-Tn represents each time point corresponding to the execution of the training task. At time T1, the collected training environment image is Real_Image1, the eye image is Eye_Image1, and the robotic arm training parameters are {a11,a21,a31,a41,a51,a61}.
[0062] By having the user control the robotic arm to perform different training tasks, the system can obtain different training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point in the time series, the eye image corresponding to each time point in the time series, and the robotic arm training parameters corresponding to each time point in the time series.
[0063] Step 302: Input the eye image corresponding to each time point into the gaze point recognition model to determine the position of the first training gaze point corresponding to each time point in the time series.
[0064] After obtaining the above data, it is necessary to determine the location of the first training gaze point for each time point based on the eye image corresponding to each time point. The location of the first training gaze point is the gaze point location of the user in the training environment image at the corresponding time point.
[0065] Since the eye image acquisition device built into the wearable device can only capture images of the user's eyes, but we need to obtain the user's gaze point location, we need to determine the gaze point location in the training environment image based on the eye image. This is done by calculating the intersection of the extended line of sight passing through the gaze origin and the gaze plane based on the eye image, thus determining the gaze point location. Figure 6 As shown, the process first estimates the gaze direction based on the eye image to determine the gaze direction; then, it models the gaze line equation based on the gaze direction and the gaze origin; finally, it determines the gaze plane position and normal vector based on the training environment image corresponding to the eye image; and then models the gaze plane based on the gaze plane position and normal vector. Finally, it calculates and outputs the gaze point position based on the gaze line equation model and the gaze plane model. With the continuous development of deep learning technology, it can also be applied to gaze point position calculation tasks. This task allows deep networks to learn the mapping relationship between the input image and the gaze point position in a two-dimensional plane, thereby estimating the gaze point in the two-dimensional plane. Compared to gaze point calculation methods based on the intersection of the line of sight and the gaze plane, deep learning-based gaze point estimation methods are more direct and simpler.
[0066] Before training the gaze recognition model, a second training dataset needs to be constructed first; then, the third neural network model is trained based on the second training dataset to obtain the gaze recognition model.
[0067] When constructing the second training dataset, the display device is controlled to display a labeled point image; the labeled point image includes multiple labeled points. The user wears a wearable device, such as VR glasses, and the labeled point image is displayed on the wearable device's display screen. Figure 7 As shown. Figure 7 The red dots in the image represent marker points, and the coordinates of each red dot are known. When the user gazes at each marker point, the corresponding training eye image and the location of the second training gaze point corresponding to the training eye image are recorded. The second training dataset includes training eye images and the locations of the second training gaze points corresponding to the training eye images. Specifically, during data acquisition, the user is asked to gaze at each marker point in the marker image sequentially, and the corresponding training eye image is recorded when each marker point is gazed at. The current coordinates of the marker point are also the locations of the second training gaze points corresponding to the training eye image. The second training dataset is shown below:
[0068] Training Eye Images Second training fixation point position 1 <![CDATA[Eye_Image1]]> <![CDATA[(x1,y1)]]> 2 <![CDATA[Eye_Image2]]> <![CDATA[(x2,y2)]]> 3 <![CDATA[Eye_Image3]]> <![CDATA[(x3,y3)]]> 4 <![CDATA[Eye_Image4]]> <![CDATA[(x4,y4)]]> ... ... ... n-3 <![CDATA[Eye_Image n-3 ]]> <![CDATA[(x5,y5)]]> n-2 <![CDATA[Eye_Image n-2 ]]> <![CDATA[(x6,y6)]]> n-1 <![CDATA[Eye_Image n-1 ]]> <![CDATA[(x7,y7)]]> n <![CDATA[Eye_Image n ]]> <![CDATA[(x8,y8)]]>
[0069] After obtaining the second training dataset, the third neural network model is trained based on the second training dataset to obtain the gaze point recognition model.
[0070] In one embodiment, the third neural network model is a deep convolutional neural network model. The third neural network model includes: multiple convolutional layers, a first fully connected layer, and a second fully connected layer; the multiple convolutional layers are connected sequentially, and the first and second fully connected layers are respectively connected to the last convolutional layer; the first fully connected layer outputs the horizontal coordinate; the second fully connected layer outputs the vertical coordinate. For example, this embodiment uses a ResNet50 network as the convolutional layer network, used as a feature extraction network for eye map texture information; two fully connected layers are concatenated after the convolutional layer network to integrate the features of the gaze in the x and y coordinates of the eye image, respectively. The two fully connected layers calculate the x and y coordinates of the gaze point position, and train the x and y coordinate features using two identical loss functions. Each loss function is a linear combination of cross-entropy loss and mean squared loss. The two fully connected layers integrate the gaze features in the x and y dimensions, respectively, and feed them into two non-linear SoftMax layers, which predict the x and y coordinates of the gaze point position, respectively. In each loss function, cross-entropy loss is used to predict the gaze classification in the corresponding dimension, achieving a coarse estimate of the gaze's current position and coarsely adjusting the parameters of the deep convolutional neural network model. Simultaneously, the expected value of the current dimension is calculated, and the mean squared loss line is used for fine-tuning the gaze in the current dimension, further refining the parameters of the deep convolutional neural network model. Finally, the cross-entropy loss and mean squared loss line are combined to improve the model's computational accuracy.
[0071] In one embodiment, the cross-entropy loss is expressed as:
[0072] H(y i p i )=-∑ i y i log p i
[0073] Where H represents the cross-entropy loss, and yi and pi represent the true value and predicted value in the corresponding dimension, respectively.
[0074] The mean square loss line is represented as:
[0075]
[0076] Where MSE represents mean squared loss, yi and pi represent the true value and predicted value in the corresponding dimension, respectively, and N represents the number of predictions.
[0077] The joint loss of the cross-entropy loss and the mean squared loss is expressed as:
[0078] CLS(y i p i )=H(y i p i )+a·MSE(y i p i )
[0079] Where H represents cross-entropy loss, MSE represents mean squared loss, CLS represents joint loss, yi and pi represent the true and predicted values in the corresponding dimensions, respectively, and a represents the regression coefficient of mean squared loss.
[0080] After calculating the location of the first training fixation point, the locations of the first training fixation points corresponding to each time point in the time series are summarized, resulting in the following table:
[0081]
[0082] Step 303: Based on the location of the first training gaze point and the training environment image, determine the training potential field map corresponding to each time point in the time series.
[0083] After obtaining the first training fixation point position corresponding to each time point in the time series, the training potential field map corresponding to each time point in the time series is determined based on the first training fixation point position and the corresponding training environment image. First, the target position coordinates in the training environment image are obtained; the target position coordinates are any position coordinates other than the first training fixation point position. Then, the relative distance is determined based on the target position coordinates and the first training fixation point position. Based on the relative distance, the potential field value of the target position coordinates is determined. Based on the potential field value of each position coordinate in the training environment image, the training potential field map corresponding to the training environment image is determined. The method of calculating the training potential field map based on the first training fixation point position and the training environment image is the same as the method of determining the attention distribution potential field map based on fixation point information and the real environment image. Specifically, for example, the first training fixation point position is (x0, y0), and each coordinate information in the training environment image is (xi, yi); the potential field value V corresponding to each coordinate information can be expressed as:
[0084]
[0085] Where k represents the potential field value coefficient, which can be set according to actual usage requirements. The potential field value corresponding to the first training fixation point position does not need to be calculated, or can be set to a large constant. The constructed training potential field diagram is as follows. Figure 8 As shown, Figure 8The highest value in the training potential field map corresponds to the first training fixation point. The closer the potential field is to the first training fixation point, the larger the potential field value; the farther away it is, the smaller the potential field value. The training potential field map clearly reflects the user's visual attention information during the grasping task.
[0086] After calculating the training potential field diagram, the training potential field diagrams corresponding to each time point in the time series are summarized, resulting in the following table:
[0087]
[0088] In this implementation, a gaze point recognition model is used to determine the position of the first training gaze point. The gaze point recognition model has a simple structure, fast calculation speed, and high solution efficiency, thereby improving the accuracy of the dataset and the computational efficiency. By determining the training potential field map through the gaze point position, the attention information of the user's gaze during the grasping task can be better reflected, further improving the training accuracy of subsequent models, and thus improving the accuracy of the robotic arm's motion trajectory for the task to be performed.
[0089] After constructing the first training dataset, the first neural network model and the second neural network model need to be trained separately based on the first training dataset to obtain the foveation point solution model and the robotic arm grasping model. Since the first training dataset contains temporal factors, RNN neural networks are the preferred network for time prediction tasks. However, RNNs are prone to gradient vanishing or gradient exploding problems. Therefore, LSTM network models, also known as Long Short-Term Memory networks, can be used. The LSTM network model is a type of recurrent neural network that introduces gating units into the traditional RNN network and adds memory units. The gating units control the information to replace the information stored in the memory units. During model training, the foveation point solution model mainly obtains the mapping relationship between the training environment image and training task information and the first training foveation point position, thus obtaining the foveation point solution model; the robotic arm grasping model mainly estimates the robotic arm training parameters based on the training environment image and training potential field map, thus obtaining the robotic arm grasping model.
[0090] In one embodiment, a first neural network model is trained based on training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training fixation point position corresponding to each time point in the first training dataset to obtain a fixation point solution model. The first neural network model includes a long short-term memory network; the first neural network includes an input layer, a hidden layer, an output layer, and an optimization network; the input layer, hidden layer, and output layer are connected sequentially; the optimization network is connected to the input layer, hidden layer, and output layer respectively; the input layer is used to partition the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training fixation point position corresponding to each time point according to the task information and the time series; the hidden layer is used to generate long-term memory features and short-term memory features based on the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training fixation point position corresponding to each time point; the output layer is used to output the training results based on the long-term memory features and short-term memory features; the optimization network includes a loss function used to optimize the parameters of the hidden layer.
[0091] The first neural network model can be divided into four parts: input layer, hidden layer, output layer, and optimization network. The input layer partitions the dataset according to task information and time series, dividing the time series corresponding to the training task information, the training environment image at each time point, and the first training fixation point position at each time point. The partitioned training data is then input into each layer of the hidden layer according to different training task information and time points. The hidden layer includes the model's structural parameters and depth, and generates long-term memory features and short-term memory features based on the input data. The output layer performs inverse normalization based on the long-term memory features and short-term memory features, outputting the training result, and evaluates the model against the first training fixation point position in the first training dataset. During model training, gradient descent is used for model optimization, with the Adam optimizer and MSE (mean squared loss) loss function. The training environment image and training task information are used as input to the first neural network model, and the first training fixation point position is used as the output. A long short-term memory network is used for model training, ultimately obtaining the mapping relationship between the training environment image and training task information and the first training fixation point position, as detailed below:
[0092] output(x i y i ) = f(Real_Imege i task_info i )
[0093] Where Real_Imagei represents the training environment image at time point i, task_infoi represents the training task information at time point i, and output(xi, yi) represents the position of the first training fixation point at time point i.
[0094] In one embodiment, a second neural network model is trained based on the training environment image, training potential field map, and robotic arm training parameters corresponding to each time point in the first training dataset to obtain a robotic arm grasping model. The second neural network model includes a Long Short-Term Memory (LSTM) network. After obtaining the gaze point position from the model based on the gaze point, the corresponding potential field map is obtained based on the gaze point position. The LTM network learns the mapping relationship between the training environment image, training potential field map, and robotic arm training parameters. The training data used by the LTM network has temporal coupling; therefore, the integrity of the time-series data needs to be maintained when training the robotic arm grasping model. For the constructed first training dataset, the data at each time point corresponding to each training task information needs to be sequentially input into the LTM network for model training, according to the time point order. During model training, the target object and action in the training task information can also be encoded and identified according to the training task information to obtain target object IDs and action IDs, which are then used as input data for training the LTM network. The target object ID can be an integer, such as 1 for a bottle, 2 for a box, etc., or it can be an independent encoding, such as an independent encoding vector [1,0,0] for a bottle and [0,1,0] for a box. This encoding can better adapt to the input structure of the neural network, making it easier for the model to recognize different target objects. The action ID refers to the specific operation type taken by the robotic arm for the target object when performing a task. For example, common robotic arm actions may include "grasping," "extending," and "placing." Each action type will have a unique ID, through which the robotic arm can distinguish the specific operation to be performed. The specific ID encoding method can refer to the target object encoding method. The training environment image and training potential field map are used as inputs to the second neural network model, and the robotic arm training parameters are used as outputs. A long short-term memory network is used for model training, ultimately obtaining the mapping relationship between the training environment image and training potential field map and the robotic arm training parameters, as follows:
[0095] output(a1 i a2 i a3 i a4 i a5 i a6 i ) = f(Real_Image i V_image i )
[0096] Where Real_Imagei represents the training environment image at time point i, V_Imagei represents the training potential field image at time point i, and output(a1i, a2i, a3i, a4i, a5i, a6i) represents the robotic arm training parameters at time point i.
[0097] In one specific embodiment, such as Figure 9 As shown, a robotic arm control process is provided. This embodiment proposes to integrate eye-tracking technology with the robotic arm's machine learning training. In the gaze point solving model and the robotic arm grasping model, the gaze point is solved as a guide, improving the success rate and efficiency of the robotic arm in completing the specified task. The process provided in this embodiment mainly includes four parts: data acquisition, data preprocessing, model training, and model application.
[0098] During the data acquisition phase, eye diagram data of the user is acquired using an eye diagram acquisition device, which can be VR glasses. Real-world image data is acquired using an image sensor located at the end effector of the robotic arm, which can be a 3D camera. Joint data is acquired using photoelectric encoders located at each joint of the robotic arm.
[0099] In the data preprocessing stage, the gaze point position corresponding to the eye map data is determined by the deep convolutional neural network corresponding to the gaze point recognition model based on the eye map data. Then, based on the gaze point position and real environment image data, the attention distribution potential field map is constructed by example calculation.
[0100] The model training phase mainly includes training the gaze point solving model and the robotic arm grasping model. Using real-world image data, task information, and gaze point positions, a Long Short-Term Memory (LSTM) network is trained to solve the mapping relationship between "real-world image data and task information" and "gaze point positions," thus obtaining the gaze point solving model. Using real-world image data, attention distribution potential field maps, and robotic arm joint data, a Long Short-Term Memory (LSTM) network is trained to solve the mapping relationship between "real-world image data and attention distribution potential field maps" and "robotic arm joint data," thus obtaining the robotic arm grasping model. The robotic arm joint data refers to the robotic arm joint angle data.
[0101] In the model application phase, the task information to be executed by the robotic arm and real-time image data of the real environment are acquired. The task information and real environment image data are input into the trained gaze point solving model to obtain the gaze point position. An attention distribution potential field map is generated based on the gaze point position. Then, the real environment image data and attention distribution potential field map are input into the trained robotic arm grasping model to obtain the joint angle data of the robotic arm. The robotic arm is motion controlled by the joint angle data to execute the task to be executed.
[0102] In the field of biomimetic design, robotic arms are inspired by the human arm, aiming to replicate its flexibility and precise grasping ability for applications in automated production and service robots. Eye-tracking technology, by capturing and analyzing eye map data, reveals the observer's visual attention distribution and decision priorities, injecting new momentum into technological development. This application integrates eye-tracking technology with machine learning training for robotic arms to improve the efficiency and intelligence of grasping tasks. During user demonstrations of grasping, the eye-tracking device records their eye map data in real time. This data not only maps the operator's focus of attention on the object but also contains the deep logic of their grasping strategy. Combining eye-tracking data with the robotic arm's grasping action data provides rich and multi-dimensional learning materials for machine learning algorithms, promoting deep learning and understanding of the model. By analyzing eye-tracking patterns in different grasping tasks, machine learning algorithms can learn the patterns of human visual attention in grasping decisions and their correlation with actions. Based on eye-tracking data-assisted learning, the robotic arm can more accurately identify targets and intelligently plan paths in grasping tasks in complex environments, significantly improving grasping efficiency and success rate.
[0103] In this embodiment, by adding gaze points to the training samples, the gaze points can accurately reflect the user's focus during the teaching process. An eye map acquisition device collects the user's eye map data in real time during the teaching process and calculates the gaze point position. The gaze point position not only maps changes in the user's attention but also reveals the internal logic of their grasping strategy for the target object. Adding gaze point information to the training samples provides richer and more comprehensive learning materials for the machine learning system, helping to enhance the model's deep learning and understanding capabilities. During model training, an attention distribution potential field map is added. Centered on the gaze point, the reciprocal of the distance from each point in the real-world image to the gaze point is calculated as its potential field value, generating a potential field map rich in deep attention information. The real-time generation of the potential field map can promptly reflect changes in the user's attention during the grasping task, providing precise guidance for the robotic arm to dynamically adjust its grasping strategy. This helps enhance the robotic arm's adaptability and flexibility in complex environments, improves the success rate and efficiency of grasping tasks, reduces ineffective movements, and saves task execution time. By integrating knowledge point data with deep learning and introducing focal point information, the model can more accurately capture key features in the task execution process, improving its understanding and execution capabilities for complex grasping tasks. This not only enhances the model's learning focus and efficiency but also reduces the amount of training data required. At the same time, it improves the model's generalization ability, enabling the model to learn effective grasping strategies more intuitively.
[0104] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0105] Based on the same inventive concept, this application also provides a robotic arm control device for implementing the robotic arm control method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the robotic arm control device provided below can be found in the limitations of the robotic arm control method described above, and will not be repeated here.
[0106] In one embodiment, such as Figure 10As shown, a robotic arm control device is provided, including: an acquisition module 100, a gaze point calculation module 200, a potential field map calculation module 300, and a control module 400, wherein:
[0107] The acquisition module 100 is used to acquire real-world images collected by the robotic arm and information on the tasks to be performed by the robotic arm.
[0108] The gaze point calculation module 200 is used to input the real environment image and the task information into the gaze point solving model to determine the corresponding gaze point information in the real environment image.
[0109] The potential field map calculation module 300 is used to determine the attention distribution potential field map based on the gaze point information and the real environment image.
[0110] The control module 400 is used to input the attention distribution potential field map and the real environment image into the robotic arm grasping model to obtain the motion parameters of the robotic arm, and control the movement of the robotic arm based on the motion parameters.
[0111] The robotic arm control device also includes: a model building model;
[0112] A model building module is used to construct a first training dataset. The first training dataset includes: multiple training task information, a time series corresponding to each training task information, and training acquisition data corresponding to each time point in the time series. The training acquisition data includes: training environment images, first training gaze point positions, training potential field maps, and robotic arm training parameters. Based on the training task information, the time series corresponding to the training task information, the training environment images corresponding to each time point, and the first training gaze point positions corresponding to each time point in the first training dataset, a first neural network model is trained to obtain the gaze point solving model. Based on the training environment images, training potential field maps, and robotic arm training parameters corresponding to each time point in the first training dataset, a second neural network model is trained to obtain the robotic arm grasping model.
[0113] The model building module is also used to acquire the training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series when the user controls the robotic arm to perform a grasping task; and to construct the first training dataset based on the multiple training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series.
[0114] The model building module is further configured to acquire, when the user controls the robotic arm to perform a grasping task, the corresponding training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point in the time series, the eye image corresponding to each time point in the time series, and the robotic arm training parameters corresponding to each time point in the time series; input the eye image corresponding to each time point into the gaze point recognition model to determine the position of the first training gaze point corresponding to each time point in the time series; and determine the training potential field map corresponding to each time point in the time series based on the position of the first training gaze point and the training environment image.
[0115] The model building module is also used to build a second training dataset; the second training dataset includes: training eye images and the second training gaze point positions corresponding to the training eye images; the third neural network model is trained based on the second training dataset to obtain the gaze point recognition model.
[0116] The model building module is also used to control the display device to display the marker point image; the marker point image includes multiple marker points; when the user gazes at each of the marker points respectively, the corresponding training eye image and the position of the second training gaze point corresponding to the training eye image are recorded.
[0117] The third neural network model is a deep convolutional neural network model; the third neural network model includes: multiple convolutional layers, a first fully connected layer, and a second fully connected layer; the multiple convolutional layers are connected sequentially, and the first fully connected layer and the second fully connected layer are respectively connected to the last convolutional layer; the first fully connected layer is used to output the horizontal coordinate; the second fully connected layer is used to output the vertical coordinate.
[0118] The model building module is further configured to obtain the target position coordinates in the training environment image; the target position coordinates are any position coordinates other than the first training gaze point position; determine the relative distance based on the target position coordinates and the first training gaze point position; determine the potential field value of the target position coordinates based on the relative distance; and determine the training potential field map corresponding to the training environment image based on the potential field value of each position coordinate in the training environment image.
[0119] The first neural network model includes a long short-term memory network; the first neural network includes an input layer, a hidden layer, an output layer, and an optimization network; the input layer, hidden layer, and output layer are connected sequentially; the optimization network is connected to the input layer, hidden layer, and output layer respectively; the input layer is used to partition the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training fixation point position corresponding to each time point according to the task information and the time series; the hidden layer is used to generate long-term memory features and short-term memory features based on the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training fixation point position corresponding to each time point; the output layer is used to output the training result based on the long-term memory features and short-term memory features; the optimization network includes a loss function for optimizing the parameters of the hidden layer.
[0120] The second neural network model includes a long short-term memory network.
[0121] Each module in the aforementioned robotic arm control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0122] In one embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data related to implementing a robotic arm control method. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a robotic arm control method.
[0123] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0124] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the robotic arm control methods described above.
[0125] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements any of the robotic arm control methods described above.
[0126] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the robotic arm control methods described above.
[0127] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0128] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0129] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A robotic arm control method, characterized in that, The method includes: Acquire real-world images captured by the robotic arm and information about the tasks to be performed by the robotic arm; The real environment image and the task information are input into the gaze point solving model to determine the corresponding gaze point information in the real environment image. Based on the gaze point information and the real environment image, determine the attention distribution potential field map; The attention distribution potential field map and the real environment image are input into the robotic arm grasping model to obtain the motion parameters of the robotic arm, and the movement of the robotic arm is controlled based on the motion parameters.
2. The method according to claim 1, characterized in that, Before acquiring the environmental images collected by the robotic arm and the task information to be performed by the robotic arm, the method further includes: Construct a first training dataset; the first training dataset includes: multiple training task information, a time series corresponding to each training task information, and training acquisition data corresponding to each time point in the time series; the training acquisition data includes: training environment image, first training gaze point position, training potential field map, and robotic arm training parameters. Based on the training task information in the first training dataset, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the first training gaze point position corresponding to each time point, the first neural network model is trained to obtain the gaze point solution model. The second neural network model is trained based on the training environment image, the training potential field diagram, and the robotic arm training parameters corresponding to each time point in the first training dataset to obtain the robotic arm grasping model.
3. The method according to claim 2, characterized in that, The construction of the first training dataset includes: When the user controls the robotic arm to perform a grasping task, the corresponding training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series are obtained. The first training dataset is constructed based on multiple training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series.
4. The method according to claim 3, characterized in that, When the user controls the robotic arm to perform a grasping task, the corresponding training task information, the time series corresponding to the training task information, and the training data collected at each time point in the time series include: When the user controls the robotic arm to perform a grasping task, the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point in the time series, the eye image corresponding to each time point in the time series, and the robotic arm training parameters corresponding to each time point in the time series are obtained. The eye image corresponding to each time point is input into the gaze point recognition model to determine the position of the first training gaze point corresponding to each time point in the time series. Based on the location of the first training gaze point and the training environment image, the training potential field map corresponding to each time point in the time series is determined.
5. The method according to claim 4, characterized in that, Before inputting the eye image corresponding to each time point into the fixation point recognition model to determine the position of the first training fixation point corresponding to each time point in the time series, the method further includes: Construct a second training dataset; the second training dataset includes: training eye images and the corresponding second training gaze point locations of the training eye images; The third neural network model is trained based on the second training dataset to obtain the gaze point recognition model.
6. The method according to claim 5, characterized in that, The construction of the second training dataset includes: The control display device displays a marker point image; the marker point image includes multiple marker points; As the user gazes at each of the marked points, the corresponding training eye image and the position of the second training gaze point corresponding to the training eye image are recorded.
7. The method according to claim 5, characterized in that, The third neural network model is a deep convolutional neural network model; The third neural network model includes: multiple convolutional layers, a first fully connected layer, and a second fully connected layer; the multiple convolutional layers are connected sequentially, and the first fully connected layer and the second fully connected layer are respectively connected to the last convolutional layer. The first fully connected layer is used to output the horizontal coordinate; The second fully connected layer is used to output the ordinate.
8. The method according to claim 4, characterized in that, The step of determining the training potential field map corresponding to each time point in the time series based on the first training fixation point position and the training environment image includes: Obtain the target location coordinates in the training environment image; the target location coordinates are any location coordinates other than the first training gaze point location; The relative distance is determined based on the target location coordinates and the location of the first training gaze point; Based on the relative distance, determine the potential field value of the target position coordinates; The training potential field map corresponding to the training environment image is determined based on the potential field value of each location coordinate in the training environment image.
9. The method according to claim 2, characterized in that, The first neural network model includes a long short-term memory network; The first neural network includes: an input layer, a hidden layer, an output layer, and an optimization network; the input layer, hidden layer, and output layer are connected in sequence; the optimization network is connected to the input layer, hidden layer, and output layer respectively. The input layer is used to divide the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the position of the first training gaze point corresponding to each time point according to the task information and the time series. The hidden layer is used to generate long-term memory features and short-term memory features based on the training task information, the time series corresponding to the training task information, the training environment image corresponding to each time point, and the position of the first training fixation point corresponding to each time point. The output layer is used to output the training results based on the long-term memory features and the short-term memory features; The optimized network includes a loss function used to optimize the parameters of the hidden layer.
10. The method according to claim 2, characterized in that, The second neural network model includes a long short-term memory network.
11. A robotic arm control device, characterized in that, The device includes: The acquisition module is used to acquire real-world images captured by the robotic arm and information about the tasks to be performed by the robotic arm. The gaze point calculation module is used to input the real environment image and the task information into the gaze point solving model to determine the corresponding gaze point information in the real environment image. The potential field map calculation module is used to determine the attention distribution potential field map based on the gaze point information and the real environment image; The control module is used to input the attention distribution potential field map and the real environment image into the robotic arm grasping model to obtain the motion parameters of the robotic arm, and control the movement of the robotic arm based on the motion parameters.
12. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 10.