Picking mechanical arm action optimization method and system combined with machine learning
By acquiring and analyzing the motion data of the harvesting robotic arm, and using machine learning models to generate motion association rules and mapping relationships, the adaptability and correlation problems of the harvesting robotic arm in motion control were solved, achieving efficient and accurate harvesting results and low damage rate.
Patent Information
- Application Number
- CN202511661944.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-11-13
AI Technical Summary
Existing robotic arms for harvesting lack adaptability and correlation in motion control, resulting in poor harvesting results, and the optimized motion parameters are difficult to convert into actual executable motion sequences.
By acquiring motion data sets of the harvesting robotic arm in different scenarios, correlation analysis is performed to generate motion association rules. A mapping relationship is established using a machine learning model to generate an initial motion optimization plan. Then, motion sequence adaptation processing is performed to ensure that the optimization plan is accurately executed in actual harvesting operations.
This improved the harvesting efficiency, accuracy, and adaptability of the robotic arm, reduced crop damage rates, and ensured the effective implementation of the optimized plan in practical applications.
Smart Images

Figure CN121105045A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, in particular to a picking mechanical arm action optimization method and system combined with machine learning. BACKGROUND
[0002] In the field of agricultural automation, picking mechanical arms are increasingly widely used, which can significantly improve picking efficiency and reduce labor costs, and are of great significance to ensure the stable supply of agricultural products. However, the existing picking mechanical arms have many deficiencies in action control.
[0003] On the one hand, the traditional picking mechanical arm action control method is often based on a preset fixed program, which lacks adaptability to different picking scenes. In actual picking process, picking scenes will vary greatly due to differences in crop types, growth states, environmental conditions and other factors. The fixed program is difficult to adjust the action of the mechanical arm in real time according to these changes, resulting in poor picking effect, and may cause incomplete picking, damage to crops and other problems.
[0004] On the other hand, the existing method usually does not fully consider the correlation between the movements of the joints of the mechanical arm and the coordination with the action of the picking execution component when optimizing the action of the picking mechanical arm. The mechanical arm is a complex motion system, and the movements of the joints influence each other. Only by comprehensively considering these factors can efficient and accurate picking action be achieved. Moreover, there is a lack of effective means to accurately convert the optimized action parameters into a mechanical arm executable action sequence, making it difficult to truly apply the optimization results to actual picking operations. SUMMARY
[0005] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide a picking mechanical arm action optimization method combined with machine learning, which comprises: obtaining a set of action data of a picking mechanical arm under different picking scenes, performing correlation analysis on the set of action data to generate action correlation rules, the set of action data comprising movement angle data, movement speed data of each joint of the mechanical arm and action duration data of the picking execution component; calling a pre-trained machine learning model, inputting the action correlation rules and a preset picking action optimization target into the machine learning model, establishing a mapping relationship between the action correlation rules and the picking action optimization target, and obtaining a mapping relationship table; generating an initial action optimization scheme of the picking mechanical arm based on the mapping relationship table, the initial action optimization scheme comprising adjustment angle parameters, adjustment speed parameters of each joint and adjustment duration parameters of the picking execution component; perform action sequence adaptation processing on the initial action optimization scheme to convert parameters in the initial action optimization scheme into an action sequence executable by the mechanical arm, to obtain an adapted action sequence; convert the adapted action sequence into picking mechanical arm action control instructions, and send the picking mechanical arm action control instructions to a picking mechanical arm control system, which drives the picking mechanical arm to perform picking actions after receiving the picking mechanical arm action control instructions.
[0006] In still another aspect, the embodiments of the present application also provide a picking mechanical arm action optimization method system combined with machine learning, which comprises a processor, a machine-readable storage medium, the machine-readable storage medium is connected with the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to realize the above-mentioned method.
[0007] Based on the above aspects, by obtaining a set of action data of the picking mechanical arm in different picking scenarios and performing correlation analysis, an action correlation rule is generated, a pre-trained machine learning model is called, a mapping relationship between the action correlation rule and a preset picking action optimization target is established, a mapping relationship table is obtained, a corresponding action adjustment strategy can be quickly found according to different optimization targets, and an initial action optimization scheme generated based on the mapping relationship table comprehensively considers the adjustment angle, speed of each joint and adjustment time length of the picking execution component and the like, so that comprehensive optimization of the picking mechanical arm action is realized. Action sequence adaptation processing is performed on the initial action optimization scheme to convert it into an action sequence executable by the mechanical arm, so that the optimization scheme can be accurately and correctly applied to actual picking work. Finally, the adapted action sequence is converted into action control instructions and sent to the picking mechanical arm control system, to drive the mechanical arm to perform the optimized picking action, effectively improving the picking efficiency, accuracy and adaptability of the picking mechanical arm and reducing the crop damage rate in the picking process. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 is an execution flow diagram of the picking mechanical arm action optimization method combined with machine learning provided by the embodiments of the present application.
[0009] Figure 2 is a schematic diagram of exemplary hardware and software components of the picking mechanical arm action optimization method system combined with machine learning provided by the embodiments of the present application. DETAILED DESCRIPTION
[0010] The present application will be specifically described below in conjunction with the drawings of the specification, Figure 1is a flowchart of a picking mechanical arm action optimization method combined with machine learning provided by an embodiment of the present application. The picking mechanical arm action optimization method combined with machine learning is described in detail below.
[0011] Step S110: Obtain a set of action data of the picking mechanical arm in different picking scenarios, perform correlation analysis on the set of action data, generate action correlation rules, and the set of action data includes motion angle data, motion speed data of each joint of the mechanical arm, and action duration data of a picking execution component.
[0012] In this embodiment, apple picking is taken as an application scenario, and the picking mechanical arm needs to perform picking actions in different apple growth states (if real size, fruit hanging position, and branch density). First, a set of action data in this scenario needs to be obtained and action correlation rules are generated.
[0013] Step S111: Receive a set of action data transmitted by a joint sensor and an action timer of the picking mechanical arm, the joint sensor is used to collect motion angle data and motion speed data of each joint of the mechanical arm, and the action timer is used to collect action duration data of the picking execution component.
[0014] In the apple picking scenario, the picking mechanical arm is configured with six rotary joints, each joint is installed with a high-precision joint sensor, which can collect motion angle data and motion speed data of each joint in the motion process in real time, and the sampling frequency is set to one hundred times per second. At the same time, an action timer is installed on the picking execution component (such as a pneumatic gripper) at the end of the mechanical arm, which is used to record the action duration data from the closure of the gripper to the grasping of the apple and then to the placing of the apple into the storage box. The joint sensor and the action timer transmit the collected data to a data processing terminal through wired Ethernet, and the data processing terminal receives these data and stores them as a set of action data, each data in the set of action data includes information such as collection time stamp, joint number, motion angle data, motion speed data, and action duration data of the picking execution component.
[0015] Step S112: Divide the set of action data according to picking scenario types, each picking scenario type corresponds to a set of scene action data subsets, and the picking scenario types are determined according to the morphological characteristics of the picking objects and the growth environment characteristics.
[0016] According to the morphological characteristics (if the real diameter, fruit weight, fruit stem length) and the growth environment characteristics (such as fruit hanging height, distance from branches and leaves, whether it is blocked by branches and leaves) of apples, the picking scene types are divided into large fruit unobstructed scene, medium fruit slightly obstructed by branches and leaves scene, small fruit severely obstructed by branches and leaves scene, high position large fruit scene and the like. The data processing terminal matches each data in the motion data set according to the preset scene classification standard, and classifies the data belonging to the same picking scene type together to form a scene motion data subset corresponding to the scene. For example, the data with fruit diameter greater than a certain value, fruit hanging height in a certain range and no branches and leaves are classified into a large fruit unobstructed scene motion data subset.
[0017] Step S113: performing data segmentation processing on each of the scene motion data subsets, and splitting each of the scene motion data subsets into a plurality of motion data blocks, each of the motion data blocks containing the motion angle data, motion speed data of each joint of the mechanical arm and the action duration data of the picking execution component within a single complete picking action cycle.
[0018] For each scene motion data subset, the data processing terminal performs segmentation according to the periodic characteristics of the picking action. A single complete picking action cycle is defined as the entire process from the initial position of the mechanical arm starting to move, positioning the apple, the action of the picking execution component, grabbing the apple and moving to the storage box position, and releasing the apple back to the initial position. By identifying the signals of the mechanical arm starting from the initial position and returning to the initial position in the action data, the starting point and the ending point of each complete picking action cycle are determined, and then the scene motion data subset is split into a plurality of motion data blocks according to these starting points and ending points. Each motion data block contains the motion angle data sequence, the motion speed data sequence of all joints and the action duration data of the picking execution component within the picking action cycle.
[0019] Step S114: extracting the associated features in each of the motion data blocks, the associated features including the corresponding relationship between the motion angle data and the motion speed data of the same joint, the cooperative relationship between the motion angle data of different joints, and the corresponding relationship between the action duration data of the picking execution component and the joint motion speed data.
[0020] For each motion data block, the data processing terminal needs to extract the associated features therein, and the specific process is as follows.
[0021] Step S1141: separating the motion angle data, motion speed data of each joint of the mechanical arm and the action duration data of the picking execution component from each of the motion data blocks, each data type corresponding to an independent data sequence.
[0022] The data processing terminal parses the action data block, extracts the motion angle data and the motion speed data of each joint according to the joint number and the data type field, forms a motion angle data sequence and a motion speed data sequence of each joint respectively, and extracts the action duration data of the picking execution component to form an independent action duration data sequence. For example, the motion angle data sequence of joint 1 is a series of angle values arranged in chronological order, the motion speed data sequence of joint 1 is a speed value corresponding to a time point, and the action duration data sequence of the picking execution component is a duration value of each picking action cycle.
[0023] Step S1142: For the motion angle data and the motion speed data of the same joint, a one-to-one correspondence is established according to the time node sequence, the motion angle data value and the motion speed data value of each time node form a data pair, and a plurality of data pairs form an angle-speed correspondence sequence of the joint. The angle-speed correspondence sequence is the correspondence between the motion angle data and the motion speed data of the same joint.
[0024] Taking joint 1 as an example, the data processing terminal matches the data in the motion angle data sequence and the motion speed data sequence of joint 1 according to the chronological order of the time stamps, and the motion angle data value and the motion speed data value of each same time node form a data pair. All these data pairs are arranged in chronological order to form an angle-speed correspondence sequence of joint 1. Other joints are also processed in the same way to generate their respective angle-speed correspondence sequences.
[0025] Step S1143: Select a joint group in the robot arm that has a motion coordination relationship, the joint group contains two or more joints that move simultaneously in the picking action, and extract the motion angle data sequence of each joint in the joint group.
[0026] According to the kinematic characteristics of the apple picking robot arm, a joint group having a motion coordination relationship is determined. For example, joint 2 (shoulder rotation joint) and joint 3 (elbow bending joint) need to move simultaneously during the extension and contraction of the robot arm, so they are divided into a joint group. The data processing terminal extracts the motion angle data sequence of joint 2 and joint 3 in the joint group from the action data block.
[0027] Step S1144: Calculate the time synchronization of the motion angle data sequence of each joint in the joint group, determine the synchronization level by comparing the time difference of the peak values in the motion angle data sequences of different joints, and the synchronization level has a corresponding relationship with the time difference.
[0028] For the motion angle data sequences of joint 2 and joint 3, the data processing terminal identifies all peak points in the two sequences, i.e. points at which the angle value reaches a local maximum. Then the time difference of each corresponding peak point appearing in the two sequences is calculated, and the average of all time differences is taken as an index of time synchronization. The smaller the time difference, the higher the level of synchronization; the larger the time difference, the lower the level of synchronization, thereby establishing a corresponding relationship between the level of synchronization and the time difference.
[0029] Step S1145: According to the level of synchronization and the amplitude ratio of the change of the motion angle data of each joint, a coordination relationship description of the motion angle data between different joints in the joint group is established, and the coordination relationship description includes a level of synchronization value and an angle change amplitude ratio value.
[0030] The total amplitude of the change of the angle in the motion angle data sequences of joint 2 and joint 3 is calculated, i.e. the difference between the maximum value and the minimum value in each sequence. Then the ratio of the total amplitude of the change of the angles of the two joints is calculated, and combined with the previously obtained level of synchronization value to form the coordination relationship description of the joint group. For example, the level of synchronization value is 0.8 (full score is 1), and the amplitude ratio of the change of the angles of joint 2 and joint 3 is 1.2:1, and these information together constitute the coordination relationship description.
[0031] Step S1146: Action duration data of the picking execution component is extracted from each of the action data blocks, the starting time point and the ending time point of the action duration are determined, the motion speed data sequences of all joints during the starting time point to the ending time point are extracted, and the average value and the change amplitude of the motion speed data of each joint in the corresponding time period are calculated.
[0032] The action duration data of the picking execution component is obtained from the action data block, and the starting time point and the ending time point of the action duration are determined according to the time stamp. Then the motion speed data sequences of all joints between the two time points are extracted, and for each joint speed sequence, the average value of all data points is calculated, and the difference between the maximum value and the minimum value is taken as the change amplitude.
[0033] Step S1147: A corresponding relationship between the action duration data of the picking execution component and the average value and the change amplitude of the motion speed data of each joint is established, each action duration data value corresponds to a set of motion speed average value and change amplitude value, forming a duration-speed corresponding relationship. The angle-speed corresponding relationship of the same joint, the coordination relationship of the motion angle data between different joints, and the duration-speed corresponding relationship of the picking execution component are integrated to form the association features of each of the action data blocks.
[0034] The action duration data value of picking execution component is associated with the average value of motion speed of each joint and the change amplitude, one action duration data value corresponds to the average value of speed and the change amplitude value of all joints, thereby forming the duration-speed correspondence. Finally, the angle-speed correspondence of the same joint, the coordination relationship of the joint group and the duration-speed correspondence are integrated to form the complete association characteristics of the action data block.
[0035] Step S115: Calculate the association characteristic similarity between different action data blocks in the same scene action data sub-block. The similarity of data sequence in different data blocks is calculated for each association characteristic to obtain independent similarity scores of each characteristic. According to the importance weight of each association characteristic in picking action, all independent similarity scores are integrated to obtain the overall association characteristic similarity.
[0036] In the large fruit unobstructed scene action data sub-block, two action data blocks A and B are selected. For the angle-speed correspondence relationship, the dynamic time warping algorithm is used to calculate the similarity of the angle-speed correspondence sequence of the same joint in the two data blocks to obtain the independent similarity score of this characteristic. For the coordination relationship characteristic, the similarity is calculated by comparing the difference between the synchronization level value and the angle change amplitude ratio to obtain the corresponding independent similarity score. For the duration-speed correspondence relationship characteristic, the difference of action duration data value and the difference of average value and change amplitude of each joint speed are calculated to obtain the independent similarity score. According to the requirements of picking action, the importance weights of the three association characteristics are assigned, for example, the weight of angle-speed correspondence relationship is 0.4, the weight of coordination relationship is 0.3, and the weight of duration-speed correspondence relationship is 0.3. The independent similarity score of each characteristic is multiplied by its corresponding weight, and then the product is added to obtain the overall association characteristic similarity between action data blocks A and B.
[0037] Step S116: According to the association characteristic similarity, the action data block group with similar association characteristics is screened out, and each action data block group contains multiple action data blocks with association characteristic similarity meeting the preset condition.
[0038] The preset condition of association characteristic similarity is set to be greater than or equal to 0.7. The data processing terminal calculates the overall association characteristic similarity between all action data blocks in the same scene action data sub-block, and the action data blocks with similarity greater than or equal to 0.7 are grouped into one group to form multiple action data block groups. The action data blocks in each action data block group have similar association characteristics, which reflect the mechanical arm action mode under similar picking conditions.
[0039] Step S117: Analyze the commonalities of the associated features in each action data block group, and extract rule entries that can reflect the association patterns of the action data blocks in the action data block group. Each rule entry contains a description of the specific correspondence of the associated features.
[0040] For each motion data block group, the data processing terminal analyzes the correlation characteristics of all motion data blocks within the group to identify commonalities. For example, in a certain motion data block group, the angle-velocity correspondence sequence of joint 1 in all data blocks shows a pattern where the velocity first increases and then decreases as the angle increases. The synchronization level in the coordination relationship between joints 2 and 3 is above 0.75, the ratio of angle change amplitude is stable at around 1.2:1, and the motion duration data of the picking and executing component is positively correlated with the average velocity of joint 4. These commonalities are organized into rule entries, and each rule entry describes in detail the specific correspondence between the correlation characteristics.
[0041] Step S118: Classify and organize all the rule entries according to the picking scenario type. Each picking scenario type corresponds to a set of rule entries, forming a complete action association rule.
[0042] Step S1181: Establish a classification directory of picking scene types. The classification directory contains all the confirmed picking scene type names and corresponding scene feature descriptions. The scene feature descriptions include the specific content of the morphological characteristics of the picking objects and the characteristics of the growing environment.
[0043] The data processing terminal establishes a classification directory for harvesting scene types. This directory lists all the categorized harvesting scene types, such as large fruit unobstructed scene and medium fruit slightly obstructed by branches and leaves scene. Each scene type has a detailed description of its characteristics. For example, the characteristics of the large fruit unobstructed scene are: fruit diameter greater than a specific value, fruit weight within a specific range, moderate stem length, fruit hanging height between 1.5 meters and 2.5 meters, and no branches or leaves obstructing the fruit, allowing for direct harvesting.
[0044] Step S1182: Traverse each rule entry and extract the scene information corresponding to the associated features involved in the rule entry. The scene information includes the picking scene type to which the action data block from which the rule entry originates belongs.
[0045] The data processing terminal iterates through all generated rule entries one by one, extracting the harvesting scenario type of the action data block from which each rule entry originates from its metadata. This information is the scenario information corresponding to the rule entry.
[0046] Step S1183: Based on the extracted scene information, assign the rule entries to the corresponding picking scene type in the classification directory to form the initial rule group for each scene type.
[0047] According to the extracted scene information, each rule item is classified into the corresponding picking scene type in the classification catalog. For example, the rule item from the large fruit unobstructed scene action data block is assigned to this scene type, and all rule items assigned to the same scene type form the initial rule group of this scene type.
[0048] Step S1184: Rule conflict detection is performed on the initial rule group of each scene type, and the associated feature descriptions and corresponding relationships between different rule items in the initial rule group are compared. If there are rule items with consistent associated feature descriptions but different corresponding relationships, they are marked as conflict rule pairs.
[0049] For the initial rule group of each scene type, the data processing terminal compares the rule items in the group two by two. The associated feature descriptions between the rule items are compared to see if they are consistent, and the corresponding relationships between the associated features are compared to see if they are the same. If the two rule items describe the same associated features but have different corresponding relationships, for example, both describe the angle-velocity relationship of joint 1, but one rule states that the velocity increases when the angle increases, and the other rule states that the velocity decreases when the angle increases, then the two rule items are marked as conflict rule pairs.
[0050] Step S1185: For each conflict rule pair, the associated feature similarity statistical results of the action data block group from which it comes are queried, and the rule item with the largest data block group size and the largest average associated feature similarity is retained, and the other conflicting rule item is deleted.
[0051] For each conflict rule pair, the data processing terminal queries the relevant statistical information of the action data block group from which each rule item comes, including the size of the data block group (i.e. the number of action data blocks in the group) and the average value of the associated feature similarity in the group. The size and similarity average value of the two data block groups are compared, the rule item with the larger data block group size and higher similarity average value is retained, and the other conflicting rule item is deleted from the initial rule group.
[0052] Step S1186: Supplement the missing rule item corresponding to the associated feature in the initial rule group of each scene type. If there is no corresponding rule item for any associated feature in this scene type, select the rule item with the highest association degree from the rule group of the similar picking scene type and adapt it to form a supplementary rule item.
[0053] Check the initial rule set of each scene type to see if it covers all the corresponding rule entries of the associated features. If a certain associated feature is found to be missing in the rule entry under this scene type, for example, the joint 5 and joint 6 cooperative relationship rule entry is missing in the medium-sized fruit branch and leaf slightly obstructed scene, find a similar scene type to this scene type, such as the rule set of the large-sized fruit unobstructed scene, calculate the correlation degree of each rule entry in the scene rule set to the missing associated feature, and select the rule entry with the highest correlation degree. According to the characteristics of the medium-sized fruit branch and leaf slightly obstructed scene, the selected rule entry is modified, for example, the synchronization level threshold and the angle change amplitude ratio in the cooperative relationship are adjusted, a supplementary rule entry is formed, and is added to the initial rule set of this scene type.
[0054] Step S1187: Logically sort the rule set of each scene type according to the action execution order involved by the associated features, divide the rule entries into joint motion association rules, execution component action association rules, and cooperative action association rules, and after sorting according to the complexity of the associated features, add the sorted rule set to the scene identifier, which contains the picking scene type name and the rule set version number.
[0055] For the rule set of each scene type, classify according to the action execution order involved by the associated features. Rule entries describing the relationship between single joint angle and speed are classified as joint motion association rules; rule entries describing the relationship between picking execution component action duration and joint speed are classified as execution component action association rules; rule entries describing the cooperative motion relationship of multiple joints are classified as cooperative action association rules. Within each classification, sort according to the complexity of the associated features from low to high, for example, first arrange the angle-speed corresponding rules of single joint, and then arrange the cooperative rules of multiple joints. After sorting, add the scene identifier to the rule set, which consists of the picking scene type name and the rule set version number, for example, "large-sized fruit unobstructed scene_V1.0".
[0056] Step S1188: Summarize the rule sets of all picking scene types to generate a rule index table, which contains the scene type name, the rule set version number, and the number of rule entries.
[0057] The data processing terminal summarizes the rule sets of all scene types, and then generates a rule index table. The rule index table records the name of each scene type, the corresponding rule set version number, and the number of rule entries contained in the rule set, facilitating subsequent quick query and call of the corresponding rule set.
[0058] Step S1189: Integrate the summarized rule sets and the rule index table to form a complete action association rule, and store it in the rule database.
[0059] All the scene type rule groups and rule index tables after summarization are integrated together to form a complete action association rule system, which is then stored in a pre-established rule database. The rule database adopts a relational database structure, and corresponding data tables are established for rule groups and rule entries, facilitating data management, updating and querying.
[0060] Step S120: The pre-trained machine learning model is called, the action association rule and the preset picking action optimization target are input into the machine learning model, the mapping relationship between the action association rule and the picking action optimization target is established, and a mapping relationship table is obtained.
[0061] In the apple picking scene, after the picking action optimization target is preset, the corresponding action association rule is called from the rule database, and the two are input into the pre-trained machine learning model to establish a mapping relationship and generate a mapping relationship table through model operation.
[0062] Step S121: The preset picking action optimization target is determined, and the picking action optimization target includes a picking action completion efficiency target, a picking action smoothness target and a picking object damage rate control target. Each optimization target has a corresponding description index.
[0063] The preset picking action optimization target is specifically set as follows: the description index of the picking action completion efficiency target is the number of picking actions completed per unit time, which is required to achieve a specific picking frequency under the premise of ensuring picking quality; the description index of the picking action smoothness target is the speed fluctuation amplitude of each joint movement of the mechanical arm, which is required to be controlled within a specific range; and the description index of the picking object damage rate control target is the proportion of damage such as indentation and scratch on the surface of the apple during picking, which is required to be lower than a specific threshold.
[0064] Step S122: The action association rule is disassembled into a plurality of rule units, each of which corresponds to a specific association feature description, and the picking action optimization target is disassembled into a plurality of target units, each of which corresponds to a description index of an optimization target.
[0065] The data processing terminal disassembles the action association rule, and for each rule group under each picking scene type, each rule entry is split into an independent rule unit according to the association feature description. For example, the rule entry “when the angle of joint 1 increases, the speed first increases and then decreases” is disassembled into an independent rule unit, which only corresponds to the specific association feature description of the angle-velocity correspondence of joint 1. Meanwhile, the picking action completion efficiency target, the picking action smoothness target, and the picking object damage rate control target are disassembled into target units, and each target unit corresponds to a description index. For example, the “picking frequency per unit time” target unit corresponds to the description index of the picking action completion efficiency target, the “joint speed fluctuation amplitude” target unit corresponds to the description index of the picking action smoothness target, and the “apple damage ratio” target unit corresponds to the description index of the picking object damage rate control target.
[0066] Step S123: input the rule unit and the target unit into a pre-trained machine learning model, and the machine learning model includes a feature input layer, an association modeling layer, and a mapping output layer.
[0067] After all the rule units and target units obtained by disassembly are arranged according to a preset format, they are input into a pre-trained machine learning model. The machine learning model is a three-layer network structure, including a feature input layer, an association modeling layer, and a mapping output layer, and the layers are connected by full connection to realize the association modeling and mapping relationship output between the rule units and the target units.
[0068] Step S124: in the feature input layer, the rule unit and the target unit are subjected to feature coding processing, and the rule unit and the target unit in text form are converted into feature data in vector form, each rule unit corresponds to a rule feature vector, and each target unit corresponds to a target feature vector.
[0069] The feature input layer first pre-processes the rule unit and the target unit in text form, including removing redundant expressions, unifying term expressions, etc. Then, the word embedding method is used to code the pre-processed text, each word is converted into a fixed-dimensional vector, and then the word vector is subjected to average pooling processing to generate a rule feature vector corresponding to the rule unit and a target feature vector corresponding to the target unit. Each rule feature vector and target feature vector is a fixed-length numerical sequence, which can represent the core feature information of the corresponding unit.
[0070] Step S125: input the rule feature vector and the target feature vector into the association modeling layer, and the association modeling layer adopts a multi-layer perceptron structure to calculate the association degree value between the rule feature vector and the target feature vector through linear transformation and a nonlinear activation function.
[0071] The association modeling layer is composed of three hidden layers and adopts a multi-layer perceptron structure. The rule feature vector and the target feature vector are spliced and input into the first hidden layer, linearly transformed by the weight matrix of the first hidden layer, and then output to the second hidden layer after being processed by a nonlinear activation function. The second hidden layer and the third hidden layer repeat the linear transformation and nonlinear activation processing process, gradually mining the deep association between the rule feature vector and the target feature vector. Finally, the association degree value between the two is calculated through the output layer of the association modeling layer, and the association degree value is used to represent the association closeness between the rule unit and the target unit.
[0072] Step S126: determining the matching relationship between the rule unit and the target unit according to the association degree value, and the rule unit and the target unit forming a matching pair when the association degree value meets a preset condition, each matching pair including one rule unit and one corresponding target unit.
[0073] The preset condition of the association degree value is set to be greater than or equal to 0.6. The data processing terminal compares the association degree value output by the association modeling layer with the preset condition. If the association degree value of a rule unit and a target unit is greater than or equal to 0.6, it is determined that the two have a matching relationship and form a matching pair. For example, the rule unit of “joint 2 and joint 3 synchronization level 0.75 or above” and the target unit of “joint speed fluctuation amplitude” have an association degree value of 0.72, which meets the preset condition, and the two form a matching pair.
[0074] Step S127: grouping all the matching pairs according to the picking scene type, and the matching pairs under each picking scene type forming a local mapping relationship of the picking scene type.
[0075] All the matching pairs are grouped according to the picking scene type to which the rule unit in each matching pair belongs. For example, the matching pairs of the rule unit of the large fruit unobstructed scene are grouped into one group, forming a local mapping relationship of this scene type, which only reflects the matching relationship between the rule unit and the target unit in the large fruit unobstructed scene.
[0076] Step S128: integrating the local mapping relationships of each picking scene type, supplementing the association connection relationship of the matching pairs between different scenes, and forming a global mapping relationship covering all picking scene types.
[0077] The local mapping relationships of all picking scene types are summarized, and the commonalities and differences between different matching pairs are analyzed. For matching pairs in different scenes that have a correlation, the correlation and connection relationship between them is supplemented. For example, in the medium-sized fruit branch and leaf slightly blocked scene, the matching pair "the action duration of the picking execution component is positively correlated with the joint 4 speed" is correlated with the similar matching pair in the large fruit unblocked scene, and the connection relationship between the two in the speed threshold and duration range is supplemented to form a global mapping relationship covering all picking scene types.
[0078] Step S129: The global mapping relationship is presented in table form to obtain a mapping relationship table. The row dimension of the table is a rule unit, and the column dimension is a target unit. The corresponding correlation value and matching condition are recorded in the table cell.
[0079] The global mapping relationship is converted into a table form. The row dimension of the table lists all rule units in turn, and each rule unit occupies a row. The column dimension lists all target units in turn, and each target unit occupies a column. In each cell of the table, the correlation value and matching condition (i.e., whether the correlation value meets the preset condition) of the corresponding rule unit and target unit are recorded. For example, a certain cell records "correlation value: 0.68, matching condition: meets", which clearly shows the mapping relationship between the rule unit and the target unit, thereby obtaining a mapping relationship table.
[0080] Step S130: An initial action optimization scheme of the picking robot arm is generated based on the mapping relationship table. The initial action optimization scheme includes adjustment angle parameters of each joint, adjustment speed parameters, and adjustment duration parameters of the picking execution component.
[0081] In the apple picking scene, according to the matching relationship and correlation value of the rule unit and the target unit in the mapping relationship table, the initial action optimization scheme including the adjustment parameters of each joint and the adjustment parameters of the picking execution component is calculated and generated based on the original action parameters.
[0082] Step S131: The rule unit corresponding to the current picking scene type and the matched target unit are extracted from the mapping relationship table to determine the correlation features that need to be optimized and the corresponding optimization target description index in the current picking scene type.
[0083] According to the scene type of the current apple picking (such as the large fruit unblocked scene), all rule units corresponding to the scene type and the target units matched with these rule units are selected from the mapping relationship table. By analyzing these rule units and target units, the correlation features that need to be optimized in the current scene are determined, such as the angle-speed correspondence of joint 1, the cooperative relationship between joint 2 and joint 3, etc., and the corresponding optimization target description index is also determined, such as the number of picking times per unit time, the joint speed fluctuation amplitude, etc.
[0084] Step S132: According to the associated feature description in the rule unit, locate the original action parameters corresponding to the associated feature in the action data set, which include the original motion angle data, original motion speed data of each joint of the mechanical arm, and the original action duration data of the picking execution component.
[0085] According to the specific description of the associated feature in the rule unit, search and match in the action data set. For example, for the rule unit of "joint 1 angle increases, speed first increases and then decreases", locate the original motion angle data sequence and original motion speed data sequence of joint 1 in the action data set; for the rule unit of "synchronization level of joint 2 and joint 3 is above 0.75", locate the original motion angle data sequence of joint 2 and joint 3; at the same time, locate the original action duration data sequence of the picking execution component, which together constitute the original action parameters.
[0086] Step S133: Refer to the optimization target description index in the target unit, input the original action parameters into a predefined performance evaluation model, which outputs the estimated performance index value based on the current action parameters, and calculates the difference between the estimated performance index value and the optimization target description index.
[0087] The predefined performance evaluation model is trained based on historical picking data, which can predict the picking performance according to the input original action parameters. After inputting the original action parameters into the model, the model outputs the estimated performance index values such as the number of picking times per unit time, joint speed fluctuation amplitude, and apple damage ratio. Compare the above estimated performance index values with the corresponding optimization target description index respectively, calculate the difference between them, and get the difference between the estimated performance index value and the optimization target description index.
[0088] Step S134: According to the difference and the correlation degree value in the mapping relationship table, determine the adjustment direction of each original action parameter, which is determined according to the positive and negative attributes of the difference and the size of the correlation degree value.
[0089] The positive and negative attributes of the difference are analyzed. If the estimated performance index value is lower than the optimization target description index (for example, the estimated picking frequency per unit time is lower than the target frequency), the original action parameters need to be adjusted to improve performance. If the estimated performance index value is higher than the optimization target description index (for example, the estimated joint speed fluctuation amplitude is greater than the target amplitude), the original action parameters need to be adjusted to reduce the fluctuation. At the same time, the correlation degree value of the corresponding rule unit and the target unit in the mapping relationship table is combined. The greater the correlation degree value, the greater the influence of the rule unit on the target unit. The determination of the adjustment direction needs to give priority to the original action parameters corresponding to the rule unit. Based on the above factors, the adjustment direction of each original action parameter is determined, such as increasing the movement speed of joint 1, reducing the angle change amplitude ratio of joint 2 and joint 3, and the like.
[0090] Step S135: Based on the adjustment direction and the preset adjustment step rule, the adjustment angle parameter, the adjustment speed parameter and the adjustment time length parameter of the picking execution component of each joint are calculated. The adjustment angle parameter is the superposition result of the original movement angle data and the adjustment amount, the adjustment speed parameter is the superposition result of the original movement speed data and the adjustment amount, and the adjustment time length parameter is the superposition result of the original action time length data and the adjustment amount.
[0091] Step S1351: The preset adjustment step rule is obtained, and the adjustment step rule includes adjustment step coefficients corresponding to different correlation degree values. The correlation degree value and the adjustment step coefficient have a corresponding relationship.
[0092] The preset adjustment step rule is obtained from the system configuration file. The adjustment step rule clearly shows the adjustment step coefficients corresponding to different correlation degree value intervals. For example, when the correlation degree value is between 0.9 and 1.0, the adjustment step coefficient is 0.1; when the correlation degree value is between 0.8 and 0.9, the adjustment step coefficient is 0.08, and so on, establishing the corresponding relationship between the correlation degree value and the adjustment step coefficient.
[0093] Step S1352: According to the correlation degree value of the rule unit and the target unit in the mapping relationship table, the corresponding adjustment step coefficient is selected from the adjustment step rule.
[0094] For each original action parameter that needs to be adjusted, the correlation degree value of the corresponding rule unit and the target unit in the mapping relationship table is found, and the corresponding adjustment step coefficient is matched and selected from the adjustment step rule according to the correlation degree value. For example, the correlation degree value of a certain original action parameter is 0.85, and the selected adjustment step coefficient is 0.08.
[0095] Step S1353: According to the adjustment direction, the positive and negative attributes of the adjustment amount are determined for the original movement angle data of each joint of the mechanical arm. When the adjustment direction is to increase the angle, the adjustment amount is positive. When the adjustment direction is to decrease the angle, the adjustment amount is negative.
[0096] For the original motion angle data of each joint of the mechanical arm, the positive and negative of the adjustment amount is determined according to the determined adjustment direction. If the adjustment direction is to increase the joint motion angle, the adjustment amount is set to a positive value; if the adjustment direction is to reduce the joint motion angle, the adjustment amount is set to a negative value.
[0097] Step S1354: Calculate the adjustment amount of the joint adjustment angle, which is equal to the original motion angle data multiplied by the adjustment step coefficient, and the adjustment angle parameter is equal to the original motion angle data plus the adjustment amount.
[0098] Taking the original motion angle data of joint 1 as an example, the original data is multiplied by the selected adjustment step coefficient to obtain the adjustment amount of the adjustment angle of joint 1. If the original motion angle data is A, the adjustment step coefficient is k, the adjustment amount is A x k, and the adjustment angle parameter is A + (A x k). Other joints obtain their respective adjustment angle parameters according to the same calculation method.
[0099] Step S1355: For the original motion speed data of each joint of the mechanical arm, the positive and negative of the adjustment amount is determined according to the adjustment direction. The adjustment amount is positive when the adjustment direction is to increase the speed, and the adjustment amount is negative when the adjustment direction is to reduce the speed.
[0100] For the original motion speed data of each joint of the mechanical arm, the positive and negative of the adjustment amount is determined according to the adjustment direction. The adjustment amount is positive when the adjustment direction is to increase the joint motion speed, and the adjustment amount is negative when the adjustment direction is to reduce the joint motion speed.
[0101] Step S1356: Calculate the adjustment amount of the joint adjustment speed, which is equal to the original motion speed data multiplied by the adjustment step coefficient, and the adjustment speed parameter is equal to the original motion speed data plus the adjustment amount.
[0102] Taking the original motion speed data of joint 2 as an example, it is multiplied by the corresponding adjustment step coefficient to obtain the adjustment amount of the adjustment speed of joint 2. If the original motion speed data is V, the adjustment step coefficient is k, the adjustment amount is V x k, and the adjustment speed parameter is V + (V x k). The remaining joints are calculated in the same way to obtain the adjustment speed parameter.
[0103] Step S1357: For the original action duration data of the picking execution component, the positive and negative of the adjustment amount is determined according to the adjustment direction. The adjustment amount is positive when the adjustment direction is to extend the duration, and the adjustment amount is negative when the adjustment direction is to shorten the duration.
[0104] For the original action duration data of the picking execution component, the positive and negative of the adjustment amount is determined according to the adjustment direction. The adjustment amount is positive when the adjustment direction is to extend the action duration, and the adjustment amount is negative when the adjustment direction is to shorten the action duration.
[0105] Step S1358: Calculate the adjustment amount of the picking execution component adjustment duration, which is equal to the original action duration data multiplied by the adjustment step coefficient. The adjustment duration parameter is equal to the original action duration data plus the adjustment amount.
[0106] The original action duration data of the picking execution component is multiplied by the corresponding adjustment step coefficient to obtain the adjustment amount of the adjustment duration. If the original action duration data is T, the adjustment step coefficient is k, the adjustment amount is T x k, and the adjustment duration parameter is T + (T x k).
[0107] Step S1359: Record the calculation process of the adjustment angle parameter, the adjustment speed parameter of each joint, and the adjustment duration parameter of the picking execution component, including the original data value, the adjustment step coefficient, the adjustment amount, and the final parameter value.
[0108] The original motion angle data of each joint, the adjustment step coefficient, the adjustment amount, the adjustment angle parameter, and the original motion speed data, the adjustment step coefficient, the adjustment amount, the adjustment speed parameter, and the original action duration data of the picking execution component, the adjustment step coefficient, the adjustment amount, and the adjustment duration parameter are stored in the log file for subsequent tracing and verification.
[0109] Step S136: Arrange the calculated adjustment angle parameters, adjustment speed parameters of each joint, and adjustment duration parameters of the picking execution component according to the action execution sequence. Each action execution stage corresponds to a set of parameter combinations.
[0110] According to the action execution sequence of apple picking, the picking process is divided into initial positioning stage, mechanical arm stretching stage, picking execution stage, mechanical arm recovery stage, fruit release stage, etc. The calculated adjustment angle parameters, adjustment speed parameters of each joint, and adjustment duration parameters of the picking execution component are distributed according to the requirements of each action execution stage, and each stage corresponds to a set of parameter combinations containing related joint parameters and execution component parameters. For example, the mechanical arm stretching stage corresponds to the adjustment angle and speed parameter combination of joints 2 and 3, and the picking execution stage corresponds to the adjustment duration parameter of the picking execution component and the adjustment speed parameter combination of joint 4.
[0111] Step S137: Supplement the transition connection parameters between each parameter combination, which include the transition time and speed change gradient of joint adjustment between adjacent action stages.
[0112] The parameter change requirement between adjacent action execution stages is analyzed, and the transition connection parameter is supplemented. The transition time is set as the time interval of parameter adjustment between adjacent two stages, to ensure smooth transition of the robot arm action; the speed change gradient is set as the speed change rate of the joint between adjacent stages, to avoid action instability caused by sudden speed change. For example, the transition time from the robot arm stretching stage to the picking execution stage is set to a specific value, and the speed change gradient of joint 2 is set to a specific change rate, to ensure natural action connection.
[0113] Step S138: integrate the parameter combination with the transition connection parameter to form an initial action optimization scheme containing a complete action optimization parameter system.
[0114] The parameter combination of each action execution stage and the transition connection parameter are integrated in time sequence to form a complete parameter system covering the entire picking action process. The complete parameter system contains all parameters such as the adjustment angle, adjustment speed of each joint in different stages, the adjustment time of the picking execution component, and the transition time and speed change gradient between stages, which together constitute the initial action optimization scheme.
[0115] Step S140: perform action sequence adaptation processing on the initial action optimization scheme to convert the parameters in the initial action optimization scheme into an action sequence executable by the robot arm, to obtain an adapted action sequence.
[0116] For each parameter in the initial action optimization scheme, format conversion and time sequence arrangement are performed to convert it into an action sequence that can be recognized and executed by the robot arm control system, i.e. an adapted action sequence.
[0117] Step S141: extract the adjustment angle parameter, adjustment speed parameter, adjustment time parameter of the picking execution component, and transition connection parameter of each joint in the initial action optimization scheme.
[0118] From the initial action optimization scheme, the adjustment angle parameter sequence, adjustment speed parameter sequence, adjustment time parameter of the picking execution component, and transition time and speed change gradient in the transition connection parameter are extracted one by one, to prepare for subsequent action sequence conversion.
[0119] Step S142: convert the adjustment angle parameter of each joint into an angle control sequence of the robot arm joint movement, each angle control sequence containing joint angle target values corresponding to multiple time nodes, and the interval of time nodes is determined according to the adjustment speed parameter.
[0120] For each joint, the interval of time nodes is determined according to the adjustment speed parameter of the joint. The faster the adjustment speed, the smaller the interval of time nodes; the slower the adjustment speed, the larger the interval of time nodes. At each time node, the corresponding joint angle target value is determined, and the above time nodes and angle target values are arranged in order to form the angle control sequence of the joint. For example, the adjustment speed parameter of joint 1 is high, the interval of time nodes is set to a small value, and each time node corresponds to an angle target value that gradually increases to form the angle control sequence of joint 1.
[0121] Step S143: converting the adjustment speed parameter of each joint into a speed control sequence of the joint movement of the robot arm, each speed control sequence containing joint speed target values corresponding to multiple time nodes, and the speed control sequence and the angle control sequence being synchronized at the time nodes.
[0122] The adjustment speed parameter of each joint is assigned to the corresponding time node based on the time nodes of the angle control sequence, the joint speed target value at each time node is determined, and the speed control sequence is formed. The time nodes of the speed control sequence are completely consistent with the time nodes of the angle control sequence to realize synchronous control of the angle and the speed. For example, the angle control sequence of joint 2 is set to T1, T2, T3, … at a certain time node, and the speed control sequence of joint 2 is set to speed target values at the corresponding time nodes T1, T2, T3, …, so that at each time node, joint 2 can reach the preset angle target and maintain the corresponding speed target, avoiding motion jamming or overshoot caused by mismatch between the angle and the speed.
[0123] Step S144: converting the adjustment duration parameter of the picking execution component into a motion control sequence of the picking execution component, the motion control sequence containing a start time node, a motion maintenance time node, and a stop time node of the picking execution component, and the motion maintenance time node corresponding to the adjustment duration parameter.
[0124] According to the adjustment duration parameter of the picking execution component, the starting time node, the action maintenance time node and the stopping time node of the picking execution component are determined in combination with the time axis of the overall motion of the mechanical arm. The starting time node is set as the time when the mechanical arm completes positioning and the picking execution component starts to act; the action maintenance time node is determined according to the adjustment duration parameter, that is, the time node reached after a time period equal to the adjustment duration parameter from the starting time node, which represents the time when the picking execution component completes the core action such as grabbing or shearing; and the stopping time node is set as the time when the picking execution component returns to the initial state after completing the action. The three time nodes are arranged in sequence, and the action state (start, maintain, stop) of each node corresponding to the execution component is labeled to form the action control sequence of the picking execution component. For example, if the adjustment duration parameter of the picking execution component is a specific duration, the starting time node is set as the time T5 when the joint of the mechanical arm is adjusted to the picking position, the action maintenance time node is T5 plus the time T6 corresponding to the adjustment duration parameter, and the stopping time node is set as the time T7 when the execution component is reset after T6, thereby forming the action control sequence of the execution component.
[0125] Step S145: converting the transition connection parameter into a transition control sequence, the transition control sequence including a joint angle transition target value, a speed transition target value and a transition time node between adjacent action stages, the transition control sequence being used to connect the adjacent angle control sequence and the speed control sequence.
[0126] For the transition time and speed change gradient in the transition connection parameter, the transition time node, the joint angle transition target value and the speed transition target value are determined in combination with the angle control sequence and the speed control sequence of the adjacent action stages. The transition time node is set as the connection time between the end of the previous action stage and the start of the next action stage; the joint angle transition target value is set as the transition value between the final angle target value of the previous stage and the initial angle target value of the next stage, to ensure smooth change of the angle; and the speed transition target value is determined according to the speed change gradient, that is, the intermediate value gradually adjusted from the current speed target value to the next stage speed target value according to the gradient. The transition time node, the angle transition target value and the speed transition target value of each joint are integrated in time sequence to form the transition control sequence. For example, in the transition process between the end of the mechanical arm stretching stage and the start of the picking execution stage, the transition time node is set as the final time node T8 of the stretching stage, the angle transition target value of joint 2 is set as the intermediate value between the final angle of the stretching stage and the initial angle of the picking stage, and the speed transition target value is set as the transition value from the stretching stage speed to the initial speed of the picking stage according to the preset gradient, which together constitute the transition control sequence of the transition process.
[0127] Step S146: integrate the angle control sequence, the speed control sequence, the action control sequence of the picking execution component, and the transition control sequence according to the time node sequence, each time node corresponding to a complete set of control parameters.
[0128] For example, step S1461: extract all time nodes in the angle control sequence to form an angle time node set; extract all time nodes in the speed control sequence to form a speed time node set; extract all time nodes in the action control sequence of the picking execution component to form an execution time node set; extract all time nodes in the transition control sequence to form a transition time node set.
[0129] Traverse the angle control sequence of each joint, collect all time nodes therein, and form an angle time node set after removing duplicate nodes; in the same way, extract time nodes from the speed control sequence, the action control sequence of the picking execution component, and the transition control sequence, respectively, to form a speed time node set, an execution time node set, and a transition time node set. The time nodes in each set are preliminarily arranged in chronological order to prepare for subsequent merging.
[0130] Step S1462: merge the angle time node set, the speed time node set, the execution time node set, and the transition time node set, remove duplicate time nodes, and sort them in chronological order to form a unified time node sequence.
[0131] Put all nodes in the four time node sets into a temporary set, delete duplicate nodes by comparing the time information of the nodes, and retain unique time nodes. Then, sort the above unique nodes in chronological order to form a unified time node sequence throughout the picking action process, ensuring that the time bases of all control sequences are consistent.
[0132] Step S1463: traverse each time node in the unified time node sequence, find the joint angle target value corresponding to the time node from the angle control sequence, and if the time node does not exist directly in the angle control sequence, perform linear interpolation according to the joint angle target values of adjacent time nodes to obtain the supplementary angle target value of the time node.
[0133] Each node in the uniform time node sequence is traversed one by one, and for each node, it is checked in the angle control sequence of each joint whether there is a same time node. If there is, the joint angle target value corresponding to the node is directly extracted; if not, the two adjacent time nodes before and after the node (i.e. the previous time node and the subsequent time node) are found, the joint angle target values of the two nodes in the angle control sequence are obtained, and the supplementary angle target value of the current time node is calculated by linear interpolation, so as to ensure the continuity of the joint angle on the time axis. For example, the node T9 in the uniform time node sequence is located between T8 and T10 in the angle control sequence, the joint 3 angle target value corresponding to T8 is A8, and the angle target value corresponding to T10 is A10, then the supplementary angle target value corresponding to T9 is calculated by linear interpolation.
[0134] Step S1464: The joint speed target value corresponding to each time node is found in the speed control sequence, and if the time node does not exist directly in the speed control sequence, the supplementary speed target value of the time node is obtained by linear interpolation according to the joint speed target values of the adjacent time nodes.
[0135] The same as the acquisition method of the angle target value, for each node in the uniform time node sequence, a search is performed in the speed control sequence of each joint. When the corresponding node exists, the speed target value is directly extracted; when the corresponding node does not exist, the supplementary speed target value is obtained by linear interpolation according to the speed target values of the adjacent nodes, so as to ensure the smooth transition of the joint speed in the entire motion process and avoid sudden changes in speed.
[0136] Step S1465: The execution component action state corresponding to each time node is found in the picking execution component action control sequence. The execution component action state includes a start state, a maintenance state and a stop state. If the time node is between the start time node and the maintenance time node, the state is the maintenance state; if the time node is between the maintenance time node and the stop time node, the state is the stop state.
[0137] For each time node in the uniform time node sequence, it is compared with the start time node, the action maintenance time node and the stop time node in the picking execution component action control sequence. If the time node is equal to the start time node, the action state is the start state; if the time node is between the start time node and the action maintenance time node, the action state is the maintenance state; if the time node is equal to the action maintenance time node, the action state is the critical state of the transition from the maintenance state to the stop state; if the time node is between the action maintenance time node and the stop time node, the action state is the stop state; if the time node is equal to or later than the stop time node, the action state is the stop completion state.
[0138] Step S1466: Find the transition parameter corresponding to each time node in the transition control sequence, the transition parameter includes the joint angle transition target value, the speed transition target value, if the time node is in the transition time node range, the corresponding transition parameter is extracted, if it is out of the transition time node range, the transition parameter is set to zero.
[0139] Match each node in the unified time node sequence with the transition time node range in the transition control sequence, the transition time node range is the interval from the transition start time node to the transition end time node. If the current time node is in the interval, the corresponding joint angle transition target value and speed transition target value are extracted from the transition control sequence; if the current time node is before the transition start time node or after the transition end time node, the joint angle transition target value and speed transition target value corresponding to the node are both set to zero, indicating that no transition control needs to be performed at this time.
[0140] Step S1467: Combine the joint angle target value or supplementary angle target value, joint speed target value or supplementary speed target value, execution component action state and transition parameter corresponding to each time node to form the complete control parameter group of the time node, arrange the complete control parameter groups of all time nodes in chronological order to form the integrated control parameter group sequence, each time node corresponds to a complete control parameter group.
[0141] For each node in the unified time node sequence, the joint angle target value (or supplementary angle target value), joint speed target value (or supplementary speed target value), picking execution component action state and transition parameter obtained before are summarized together to form the complete control parameter group of the time node. For example, the control parameter group of time node T11 includes the angle target values of joints 1 to 6, the speed target values of joints 1 to 6, the maintenance state of the picking execution component, and the transition parameters of each joint (if in the transition stage). Arrange the complete control parameter groups of all time nodes in chronological order to form the integrated control parameter group sequence, ensuring that each time node has unique and complete control parameters.
[0142] Step S147: Delete the repeated time nodes and repeated control parameters that appear in the integration process, keep the unique control parameter group corresponding to each time node, time axis calibration is performed on the integrated control parameter group, so that all control parameter groups are arranged according to a unified time interval to form a continuous action control parameter sequence, convert the action control parameter sequence into an instruction format sequence that can be recognized by the mechanical arm control system, and obtain the adapted action sequence.
[0143] In the integrated control parameter group sequence, it is checked whether there is a repeated time node, if there is, one of the control parameter groups is retained and the rest of the parameter groups corresponding to the repeated nodes are deleted. At the same time, the control parameter groups of different time nodes are compared, if there are repeated parameter groups with the same parameters, the redundant groups can be retained or deleted according to the action continuity requirement. Then, taking the unified time interval as the benchmark, the time axis of the control parameter group sequence is calibrated, if the time interval of adjacent parameter groups does not meet the unified standard, the intermediate parameter groups are supplemented by interpolation or the node time is adjusted to ensure consistent interval, forming a continuous and uninterrupted action control parameter sequence. Finally, according to the communication protocol and instruction format requirements of the robot arm control system, each parameter group in the action control parameter sequence is converted into the corresponding instruction code, for example, the joint angle target value is converted into a numerical code that can be recognized by the control system, and the execution component action state is converted into the corresponding state instruction code. These instruction codes are arranged in chronological order to form an adaptive action sequence.
[0144] Step S150: converting the adaptive action sequence into a picking robot arm action control instruction, and sending the picking robot arm action control instruction to a picking robot arm control system. After receiving the picking robot arm action control instruction, the picking robot arm control system drives the picking robot arm to perform the picking action.
[0145] The instruction codes in the adaptive action sequence are packaged according to the instruction frame format of the robot arm control system, each instruction frame contains time stamp, joint number, parameter type, parameter value and check code and other information, to ensure the accuracy and integrity of the instruction transmission. After packaging, the picking robot arm action control instruction is sent to the picking robot arm control system through industrial Ethernet or special communication bus. After receiving the instruction, the robot arm control system first checks the instruction frame, and after the check is passed, the control parameters of each time node are parsed, the servo motor of each joint is driven to move at the preset angle and speed according to the parameters, and the picking execution component is controlled to start, maintain and stop the action at the specified time node, finally completing the whole apple picking process.
[0146] Step S210: collecting historical action association rules and historical picking action optimization targets of the picking robot arm in historical picking scenarios, decomposing the historical action association rules into historical rule units, and decomposing the historical picking action optimization targets into historical target units to form a model training data set, the model training data set contains corresponding samples of multiple historical rule units and historical target units.
[0147] Extract historical action association rules accumulated in past apple picking processes from the rule database, which cover different historical picking scene types (such as different seasons, different orchard apple picking scenes). At the same time, collect the corresponding historical picking action optimization goals, including historical picking efficiency goals, action smoothness goals, and damage rate control goals. Split the historical action association rules into multiple historical rule units according to the same disassembly method as the current rule unit, and each historical rule unit corresponds to a specific historical association feature description; split the historical picking action optimization goals into historical goal units, and each historical goal unit corresponds to a historical optimization index. Match each historical rule unit with the corresponding historical goal unit to form a set of samples, and multiple sets of samples collectively constitute the model training dataset.
[0148] Step S220: Divide the model training dataset, and divide the samples into a training sample subset and a validation sample subset according to a preset proportion. The training sample subset is used for model parameter training, and the validation sample subset is used for model performance verification.
[0149] The preset proportion is set to 70% for the training sample subset and 30% for the validation sample subset. Random sampling is used to extract 70% of the samples from the model training dataset to form the training sample subset, and the remaining 30% of the samples form the validation sample subset. During the sampling process, ensure that the training sample subset and the validation sample subset both cover samples of different historical picking scene types to avoid sample distribution bias during model training.
[0150] Step S230: Input the training sample subset into the feature input layer of the machine learning model to perform feature encoding processing on the historical rule unit and the historical goal unit, and generate a historical rule feature vector and a historical goal feature vector.
[0151] The historical rule unit and the historical goal unit in the training sample subset are input into the feature input layer of the machine learning model. The feature input layer first preprocesses the text content of the historical rule unit and the historical goal unit, including removing meaningless auxiliary words, unifying the expression method of professional terms, and splitting long text into short sentences. After preprocessing, the bag-of-words model combined with the TF-IDF algorithm is used to extract features from the text, convert each historical rule unit into a fixed-dimensional historical rule feature vector, and convert each historical goal unit into a historical goal feature vector of the same dimension. Each element in the vector corresponds to the weight value of a feature word, representing the importance of the feature word in the text.
[0152] Step S240: Input the historical rule feature vector and the historical goal feature vector into the association modeling layer, initialize the multi-layer perceptron structure parameters of the association modeling layer, including the weight matrix and the bias vector of each layer, set the number of training iterations and the learning rate.
[0153] The generated historical rule feature vector and the historical target feature vector are spliced to form a combined feature vector, which is input into the association modeling layer. The multi-layer perception of the association modeling layer includes three hidden layers, the first layer has a specific number of neurons, and the number of neurons of the second and third layers decreases in turn. The weight matrix and bias vector of each hidden layer are initialized, the initial value of the weight matrix is generated by the Xavier initialization method, and the initial value of the bias vector is set to a zero vector. The number of training iterations is set to a specific number, and the learning rate is set to a specific value. The learning rate is used to control the amplitude of each parameter update, and the number of iterations is used to control the total number of model training rounds.
[0154] Step S250: In each iteration process, the association degree prediction value of the historical rule feature vector and the historical target feature vector is calculated by the association modeling layer, and the association degree prediction value is compared with the preset association degree true value in the sample to calculate the prediction error.
[0155] In each iteration, the combined feature vector passes through the first hidden layer of the association modeling layer, performs matrix multiplication operation with the weight matrix of the layer, and then adds the bias vector to obtain the linear transformation result, and then performs nonlinear transformation through the ReLU activation function to generate the first hidden layer output feature. The first hidden layer output feature is input into the second hidden layer, and the above linear transformation and ReLU activation function processing process is repeated to obtain the second hidden layer output feature. The second hidden layer output feature is input into the third hidden layer, and after linear transformation and Sigmoid activation function processing, the association degree prediction value is obtained. The range of the association degree prediction value is between 0 and 1, representing the predicted association degree of the historical rule unit and the historical target unit. The association degree prediction value is compared with the manually labeled association degree true value in the sample, and the prediction error between the two is calculated using the mean square error loss function. The calculation method of the mean square error loss function is the average value of the square of the difference between the predicted value and the true value.
[0156] Step S260: According to the prediction error, the weight matrix and bias vector of the association modeling layer are adjusted using the gradient descent algorithm, and the encoding parameters of the feature input layer are updated at the same time, so as to reduce the overall prediction error of the training sample subset.
[0157] According to the calculated prediction error, the error signal is back-propagated using the stochastic gradient descent algorithm. First, the partial derivative of the loss function with respect to the output of the association modeling layer is calculated, and then the partial derivatives of the weight matrices and bias vectors of the third hidden layer, the second hidden layer, and the first hidden layer with respect to the loss function are calculated in turn according to the chain rule. According to the direction and size of the partial derivative, the weight matrices and bias vectors of each layer are adjusted by combining the preset learning rate, so that the prediction error changes in the direction of reduction. At the same time, the error signal is used to adjust the feature word weight calculation parameters of the TF-IDF algorithm in the feature input layer, optimize the feature encoding effect, and further reduce the overall prediction error of the training sample subset.
[0158] Step S270: After a preset number of iterations are completed, the verification sample subset is input into the trained machine learning model, the correlation degree prediction error of the verification sample subset is calculated, and if the prediction error meets the preset requirement, it is determined that the model training is completed.
[0159] When the model completes a preset number of iterations, the historical rule units and historical target units in the verification sample subset are input into the trained machine learning model, and the historical rule feature vectors and historical target feature vectors are generated according to the same processing flow as the training samples, and the correlation degree prediction value is calculated through the association modeling layer. The average prediction error of the verification sample subset is calculated, and if the error is less than the preset error threshold (i.e., meets the preset requirement), it indicates that the model has good generalization ability and can accurately predict the correlation degree of different samples, and it is determined that the model training is completed.
[0160] Step S280: If the prediction error of the verification sample subset does not meet the preset requirement, the training iteration number and the learning rate are adjusted, and the model training is performed again until the prediction error of the verification sample subset meets the preset requirement. The parameters of each layer of the trained machine learning model are saved to form a callable pre-trained machine learning model.
[0161] If the average prediction error of the verification sample subset is greater than or equal to the preset error threshold, it indicates that the model may have underfitting or overfitting problems. At this time, the training iteration number is adjusted, and the iteration number is appropriately increased to solve the underfitting problem; at the same time, the learning rate is adjusted, if the learning rate is too high before, the parameter oscillation is caused, then the learning rate is appropriately reduced, if the learning rate is too low, the convergence is slow, then the learning rate is appropriately increased. After the adjustment is completed, the model training is performed again using the adjusted parameters, and the process of steps S230 to S270 is repeated until the prediction error of the verification sample subset meets the preset requirement. After the model training is completed, all model parameters such as the encoding parameters of the feature input layer, the weight matrices and bias vectors of each hidden layer of the association modeling layer are saved to the model file to form a callable pre-trained machine learning model.
[0162] Figure 2A schematic diagram showing exemplary hardware and software components of the picking robot action optimization method system 100 provided by some embodiments of the present application that can implement the idea of the present application is shown. For example, a processor 120 can be used in the picking robot action optimization method system 100 in combination with machine learning and for performing the functions in the present application.
[0163] The picking robot action optimization method system 100 in combination with machine learning can be a general server or a special-purpose server, both of which can be used to implement the picking robot action optimization method in combination with machine learning of the present application. Although only one server is shown in the present application, for the sake of convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0164] For example, the picking robot action optimization method system 100 in combination with machine learning can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, a ROM, or a RAM, or any combination thereof. Exemplarily, the picking robot action optimization method system 100 in combination with machine learning can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The methods of the present application can be implemented according to these program instructions. The picking robot action optimization method system 100 in combination with machine learning also includes an I / O interface 150 between the computer and other input / output devices.
[0165] For the sake of illustration, only one processor is described in the picking robot action optimization method system 100 in combination with machine learning. However, it should be noted that the picking robot action optimization method system 100 in combination with machine learning in the present application can also include multiple processors, so the steps performed by one processor described in the present application can also be jointly performed or individually performed by multiple processors. For example, if the processor of the picking robot action optimization method system 100 in combination with machine learning performs steps A and B, it should be understood that steps A and B can also be jointly performed by two different processors or individually performed in one processor. For example, a first processor performs step A, a second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0166] In addition, the present application also provides a readable storage medium, in which computer executable instructions are pre-set, and when a processor executes the computer executable instructions, the picking robot action optimization method in combination with machine learning is implemented.
[0167] It should be noted that the foregoing description of embodiments of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form described, and many modifications, variations, and alternatives are possible.
Claims
1. A method for optimizing the movements of a harvesting robotic arm by incorporating machine learning, characterized in that, The method includes: Obtain motion data sets of the picking robotic arm under different picking scenarios, perform correlation analysis on the motion data sets, and generate motion correlation rules. The motion data sets include motion angle data, motion speed data, and motion duration data of each joint of the robotic arm and the picking execution component. Call a pre-trained machine learning model, input the action association rules and the preset picking action optimization target into the machine learning model, establish the mapping relationship between the action association rules and the picking action optimization target, and obtain a mapping relationship table; Based on the mapping table, an initial motion optimization scheme for the picking robot arm is generated. The initial motion optimization scheme includes adjustment angle parameters, adjustment speed parameters, and adjustment duration parameters for each joint and the picking execution component. The initial motion optimization scheme is subjected to motion sequence adaptation processing, and the parameters in the initial motion optimization scheme are converted into a motion sequence that can be executed by the robotic arm to obtain the adapted motion sequence. The adapted action sequence is converted into action control instructions for the harvesting robotic arm, and the action control instructions for the harvesting robotic arm are sent to the harvesting robotic arm control system. After receiving the action control instructions for the harvesting robotic arm, the harvesting robotic arm control system drives the harvesting robotic arm to perform harvesting actions.
2. The method for optimizing the actions of a harvesting robotic arm by combining machine learning according to claim 1, characterized in that, The process of acquiring motion data sets of the harvesting robotic arm under different harvesting scenarios, performing correlation analysis on the motion data sets, and generating motion correlation rules includes: The system receives a set of motion data transmitted by the picking robotic arm through joint sensors and motion timers. The joint sensors are used to collect motion angle data and motion speed data of each joint of the robotic arm, and the motion timers are used to collect motion duration data of the picking execution components. The action data set is divided according to the picking scene type. Each picking scene type corresponds to a set of scene action data subsets. The picking scene type is determined according to the morphological characteristics and growth environment characteristics of the picking object. Each subset of scene action data is segmented into multiple action data blocks. Each action data block contains motion angle data, motion speed data, and action duration data of each joint of the robotic arm within a single complete picking action cycle. Extract the associated features from each of the motion data blocks. The associated features include the correspondence between motion angle data and motion speed data of the same joint, the coordination relationship between motion angle data of different joints, and the correspondence between motion duration data of the picking and executing component and joint motion speed data. Calculate the similarity of associated features between different action data blocks in the same subset of action data for the same scene. For each associated feature, calculate the similarity of its data sequence in different data blocks separately to obtain the independent similarity score of each feature. Based on the importance weight of each associated feature in the picking action, combine all independent similarity scores to obtain the overall similarity of associated features. Based on the correlation feature similarity, action data block groups with similar correlation features are selected, and each action data block group contains multiple action data blocks whose correlation feature similarity meets preset conditions. Analyze the commonalities of the associated features in each action data block group, and extract rule entries that can reflect the association patterns of the action data blocks in the action data block group. Each rule entry contains a description of the specific correspondence between the associated features. All the rule entries are categorized and organized according to the picking scenario type. Each picking scenario type corresponds to a set of rule entries, forming a complete action association rule.
3. The method for optimizing the actions of a harvesting robotic arm by combining machine learning according to claim 2, characterized in that, The extraction of associated features from each action data block includes: From each of the motion data blocks, the motion angle data, motion speed data, and motion duration data of each joint of the robotic arm and the picking execution component are separated, and each data type corresponds to an independent data sequence; For the motion angle data and motion velocity data of the same joint, a one-to-one correspondence is established according to the time node sequence. The motion angle data value and motion velocity data value of each time node form a data pair. Multiple data pairs form the angle-velocity correspondence sequence of the joint. The angle-velocity correspondence sequence is the correspondence between the motion angle data and motion velocity data of the same joint. Select joint groups in the robotic arm that have a motion coordination relationship, wherein the joint group contains two or more joints that move simultaneously during the picking action, and extract the motion angle data sequence of each joint in the joint group; The temporal synchronicity of the joint motion angle data sequences in the joint group is calculated, and the synchronicity level is determined by comparing the time difference of the peak occurrence in different joint motion angle data sequences. The synchronicity level is related to the time difference. Based on the synchronization level and the proportion of changes in the motion angle data of each joint, a description of the collaborative relationship between the motion angle data of different joints in the joint group is established. The collaborative relationship description includes the synchronization level value and the proportion of the angle change. Extract the action duration data of the picking execution component from each action data block, determine the start and end time points of the action duration, extract the motion speed data sequence of all joints during the period from the start time point to the end time point, and calculate the average value and change range of the motion speed data of each joint in the corresponding time period. Establish a correspondence between the action duration data of the picking execution component and the average value and variation range of the joint movement speed data. Each action duration data value corresponds to a set of average joint movement speed and variation range values, forming a duration-speed correspondence. Integrate the angle-speed correspondence of the same joint, the coordination relationship of movement angle data between different joints, and the duration-speed correspondence of the picking execution component to form the associated features of each action data block.
4. The method for optimizing the actions of a harvesting robotic arm by combining machine learning according to claim 2, characterized in that, The rule entries are categorized and organized according to the picking scenario type, with each picking scenario type corresponding to a set of rule entries, forming a complete action association rule, including: Establish a classification directory of picking scene types. The classification directory contains all the identified picking scene type names and corresponding scene feature descriptions. The scene feature descriptions include specific content such as the morphological characteristics of the picking objects and the characteristics of the growing environment. Traverse each rule entry and extract the scene information corresponding to the associated features involved in the rule entry. The scene information includes the picking scene type to which the action data block from which the rule entry originates belongs. Based on the extracted scene information, the rule entries are assigned to the corresponding picking scene type in the category directory to form the initial rule group for each scene type; For each scenario type, an initial rule group is used to detect rule conflicts. The association feature descriptions and correspondences between different rule entries in the initial rule group are compared. If there are rule entries with the same association feature descriptions but different correspondences, they are marked as conflicting rule pairs. For each conflicting rule pair, query the correlation feature similarity statistics of the action data block group from which it originates, retain the rule entry with the largest size of the source data block group and the largest average correlation feature similarity, and delete the other conflicting rule entry. Supplement the rule entries corresponding to the missing related features in the initial rule group for each scene type. If any related feature has no corresponding rule entry under the scene type, select the rule entry with the highest correlation from the rule group of similar picking scene types and adapt and modify it to form a supplementary rule entry. Logically sort the rule groups for each scenario type. According to the execution order of the actions involved in the associated features, the rule entries are divided into joint motion association rules, execution component action association rules, and collaborative action association rules. After sorting according to the complexity of the associated features, a scenario identifier is added to the sorted rule groups. The scenario identifier includes the picking scenario type name and the rule group version number. Summarize the rule groups for all picking scenario types to generate a rule index table, which includes the scenario type name, rule group version number, and number of rule entries; The summarized rule groups and rule index tables are integrated to form complete action association rules, which are then stored in the rule database.
5. The method for optimizing the actions of a harvesting robotic arm by combining machine learning according to claim 1, characterized in that, The process involves calling a pre-trained machine learning model, inputting the action association rules and preset picking action optimization targets into the machine learning model, establishing a mapping relationship between the action association rules and the picking action optimization targets, and obtaining a mapping relationship table, including: Determine the preset optimization goals for the picking action. The optimization goals for the picking action include the efficiency goal of completing the picking action, the stability goal of the picking action, and the damage rate control goal of the picking object. Each optimization goal has a corresponding descriptive index. The action association rule is decomposed into multiple rule units, each rule unit corresponding to a specific association feature description. At the same time, the picking action optimization target is decomposed into multiple target units, each target unit corresponding to a description index of the optimization target. The rule unit and the target unit are input into a pre-trained machine learning model, which includes a feature input layer, an association modeling layer, and a mapping output layer. In the feature input layer, feature encoding processing is performed on the rule units and the target units to convert the text-based rule units and target units into vector-based feature data. Each rule unit corresponds to a rule feature vector, and each target unit corresponds to a target feature vector. The regular feature vector and the target feature vector are input into the association modeling layer. The association modeling layer adopts a multilayer perceptron structure and calculates the correlation degree between the regular feature vector and the target feature vector through linear transformation and nonlinear activation function. The matching relationship between the rule unit and the target unit is determined based on the correlation value. The rule unit and the target unit that meet the preset conditions form a matching pair. Each matching pair contains a rule unit and a corresponding target unit. All the matching pairs are grouped according to the picking scene type, and the matching pairs under each picking scene type form a local mapping relationship for that picking scene type; The local mapping relationships of each picking scene type are integrated, and the association and connection relationships of matching pairs between different scenes are supplemented to form a global mapping relationship covering all picking scene types. The global mapping relationship is presented in tabular form to obtain the mapping relationship table. The row dimension of the table is the rule unit, and the column dimension is the target unit. The corresponding correlation value and matching condition are recorded in the table cell.
6. The method for optimizing the actions of a harvesting robotic arm by combining machine learning according to claim 5, characterized in that, The step of inputting the rule unit and the target unit into a pre-trained machine learning model includes: Collect historical action association rules and historical action optimization targets of the harvesting robot arm in historical harvesting scenarios. Decompose the historical action association rules into historical rule units and the historical action optimization targets into historical target units to form a model training dataset. The model training dataset contains multiple sets of corresponding samples of historical rule units and historical target units. The model training dataset is divided into a training sample subset and a validation sample subset according to a preset ratio. The training sample subset is used for model parameter training, and the validation sample subset is used for model performance verification. The training sample subset is input into the feature input layer of the machine learning model to perform feature encoding on the historical rule unit and the historical target unit, thereby generating historical rule feature vector and historical target feature vector. Input the historical rule feature vector and the historical target feature vector into the association modeling layer, initialize the multilayer perceptron structure parameters of the association modeling layer, including the weight matrix and bias vector of each layer, and set the number of training iterations and the learning rate. In each iteration, the correlation prediction value between the historical rule feature vector and the historical target feature vector is calculated through the correlation modeling layer. The correlation prediction value is then compared with the preset true correlation value in the sample to calculate the prediction error. Based on the prediction error, the gradient descent algorithm is used to adjust the weight matrix and bias vector of the association modeling layer, and the encoding parameters of the feature input layer are updated to reduce the overall prediction error of the training sample subset. After completing a preset number of iterations, the validation sample subset is input into the trained machine learning model, and the correlation prediction error of the validation sample subset is calculated. If the prediction error meets the preset requirements, the model training is determined to be complete. If the prediction error of the validation sample subset does not meet the preset requirements, adjust the number of training iterations and the learning rate, and retrain the model until the prediction error of the validation sample subset meets the preset requirements. Then, save the parameters of each layer of the trained machine learning model to form a pre-trained machine learning model that can be called.
7. The method for optimizing the actions of a harvesting robotic arm by combining machine learning according to claim 1, characterized in that, The initial motion optimization scheme for the harvesting robotic arm, generated based on the mapping table, includes: Extract the rule units and matching target units corresponding to the current picking scenario type from the mapping relationship table, and determine the associated features that need to be optimized and the corresponding optimization target description indicators under the current picking scenario type; Based on the association feature description in the rule unit, locate the original motion parameters corresponding to the association feature in the motion data set. The original motion parameters include the original motion angle data, original motion speed data, and original motion duration data of each joint of the robotic arm and the picking execution component. Referring to the optimized target description index in the target unit, the original action parameters are input into a predefined performance evaluation model. The performance evaluation model outputs a predicted performance index value based on the current action parameters, and the difference between the predicted performance index value and the optimized target description index is calculated. Based on the differences and the correlation values in the mapping table, the adjustment direction of each original action parameter is determined. The adjustment direction is determined according to the positive or negative attribute of the differences and the magnitude of the correlation values. Based on the adjustment direction and the preset adjustment step length rule, the adjustment angle parameter, adjustment speed parameter and adjustment duration parameter of each joint are calculated. The adjustment angle parameter is the result of superimposing the original motion angle data and the adjustment amount, the adjustment speed parameter is the result of superimposing the original motion speed data and the adjustment amount, and the adjustment duration parameter is the result of superimposing the original action duration data and the adjustment amount. The calculated adjustment angle parameters, adjustment speed parameters, and adjustment duration parameters of the picking execution components are arranged according to the action execution sequence, with each action execution stage corresponding to a set of parameter combinations; Supplement the transition parameters between the various parameter combinations, which include the transition time and speed change gradient of joint adjustment between adjacent action phases; By integrating parameter combinations with transition parameters, an initial motion optimization scheme containing a complete motion optimization parameter system is formed.
8. The method for optimizing the actions of a harvesting robotic arm by incorporating machine learning according to claim 7, characterized in that, Based on the adjustment direction and the preset adjustment step length rule, the adjustment angle parameters, adjustment speed parameters, and adjustment duration parameters of the picking execution component for each joint are calculated. The adjustment angle parameters are the superposition result of the original motion angle data and the adjustment amount, including: Obtain a preset adjustment step size rule, wherein the adjustment step size rule includes adjustment step size coefficients corresponding to different correlation values, and the correlation values and adjustment step size coefficients are in a corresponding relationship; Based on the correlation values between the rule units and the target units in the mapping table, the corresponding adjustment step size coefficient is selected from the adjustment step size rules; For the original motion angle data of each joint of the robotic arm, the positive or negative attribute of the adjustment amount is determined according to the adjustment direction. When the adjustment direction is to increase the angle, the adjustment amount is positive, and when the adjustment direction is to decrease the angle, the adjustment amount is negative. Calculate the adjustment amount of the joint adjustment angle. The adjustment amount is equal to the original motion angle data multiplied by the adjustment step length coefficient. The adjustment angle parameter is equal to the original motion angle data plus the adjustment amount. For the original motion speed data of each joint of the robotic arm, the positive and negative attributes of the adjustment amount are determined according to the adjustment direction. When the adjustment direction is to increase the speed, the adjustment amount is positive, and when the adjustment direction is to decrease the speed, the adjustment amount is negative. Calculate the adjustment amount of the joint adjustment speed. The adjustment amount is equal to the original motion speed data multiplied by the adjustment step length coefficient. The adjustment speed parameter is equal to the original motion speed data plus the adjustment amount. For the original action duration data of the picking execution component, the positive or negative attribute of the adjustment amount is determined according to the adjustment direction. When the adjustment direction is to extend the duration, the adjustment amount is positive, and when the adjustment direction is to shorten the duration, the adjustment amount is negative. Calculate the adjustment amount for the adjustment time of the picking execution component. The adjustment amount is equal to the original action time data multiplied by the adjustment step coefficient. The adjustment time parameter is equal to the original action time data plus the adjustment amount. Record the calculation process of adjustment angle parameters, adjustment speed parameters, and adjustment duration parameters of each joint, including original data values, adjustment step coefficients, adjustment amounts, and final parameter values.
9. The method for optimizing the actions of a harvesting robotic arm by combining machine learning according to claim 1, characterized in that, The step of performing motion sequence adaptation processing on the initial motion optimization scheme, converting the parameters in the initial motion optimization scheme into a motion sequence executable by the robotic arm, to obtain an adapted motion sequence includes: Extract the adjustment angle parameters, adjustment speed parameters, adjustment duration parameters of the picking execution component, and transition connection parameters of each joint in the initial motion optimization scheme; The adjustment angle parameters of each joint are converted into angle control sequences for the movement of the robotic arm joints. Each angle control sequence contains target joint angle values corresponding to multiple time nodes, and the interval between time nodes is determined according to the adjustment speed parameters. The adjustment speed parameters of each joint are converted into speed control sequences for the movement of the robotic arm joints. Each speed control sequence contains target joint speed values corresponding to multiple time nodes. The speed control sequence and the angle control sequence are kept synchronized at the time nodes. The adjustment duration parameter of the picking execution component is converted into an action control sequence of the picking execution component. The action control sequence includes the start time node, action maintenance time node and stop time node of the picking execution component. The action maintenance time node corresponds to the adjustment duration parameter. The transition parameters are converted into a transition control sequence, which includes joint angle transition target values, speed transition target values and transition time nodes between adjacent motion phases. The transition control sequence is used to connect adjacent angle control sequences and speed control sequences. The angle control sequence, speed control sequence, action control sequence of the picking execution component, and transition control sequence are integrated according to the time node sequence, with each time node corresponding to a complete set of control parameters; Duplicate time nodes and duplicate control parameters that occur during the integration process are deleted, and the unique control parameter group corresponding to each time node is retained. The integrated control parameter group is time-axis calibrated so that all control parameter groups are arranged according to a uniform time interval to form a continuous motion control parameter sequence. The motion control parameter sequence is then converted into an instruction format sequence that the robotic arm control system can recognize to obtain the adapted motion sequence.
10. A method system for optimizing the movements of a harvesting robotic arm by combining machine learning, characterized in that, The machine learning-integrated robotic arm motion optimization method system includes a processor and a memory, the memory and the processor being connected, the memory being used to store programs, instructions or code, and the processor being used to execute the programs, instructions or code in the memory to implement the machine learning-integrated robotic arm motion optimization method as described in any one of claims 1-9.
Citation Information
Patent Citations
String type fruit distributed visual active sensing method and application thereof
CN111602517A
Motion learning method and device, medium and electronic equipment
CN112580582A
Autonomous operation decision-making method for picking manipulator
CN117621046A
Picking mechanical arm motion planning method and system
CN119188719A
Clamping control method and device, computer equipment and readable storage medium
CN119973992A