Battery identification model training method, disorderly grasping method, device and storage medium

By building simulation training samples and CrossTransformer models, the calculation amount is reduced and the training accuracy is improved, the problem of insufficient data on the site recycling of lead-acid batteries is solved, and more efficient battery identification and disorderly capture is achieved.

CN116452913BActive Publication Date: 2025-07-04HUNAN RUIYI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310321215.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-07-04
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

In the absence of training data on the lead-acid battery recycling site, when using the Transformer model for battery recognition, more training samples and training time are required, resulting in poor recognition results.

Method used

By constructing a battery model to simulate stacked pictures of different poses, a simulation training sample is generated, and the CrossTransformer model is used to calculate the attention coefficient between the pixel points and other pixel points in the row and column. The battery recognition model is trained in combination with simulation and real training samples, reducing the calculation amount and improving training accuracy.

Benefits of technology

The recognition effect of the battery, battery gloss surface, battery electrode surface and belt lifting is improved, and more accurate disorderly capture is achieved, solving the problem of poor model recognition effect caused by the lack of on-site data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452913B_ABST
    Figure CN116452913B_ABST
Patent Text Reader

Abstract

The present invention discloses a battery recognition model training method, a disordered grasping method, a device and a storage medium. The training method includes simulating different battery stacking simulation diagrams; forming simulation training samples from pixel matrices obtained by processing the battery stacking simulation diagrams; constructing a battery recognition model, which includes N CrossTransformer models. The CrossTransformer model only calculates the attention coefficients between each pixel point and other pixel points in the row and column where the pixel point is located, and replaces each pixel point with the attention coefficient of the pixel point. The pixel matrix composed of the attention coefficients is used as the input of the next CrossTransformer model; training the battery recognition model using the simulation training samples; and retraining the battery recognition model using real training samples to obtain a target battery recognition model. The present invention has better recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of visual recognition and unordered grasping, and particularly relates to a battery recognition model training method, an unordered grasping method, a device and a storage medium. Background Art

[0002] The recycling production line of lead-acid batteries includes three links: truck unloading, acid discharging, and strap recycling. After the waste batteries enter the factory, they need to be transferred from the truck to the conveyor belt, which is the truck unloading link; after the batteries are regularly placed on the conveyor belt, the straps used for carrying handles at both ends of the batteries need to be cut off for recycling, which is the strap recycling link; subsequently, before entering the sawing machine, it is necessary to ensure that the side of the battery with electrodes faces up to ensure that the blade on the conveyor belt can smoothly saw open the bottom of the battery and discharge the acid, which is the acid discharging link.

[0003] The waste batteries weigh more than 100 catties, and the production environment of the factory is relatively harsh, with a strong demand for unmanned operation. To automate the entire production line, the key to truck unloading, acid discharging, and strap recycling lies in achieving visual recognition of the batteries. The visual recognition task is essentially a four-classification task, that is, to classify the battery, the smooth surface of the battery, the electrode surface of the battery, and the strap of the battery. As long as the classification of the battery, the smooth surface of the battery, the electrode surface of the battery, and the strap of the battery is achieved, the truck unloading, acid discharging, and strap recycling can be completed by the robotic arm.

[0004] Although the visual recognition technology has been increasingly perfected, there are still many feasibility problems at the application level. For example, the training of a battery recognition model (which can recognize the battery, the smooth surface of the battery, the electrode surface of the battery, and the strap of the battery) relies on supervised learning, and supervised learning requires a large number of training samples. The application of visual recognition technology in the field of lead-acid battery recycling is a new technology, and currently no factory uses visual recognition technology to achieve lead-acid battery recycling. Therefore, the on-site data is very scarce, resulting in a serious shortage of training samples. At the same time, the on-site data is essentially pictures. The batteries are stacked disorderly and in large quantities, resulting in a relatively messy pixel distribution in the pictures, which poses very high requirements for the battery recognition model to extract local detail features.

[0005] Currently, the mainstream image processing models are all built based on convolutional neural networks (CNNs). However, the feature extraction of CNNs is affected by the downsampling of convolutional kernels, resulting in a large amount of detail loss, which is not suitable for battery identification and recycling. The Transformer model has advantages in grasping global information and details during feature extraction. However, the Transformer model requires more training samples because its structure is more complex than that of convolutional neural networks. In a convolutional neural network, the model is mainly composed of convolutional layers and pooling layers, and these layers can share weights and parameters to a certain extent. Therefore, a convolutional neural network can effectively learn features with a small number of training samples. In contrast, the Transformer model uses self-attention mechanisms to learn the dependencies between sequences, which means that the Transformer model needs to consider all interactions between each position in the sequence and has no shared weights or parameters. This leads to the Transformer model requiring more training samples for training compared to convolutional neural networks. In addition, the Transformer model also requires a longer training time because during the training process, the Transformer model needs to gradually understand the structure in the sequence and gradually capture more subtle relationships, which also requires more training samples to achieve better performance.

[0006] Therefore, in the case of a serious lack of on-site data in the lead-acid battery recycling production line, using the Transformer model, which requires more training samples and training time, for battery identification cannot achieve good identification results. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for training a battery identification model, a method for disorderly grasping, a device, and a storage medium to solve the problem that in the case of a serious lack of on-site data in lead-acid battery recycling, using the Transformer model, which requires more training samples and training time, for battery identification cannot achieve good identification results.

[0008] The present invention solves the above technical problems through the following technical solutions: A method for training a battery identification model, the method comprising the following steps:

[0009] Construct a battery model, and simulate stacking pictures of batteries in different postures according to the battery model to obtain battery stacking simulation pictures;

[0010] Annotate each of the battery stacking simulation pictures, and the annotation content includes batteries, battery smooth surfaces, battery electrode surfaces, and lifting straps;

[0011] Preprocess each annotated battery stacking simulation picture to obtain a corresponding pixel matrix, and form a simulation training sample from all the pixel matrices;

[0012] Construct a battery recognition model, where the battery recognition model includes N CrossTransformer models. When calculating the attention coefficients, each CrossTransformer model only calculates the attention coefficients between each pixel and other pixels in the same row and the same column where the pixel is located; replace each pixel in the pixel matrix with its attention coefficient, and the pixel matrix composed of the attention coefficients is used as the input of the next CrossTransformer model;

[0013] Use the simulation training samples to train the battery recognition model to obtain a trained battery recognition model;

[0014] Obtain a real image of the battery stacking at the battery recycling site, and label the real image of the battery stacking. The labeling content includes batteries, battery smooth surfaces, battery electrode surfaces, and carrying straps;

[0015] Preprocess each labeled real image of the battery stacking to obtain a corresponding pixel matrix, and the real training samples are composed of all pixel matrices;

[0016] Use the real training samples to retrain the trained battery recognition model to obtain a target battery recognition model.

[0017] Further, the specific implementation process of preprocessing the battery stacking simulation image or the real image of the battery stacking includes:

[0018] Convert each image into a pixel matrix;

[0019] Perform normalization processing on each pixel matrix to obtain a normalized pixel matrix;

[0020] Perform data augmentation on each normalized pixel matrix to obtain a pixel matrix after data augmentation processing.

[0021] Further, the data augmentation includes mirroring, rotation, scaling, cropping, translation, and Gaussian noise.

[0022] Further, the specific implementation process of training the battery recognition model using the simulation training samples or the real training samples includes:

[0023] In the first CrossTransformer model, feature extraction is performed on each pixel point in the pixel matrix to obtain the feature vector of each pixel point in the pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; according to the query matrix Q, key matrix K, and value matrix V of each pixel point, the attention coefficient between this pixel point and other pixel points in the row and column where this pixel point is located is calculated; each pixel point is replaced with the corresponding attention coefficient, and the first pixel matrix is formed by the attention coefficients.

[0024] In the second CrossTransformer model, feature extraction is performed on each pixel point in the first pixel matrix to obtain the feature vector of each pixel point in the first pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; according to the query matrix Q, key matrix K, and value matrix V of each pixel point, the attention coefficient between this pixel point and other pixel points in the row and column where this pixel point is located is calculated; each pixel point is replaced with the corresponding attention coefficient, and the second pixel matrix is formed by the attention coefficients.

[0025] By analogy, in the Nth CrossTransformer model, feature extraction is performed on each pixel point in the (N - 1)th pixel matrix to obtain the feature vector of each pixel point in the (N - 1)th pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; according to the query matrix Q, key matrix K, and value matrix V of each pixel point, the attention coefficient between this pixel point and other pixel points in the row and column where this pixel point is located is calculated; each pixel point is replaced with the corresponding attention coefficient, and the Nth pixel matrix is formed by the attention coefficients.

[0026] The Nth pixel matrix is passed through the YOLO detection head for target prediction to obtain a target prediction vector, where the target prediction vector includes the coordinate values of the upper left and lower right corners of the detection box, the confidence score, the label index, and the label score. Among them, the confidence score represents the probability score that the detection box contains a target, the label index represents the category of the target, and the label score represents the probability score that the target belongs to the category.

[0027] Furthermore, the calculation formula for the attention coefficient is:

[0028]

[0029] Among them, A represents the attention coefficient, d k represents the number of feature channels of the query matrix Q, key matrix K, and value matrix V.

[0030] Further, each of the CrossTransformer models includes a dropout layer, and when retraining the trained battery recognition model using the true training samples, all dropout layers are turned off.

[0031] Based on the same concept, the present invention also provides a disordered grasping method, and the method includes the following steps:

[0032] Obtain the battery stacking pictures at the battery recycling site in real time;

[0033] Preprocess the battery stacking pictures to obtain a pixel matrix;

[0034] Call the target battery recognition model, and the target battery recognition model is trained according to the battery recognition model training method as described above;

[0035] Based on the target battery recognition model, identify the pixel matrix to obtain a target prediction vector;

[0036] Calculate the battery key point coordinates according to the target prediction vector, convert the key point coordinates into world coordinates, and send the world coordinates to the grasping device;

[0037] Generate a motion path according to the world coordinates of the key points and the current position of the grasping device, and control the grasping device to move and grasp according to the motion path.

[0038] Based on the same concept, the present invention provides an electronic device, including:

[0039] A memory for storing a computer program;

[0040] A processor for implementing the battery recognition model training method or the disordered grasping method as described above when executing the computer program.

[0041] Based on the same concept, the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the battery recognition model training method or the disordered grasping method as described above when executed by a processor.

[0042] Beneficial effects

[0043] Compared with the prior art, the advantages of the present invention are:

[0044] Compared with the Transformer model that needs to calculate the attention coefficients between each pixel and other pixels (all pixels except this pixel), the CrossTransformer model of the present invention only calculates the attention coefficients between a pixel and other pixels in the same row and the same column as this pixel, greatly reducing the computational complexity. Then, through recursion of multiple CrossTransformer models, each pixel is fused with pixels not in the same row and the same column as this pixel through other pixels in the same row and the same column as this pixel, which not only retains the influence of pixels not in the same row and the same column as this pixel, but also reduces the weight of pixels not in the same row and the same column as this pixel, improving the training accuracy of the model, and further improving the recognition effect of the battery, the light surface of the battery, the electrode surface of the battery, and the lifting belt, which is conducive to more accurate unordered grasping;

[0045] The present invention obtains a large number of battery stacking simulation diagrams through modeling and simulation, uses the battery stacking simulation diagrams to initially train the battery recognition model, and then uses the real battery stacking diagrams to retrain the battery recognition model, solving the problem that the model recognition effect is poor due to the lack of on-site data, further improving the training accuracy of the model and the recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only one embodiment of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is the flowchart of the battery recognition model training method in the embodiment of the present invention;

[0048] Figure 2 is the schematic diagram of the attention coefficient calculation by the Transformer model in the embodiment of the present invention;

[0049] Figure 3 is the schematic diagram of the attention coefficient calculation by the CrossTransformer model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] The technical solution of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0052] As Figure 1 shown, a method for training a battery identification model provided by an embodiment of the present invention includes the following steps:

[0053] Step 1: Construct a battery model, and simulate stacking pictures of batteries in different postures according to the battery model to obtain battery stacking simulation pictures.

[0054] In this embodiment, a 3Dmax is used to construct a battery model, and pictures of batteries stacked in different postures are simulated according to the battery model, with a total of 60,000 copies.

[0055] Step 2: Annotate each battery stacking simulation picture, and the annotation content includes batteries, battery smooth surfaces, battery electrode surfaces, and carrying straps.

[0056] The battery smooth surface refers to the non-electrode surface, and each smooth surface of the battery needs to be marked.

[0057] Step 3: Preprocess each annotated battery stacking simulation picture to obtain a corresponding pixel matrix, and all pixel matrices form a simulation training sample.

[0058] In this embodiment, the specific implementation process of preprocessing the battery stacking simulation picture includes:

[0059] Step 3.1: Use Opencv to read the battery stacking simulation picture and convert it into a pixel matrix.

[0060] The pixel matrix refers to a matrix composed of pixel points. For example, a 1082*1082 pixel matrix includes 1082*1082 pixel points, and the pixel matrix retains the position information of the pixel points in the picture.

[0061] Step 3.2: Perform normalization processing on each pixel matrix to obtain a normalized pixel matrix.

[0062] After normalization processing, all normalized pixel matrices have the same pixel size.

[0063] Step 3.3: Perform data augmentation on each normalized pixel matrix to obtain a pixel matrix after data augmentation processing.

[0064] In this embodiment, data augmentation includes mirroring, rotation, scaling, cropping, translation, and Gaussian noise. After a normalized pixel matrix undergoes 6 data augmentation processes, 7 normalized pixel matrices (including the original normalized pixel matrix) are obtained, greatly increasing the number of training samples.

[0065] Step 4: Build a battery recognition model.

[0066] The traditional Transformer model includes an encoding layer, a feature extraction layer, and an attention coefficient calculation layer. The encoding layer extracts features for each pixel point in the pixel matrix to obtain the feature vector of each pixel point. The feature extraction layer converts the feature vector of each pixel point into a query matrix Q, a key matrix K, and a value matrix V. The attention coefficient calculation layer: Based on the query matrix Q, the key matrix K, and the value matrix V of each pixel point, calculate the relationship between each pixel point and all other pixel points (such as Figure 2 a certain pixel point and all other pixel points as shown), that is, the attention coefficient. The specific calculation formula is:

[0067] (1)

[0068] where, A represents the attention coefficient, d k represents the number of feature channels of the query matrix Q, the key matrix K, and the value matrix V.

[0069] The traditional Transformer model extracts the correlation information between each pixel point and all other pixel points.

[0070] The battery recognition model includes N CrossTransformer models. When calculating the attention coefficient, each CrossTransformer model only calculates the attention coefficient between each pixel point and the other pixel points in the row and column where the pixel point is located, and no longer fuses and calculates with the query matrix Q and the key matrix K of every other pixel point. As Figure 3 shown, assuming the pixel matrix is 4*4, taking pixel point 01 as an example, only calculate the correlation between the pixel point numbered 01 and the pixel points in the row where pixel point 01 is located (i.e., numbered 02, numbered 03, numbered 04) and the pixel points in the column where it is located (i.e., numbered 05, numbered 06, numbered 07). In this way, the computational complexity of a 1082*1082 pixel matrix changes from 1170724 to 2164 (1082 + 1082), and the computational complexity is reduced by 541 times, greatly reducing the computational complexity.

[0071] In a local area, the correlation between a pixel point (e.g., numbered 01) and other pixel points in the same row and the same column as this pixel point (i.e., pixel points on the cross, e.g., numbered 02 - 07) is more important than the correlation between this pixel point (e.g., numbered 01) and pixel points not in the same row and the same column as this pixel point (i.e., pixel points not on the cross, e.g., numbered 08 - 12). This rule gradually weakens as the local area continues to expand. For example, the influence of pixel points at the four corners of a picture on the central pixel point is relatively small (or the correlation is relatively small). Although the correlation is small, it still exists.

[0072] To retain the relatively small correlation, the attention coefficient of each pixel point obtained by the CrossTransformer model is used to replace this pixel point, resulting in a pixel matrix composed of the attention coefficients of each pixel point. This pixel matrix is used as the input for the next CrossTransformer model. When calculating the attention coefficients of multiple CrossTransformer models, these pixel points have incorporated the correlation of their own row and column in the previous CrossTransformer model. In the next CrossTransformer model, this pixel point is correlated through other pixel points in its row and column and pixel points not in its row and column. Exemplarily, in the previous CrossTransformer model, pixel point 01 incorporated the correlation between pixel point 01 and pixel points 02 - 07. In the next CrossTransformer model, pixel point 01 incorporates the correlation between pixel point 01 and pixel point 08 through pixel point 02, or incorporates the correlation between pixel point 01 and pixel point 11 through pixel point 06, etc. Through the recursive calculation of multiple CrossTransformer models, the correlation fusion between each pixel point and other pixel points is achieved, and the fusion can be automatically carried out according to the magnitude of the correlation, that is, first fuse pixel points with large correlations, and then fuse pixel points with small correlations, automatically reducing the weight of pixel points not on the cross.

[0073] Step 5: Use the simulation training samples to train the battery recognition model to obtain the trained battery recognition model.

[0074] In this embodiment, N = 4. The specific implementation process of using the simulation training samples to train the battery recognition model includes:

[0075] Step 5.1: In the first CrossTransformer model, feature extraction is performed on each pixel point in the pixel matrix to obtain the feature vector of each pixel point in the pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; according to the query matrix Q, the key matrix K, and the value matrix V of each pixel point, the attention coefficient between this pixel point and other pixel points in the same row and the same column of this pixel point (i.e., the pixel points on the cross) is calculated (as shown in Equation (1)); each pixel point is replaced with the corresponding attention coefficient, and the first pixel matrix is composed of the attention coefficients.

[0076] Step 5.2: In the second CrossTransformer model, feature extraction is performed on each pixel point in the first pixel matrix to obtain the feature vector of each pixel point in the first pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; according to the query matrix Q, the key matrix K, and the value matrix V of each pixel point, the attention coefficient between this pixel point and other pixel points in the same row and the same column of this pixel point is calculated; each pixel point is replaced with the corresponding attention coefficient, and the second pixel matrix is composed of the attention coefficients.

[0077] Step 5.3: In the third CrossTransformer model, feature extraction is performed on each pixel point in the second pixel matrix to obtain the feature vector of each pixel point in the second pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; according to the query matrix Q, the key matrix K, and the value matrix V of each pixel point, the attention coefficient between this pixel point and other pixel points in the same row and the same column of this pixel point is calculated; each pixel point is replaced with the corresponding attention coefficient, and the third pixel matrix is composed of the attention coefficients.

[0078] Step 5.4: In the fourth CrossTransformer model, feature extraction is performed on each pixel point in the third pixel matrix to obtain the feature vector of each pixel point in the third pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; according to the query matrix Q, the key matrix K, and the value matrix V of each pixel point, the attention coefficient between this pixel point and other pixel points in the same row and the same column of this pixel point is calculated; each pixel point is replaced with the corresponding attention coefficient, and the fourth pixel matrix is composed of the attention coefficients.

[0079] Step 5.5: The fourth pixel matrix is passed through the YOLO detection head for target prediction to obtain the target prediction vector [x min , y min , x max , y max, confidence score, label index, label score].

[0080] The target prediction vector includes the coordinate values (x min , y max ) of the upper left corner of the detection box, the coordinate values (x max , y min ) of the lower right corner, the confidence score, the label index, and the label score. The confidence score represents the probability score that the detection box contains the target. The label index represents the category of the target. The label score represents the probability score that the target belongs to the category.

[0081] The purpose of training the battery recognition model with simulation training samples is only to enable the operator to better find the area near the minimum value. To reduce the training time and the number of training samples required, a dropout layer is added to each CrossTransformer model. When training the battery recognition model with simulation training samples, the dropout layer randomly discards some parameters. In this embodiment, the dropout rate of the dropout layer is 0.25, that is, the dropout layer randomly discards 1 / 4 of the parameters during training and only trains 3 / 4 of the parameters.

[0082] Step 6: Obtain the real image of the battery stack at the battery recycling site, and label the real image of the battery stack. The labeling content includes batteries, battery smooth surfaces, battery electrode surfaces, and carrying straps.

[0083] Step 7: Preprocess each labeled real image of the battery stack to obtain the corresponding pixel matrix, and all pixel matrices form the real training samples.

[0084] In this embodiment, the specific implementation process of preprocessing the real image of the battery stack includes:

[0085] Step 7.1: Use Opencv to read the real image of the battery stack and convert it into a pixel matrix.

[0086] Step 7.2: Normalize each pixel matrix to obtain a normalized pixel matrix.

[0087] Step 7.3: Perform data augmentation on each normalized pixel matrix to obtain the pixel matrix after data augmentation processing.

[0088] Step 8: Retrain the trained battery recognition model (i.e., the battery recognition model obtained in Step 5) with the real training samples to obtain the target battery recognition model.

[0089] The specific implementation process of retraining the trained battery recognition model with real training samples is the same as that of training the battery recognition model with simulation training samples.

[0090] When retraining the trained battery recognition model with real training samples, all dropout layers are turned off.

[0091] In the gradient backpropagation of deep learning, there are plateaus, saddle points and other flat areas in its high-dimensional space, and there are also many local minima, which are not global minima. When the operator is transmitted, it may encounter some areas that require a very long time to learn to cross this area, or even cannot escape from the local minimum. Therefore, a large number of training samples are needed to give the operator enough roaming opportunities and time in the high-dimensional space. From another perspective, if the initial position of the operator is close to the global minimum, the learning time of the model will be greatly shortened; if the initial position of the operator is far from the global minimum, a large amount of learning time is required. Training the battery recognition model with simulation training samples enables the model to learn some prior knowledge, and then retraining the battery recognition model with real training samples is conducive to finding the global minimum more quickly.

[0092] Based on the same concept, an embodiment of the present invention further provides a disorderly grasping method, including the following steps:

[0093] Step 1: Obtain the battery stacking pictures at the battery recycling site in real time.

[0094] In this embodiment, the battery stacking pictures at the battery recycling site are obtained by using the depth camera realsense d435i.

[0095] Step 2: Preprocess the battery stacking pictures to obtain a pixel matrix.

[0096] The specific implementation process of preprocessing the battery stacking pictures is as follows: convert the battery stacking pictures into a pixel matrix; perform standardization processing on the pixel matrix to obtain a standardized pixel matrix.

[0097] Step 3: Call the target battery recognition model, and the target battery recognition model is trained according to the battery recognition model training method described above.

[0098] Step 4: Based on the target battery recognition model, identify the pixel matrix in Step 2 to obtain a target prediction vector.

[0099] The target prediction vector includes the coordinate values (x min , y max ) of the upper left corner of the detection box (i.e., the bounding box), and the coordinate values (xmax , y min ), confidence score, label index, and label score, where the confidence score represents the probability score that the detection box contains the target, the label index represents the category of the target, and the label score represents the probability score that the target belongs to the category.

[0100] Step 5: Calculate the battery key point coordinates based on the target prediction vector, convert the key point coordinates into world coordinates, and send the world coordinates to the grasping device.

[0101] In this embodiment, the key point is the center point of the battery, and the grasping device is a manipulator or a robotic arm. The world coordinates of the center point are transmitted to the controller that controls the manipulator or the robotic arm through the Modbus RS485 protocol.

[0102] Step 6: Generate a motion path based on the world coordinates of the key point and the current position of the grasping device, and control the grasping device to move along the motion path and perform grasping.

[0103] The specific embodiments disclosed above are only for the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or variations, which should be covered within the protection scope of the present invention.

Claims

1. A method for training a battery recognition model, characterized in that, The method includes the following steps: Construct a battery model, simulate stacking pictures of batteries in different postures according to the battery model, and obtain battery stacking simulation pictures; Annotate each of the battery stacking simulation pictures, and the annotation content includes batteries, battery smooth surfaces, battery electrode surfaces, and lifting straps; Preprocess each annotated battery stacking simulation picture to obtain a corresponding pixel matrix, and all pixel matrices form a simulation training sample; Construct a battery recognition model, the battery recognition model includes N CrossTransformer models, and each CrossTransformer model only calculates the attention coefficients between each pixel point and other pixel points in the row and column where the pixel point is located when calculating the attention coefficients; replace each pixel point in the pixel matrix with the attention coefficient of the pixel point, and the pixel matrix composed of attention coefficients is used as the input of the next CrossTransformer model; Use the simulation training sample to train the battery recognition model to obtain a trained battery recognition model; Obtain the real picture of the battery stacking at the battery recycling site, annotate the real picture of the battery stacking, and the annotation content includes batteries, battery smooth surfaces, battery electrode surfaces, and lifting straps; Preprocess each annotated real picture of the battery stacking to obtain a corresponding pixel matrix, and all pixel matrices form a real training sample; Use the real training sample to retrain the trained battery recognition model to obtain a target battery recognition model.

2. The method for training a battery identification model according to claim 1, wherein The specific implementation process of preprocessing the battery stacking simulation picture or the real picture of the battery stacking includes: Convert each picture into a pixel matrix; Perform normalization processing on each pixel matrix to obtain a normalized pixel matrix; Perform data augmentation on each normalized pixel matrix to obtain a pixel matrix after data augmentation processing.

3. The battery identification model training method according to claim 2, wherein The data augmentation includes mirroring, rotation, scaling, cropping, translation, and Gaussian noise.

4. The method for training a battery identification model according to claim 1, wherein The specific implementation process of training the battery recognition model using the simulation training sample or the real training sample includes: In the first CrossTransformer model, extract features from each pixel point in the pixel matrix to obtain the feature vector of each pixel point in the pixel matrix; convert the feature vector of each pixel point into a query matrix Q, a key matrix K, and a value matrix V; calculate the attention coefficients between each pixel point and other pixel points in the row and column where the pixel point is located according to the query matrix Q, the key matrix K, and the value matrix V of each pixel point; replace each pixel point with the corresponding attention coefficient, and the first pixel matrix is composed of attention coefficients; In the second CrossTransformer model, feature extraction is performed on each pixel point in the first pixel matrix to obtain the feature vector of each pixel point in the first pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; attention coefficients between each pixel point and other pixel points in the row and column where the pixel point is located are calculated according to the query matrix Q, the key matrix K, and the value matrix V of each pixel point; each pixel point is replaced with the corresponding attention coefficient, and the second pixel matrix is formed by the attention coefficients. By analogy, in the Nth CrossTransformer model, feature extraction is performed on each pixel point in the (N - 1)th pixel matrix to obtain the feature vector of each pixel point in the (N - 1)th pixel matrix; the feature vector of each pixel point is converted into a query matrix Q, a key matrix K, and a value matrix V; attention coefficients between each pixel point and other pixel points in the row and column where the pixel point is located are calculated according to the query matrix Q, the key matrix K, and the value matrix V of each pixel point; each pixel point is replaced with the corresponding attention coefficient, and the Nth pixel matrix is formed by the attention coefficients. The Nth pixel matrix is passed through the YOLO detection head for target prediction to obtain a target prediction vector, where the target prediction vector includes the coordinate values of the upper left and lower right corners of the detection box, a confidence score, a label index, and a label score, where the confidence score represents the probability score that the detection box contains a target, the label index represents the category of the target, and the label score represents the probability score that the target belongs to the category.

5. The battery identification model training method according to any one of claims 1 to 4, characterized in that The calculation formula for the attention coefficient is: Among them, A represents the attention coefficient, d k represents the number of feature channels of the query matrix Q, the key matrix K, and the value matrix V.

6. The battery identification model training method according to any one of claims 1 to 4, characterized in that, Each CrossTransformer model includes a dropout layer, and when retraining the trained battery recognition model using the real training samples, all dropout layers are turned off.

7. The method for training a battery identification model according to claim 1, wherein N is 4.

8. A disordered grasping method, characterized in that, The method includes the following steps: Obtain in real time the battery stacking pictures at the battery recycling site; Preprocess the battery stacking pictures to obtain a pixel matrix; Call the target battery recognition model, where the target battery recognition model is trained according to the battery recognition model training method described in any one of claims 1 to 7; Based on the target battery recognition model, identify the pixel matrix to obtain a target prediction vector; Calculate the battery key point coordinates according to the target prediction vector, convert the key point coordinates into world coordinates, and send the world coordinates to the grasping device; Generate a motion path according to the world coordinates of the key points and the current position of the grasping device, and control the grasping device to move and grasp according to the motion path.

9. An electronic device, characterized in that, The device includes: A memory for storing a computer program; A processor for implementing the battery recognition model training method described in any one of claims 1 to 7 or the disordered grasping method described in claim 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the battery identification model training method according to any one of claims 1 to 7 or the disordered grasping method according to claim 8.

Citation Information

Patent Citations

  • Portrait matting method and device, computer equipment and readable storage medium

    CN113870283A

  • Real-time target detection method based on Pearson coefficient matrix and attention fusion

    CN114187569A