End-to-end crushing operation point determination method and device, equipment and medium

Through an end-to-end method, combining camera images and lidar point clouds, the crushing operation point information is extracted and decoded, and the problems of low intelligence and insufficient robustness of crushing operation point detection in the prior art are solved, and high accuracy and efficient detection in extreme environments are achieved.

CN120013914APending Publication Date: 2025-05-16GUANGXI LIUGONG METATHINGS TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510117903.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the existing technology, in the unmanned crushing operation, the detection of crushing operation point is low, the detection capability is insufficient in dynamic environments, and the robustness is poor, especially in extreme lighting conditions, which may lead to detection failure.

Method used

An end-to-end method is adopted to obtain the image to be broken through the camera and segment it. The target object point cloud is extracted and feature extraction is performed in combination with the lidar point cloud, the point cloud token and vehicle status token are obtained, and the broken job point position and angle prompt value are input into the decoder to decode it.

Benefits of technology

It improves the intelligence and robustness of crushing operation point determination, can maintain the accuracy and effectiveness of detection in extreme environments, and achieves fully intelligent crushing operation point determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013914A_ABST
    Figure CN120013914A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an end-to-end crushing operation point determination method and device, equipment and a medium. The method comprises the following steps: acquiring a to-be-crushed image of an unmanned crushing operation area through a camera, and segmenting the to-be-crushed image to obtain a segmentation result of a target object; laser radar point clouds are obtained through the laser radar, and target object point clouds corresponding to the segmentation result of the target object are extracted from the laser radar point clouds; performing point cloud feature extraction on the target object point cloud to obtain a plurality of point cloud tokens; obtaining vehicle information of the excavator, and performing feature extraction on the vehicle information to obtain a vehicle state token; and inputting the point cloud token and the vehicle state token into a decoder, and decoding to obtain a crushing operation point position and a crushing operation angle prompt value. According to the method, through an end-to-end crushing operation point determination mode, the intelligent degree of crushing operation point determination can be improved, and through combination of the point cloud and the image data, the problem of insufficient robustness under an extreme condition can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an end-to-end crushing operation point determination method, device, equipment and medium. Background Art

[0002] When an unmanned excavator is performing unmanned crushing operations, such as unmanned crushing operations at a fixed material port or unmanned crushing operations during autonomous driving in a fixed area, it is necessary to detect the operating points related to the crushing operation in real time to ensure the reliability of the crushing operation.

[0003] The solutions in the prior art either do not introduce intelligent methods for crushing operation point detection, or the technology used in crushing operation point detection is of low intelligence. The technical solutions provided by the prior art are not end-to-end crushing operation point determination methods, resulting in the inability of the prior art solutions to detect in dynamic environments, low intelligence, low robustness in complex scenarios, and possible detection failures, especially in extreme cases such as when the image of the crushing area is blurred due to strong light, backlight, weak light, dirty camera lens, etc., and the determined crushing operation point is unreasonable. Summary of the invention

[0004] The present invention provides an end-to-end crushing operation point determination method, device, equipment and medium to improve the intelligence and robustness of crushing operation point determination.

[0005] According to one aspect of the present invention, an end-to-end method for determining a crushing operation point is provided, the method comprising:

[0006] Acquire the image to be crushed in the unmanned crushing operation area through the camera, and segment the image to be crushed to obtain the segmentation result of the target object;

[0007] Acquire a laser radar point cloud through a laser radar, and extract a target object point cloud corresponding to the segmentation result of the target object from the laser radar point cloud;

[0008] Extracting point cloud features from the target object point cloud to obtain a plurality of point cloud tokens;

[0009] Acquire vehicle information of the excavator, and perform feature extraction on the vehicle information to obtain a vehicle status token;

[0010] The point cloud token and the vehicle status token are input into a decoder, and the position of the crushing operation point and the crushing operation angle prompt value are obtained by decoding.

[0011] According to another aspect of the present invention, an end-to-end crushing operation point determination device is provided, the device comprising:

[0012] An image segmentation module is used to obtain an image to be crushed in an unmanned crushing operation area through a camera, and to segment the image to be crushed to obtain a segmentation result of a target object;

[0013] A point cloud extraction module, used for acquiring a laser radar point cloud through a laser radar, and extracting a target object point cloud corresponding to a segmentation result of the target object from the laser radar point cloud;

[0014] A point cloud token determination module is used to extract point cloud features from the target object point cloud to obtain a plurality of point cloud tokens;

[0015] A vehicle status token determination module is used to obtain vehicle information of the excavator and perform feature extraction on the vehicle information to obtain a vehicle status token;

[0016] The crushing operation point determination module is used to input the point cloud token and the vehicle status token into a decoder, and decode to obtain the crushing operation point position and the crushing operation angle prompt value.

[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0018] at least one processor; and

[0019] a memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the end-to-end crushing operation point determination method described in any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the end-to-end crushing operation point determination method described in any embodiment of the present invention when executed.

[0022] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the end-to-end crushing operation point determination method according to any embodiment of the present invention.

[0023] The technical solution of the embodiment of the present invention is to obtain the image to be crushed in the unmanned crushing operation area through a camera, and segment the image to be crushed to obtain the segmentation result of the target object; obtain the laser radar point cloud through the laser radar, and extract the target object point cloud corresponding to the segmentation result of the target object in the laser radar point cloud; perform point cloud feature extraction on the target object point cloud to obtain multiple point cloud tokens; obtain the vehicle information of the excavator, and perform feature extraction on the vehicle information to obtain the vehicle status token; input the point cloud token and the vehicle status token into the decoder, and decode to obtain the crushing operation point position and the crushing operation angle prompt value, which solves the problem of insufficient intelligence in determining the crushing operation point. Through the end-to-end crushing operation point determination method, the intelligence level of the crushing operation point determination can be improved, and through the combination of point cloud and image data, the problem of insufficient robustness in extreme cases can be avoided.

[0024] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1a is a flow chart of an end-to-end crushing operation point determination method provided according to the first embodiment of the present invention;

[0027] Figure 1b is a schematic diagram of camera and laser radar installation provided according to Embodiment 1 of the present invention;

[0028] Figure 1c is another schematic diagram of camera and laser radar installation provided according to the first embodiment of the present invention;

[0029] Figure 1d 2 is a schematic diagram of the structure of a point cloud feature extraction backbone network provided according to the first embodiment of the present invention;

[0030] Figure 1e is a schematic diagram of an MLP structure provided according to Embodiment 1 of the present invention. An MLP is generally composed of several linear layers;

[0031] Figure 1f is a schematic diagram of a decoder structure provided according to Embodiment 1 of the present invention;

[0032] Figure 1g is a schematic diagram of image annotation provided according to an embodiment of the present invention;

[0033] Figure 1h is a point cloud annotation schematic diagram provided according to an embodiment of the present invention;

[0034] Figure 2a is a flow chart of an end-to-end crushing operation point determination method provided according to the second embodiment of the present invention;

[0035] Figure 2b is a schematic diagram of a crushing operation point determination result provided according to the second embodiment of the present invention;

[0036] Figure 2c is a schematic diagram of an excavator performing autonomous crushing according to the second embodiment of the present invention;

[0037] Figure 3 is a flow chart of another end-to-end crushing operation point determination method provided according to the second embodiment of the present invention;

[0038] Figure 4 is a schematic structural diagram of an end-to-end crushing operation point determination device provided according to Embodiment 3 of the present invention;

[0039] Figure 5 It is a structural schematic diagram of an electronic device for implementing the end-to-end crushing operation point determination method of an embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0042] Embodiment 1

[0043] Figure 1a This is a flow chart of an end-to-end method for determining a crushing operation point according to the first embodiment of the present invention. This embodiment is applicable to the case where an unmanned excavator determines a crushing operation point when performing a crushing operation autonomously. The method can be executed by an end-to-end crushing operation point determination device, which can be implemented in the form of hardware and / or software. The end-to-end crushing operation point determination device can be configured in an electronic device, such as a computer or an excavator controller. Figure 1a As shown, the method includes:

[0044] Step 110: Acquire the image to be crushed in the unmanned crushing operation area through a camera, and segment the image to be crushed to obtain a segmentation result of the target object.

[0045] The camera can be a monocular camera. The camera can be installed on the roadside of the unmanned crushing operation area or on the excavator end. In the embodiment of the present invention, while the camera is used to collect the image to be crushed, a laser radar is also used to obtain the laser radar point cloud. In order to ensure that the image and the point cloud can be accurately combined to determine the crushing operation point, the camera and the laser radar can be installed in the same range to form an overlapping field of view. The overlapping field of view is the unmanned crushing operation area. Exemplarily, Figure 1b : is a schematic diagram of the installation of a camera and a laser radar according to the first embodiment of the present invention. The camera and the laser radar can be installed on the columns on both sides of the pit at the same time to form an overlapping field of view covering the unmanned crushing operation area. Another exemplary embodiment, Figure 1c This is another camera and laser radar installation schematic diagram provided according to Embodiment 1 of the present invention. The camera and laser radar can be installed on the roof of the cockpit at the vehicle end to form a visual overlap area covering the unmanned crushing operation area.

[0046] The image to be broken captured by the camera can be segmented to obtain the segmentation result of the target object. The segmentation method can be implemented by using an image target detection algorithm. Exemplarily, it is implemented by an image instance segmentation network, such as yolov5-segment network, yolov6-segment network, yolov7-segment network, yolov8-segment network, yolov9-segment network, and yolov10-segment network.

[0047] Exemplarily, taking the yolov7-segment network as an example, the yolov7-segment network is a fully convolutional network based on a deep convolutional neural network (CNN). The yolov7-segment network consists of an input layer (image), a backbone feature extraction network (backbone), and a detection head (head). The input layer scales and normalizes the image. Scaling is to scale the input image resolution to 640×640, and there are many normalization methods. In the present invention, yolov7-segment normalizes the image pixel value from 0-255 to 0-1. For example, the normalization formula is x=x / 255. The normalized image data can be sent to the yolov7-segment network for reasoning. The yolov7-segment network adds a head instance segmentation branch based on the yolov7 network, including adding a mask output (proto) with an output size of 160×160×32. At the same time, the three outputs of the network are changed from 20×20×255, 40×40×255, and 80×80×255 to 20×20×272, 40×40×272, and 80×80×272, respectively. Each channel has an additional 32-dimensional output for multiplication with proto to obtain the instance segmentation mask. The backbone of the yolov7-segment network is built by the Efficient Layer Aggregation Network (ELAN) substructure. The ELAN network structure design conforms to the hardware characteristics of the GPU, and its calculation is efficient and time-saving. The ELAN network structure is designed based on a stacking paradigm (concatenation-based). Its input is processed by 3×3 convolution layers and cross-layer connections, then stacked together and extracted by 1×1 convolution. The head part of the yolov7-segment detection head continues the design of yolov5. The three-way backbone output with different depths is upsampled successively to perform top-down feature fusion. Each fusion result is output by REP and CBM for the final model output. Among them, REP is a reparameterizable convolution layer structure. The three data streams in its module can eventually be fused into a convolution layer for equivalence, which can improve the performance of the model while reducing the amount of calculation. CBM is a simple representation of the connection of three layers of convolution layer, batch normalization layer, and activation layer (sigmoid). Yolov7-segment uses two activation functions, Sigmoid and SiLU. The main purpose of the activation function is to introduce nonlinear changes to the neural network and increase the network's expressive power.

[0048] In an optional implementation of the embodiment of the present invention, the image to be crushed is segmented to obtain the segmentation result of the target object, including: segmenting the image to be crushed through an image segmentation network to obtain the segmentation result of at least one of the following target objects in the image: each stone to be crushed, the stone to be moved, and the grid hole stuck material area.

[0049] Among them, the stones to be crushed may be stones that need to be crushed by the excavator through the telescopic boom arm. The stones to be moved may be stones that need to be moved into the grid holes by the excavator through the telescopic boom arm. The grid hole stuck material area may be an area where the excavator needs to poke the material through the telescopic boom arm to keep the grid holes unobstructed so that the broken stones can enter the holes. The embodiment of the present invention can output the stone and grid hole stuck material target mask through yolov7-segment reasoning.

[0050] Step 120: Acquire a laser radar point cloud through the laser radar, and extract a target object point cloud corresponding to the segmentation result of the target object from the laser radar point cloud.

[0051] The target object point cloud corresponding to the target mask can be extracted from the LiDAR point cloud. For example, when the segmentation result is a stone, the target object point cloud corresponding to the stone can be extracted from the LiDAR point cloud. For another example, when the segmentation result is a grid hole stuck material area, the target object point cloud corresponding to the grid hole stuck material area can be extracted from the LiDAR point cloud. By combining the point cloud on the basis of the image segmentation result, the coordinates of the target object can be made clear. Even in extreme cases, such as when the image is blurred due to strong light, backlight, weak light, dirty camera lens, etc., because the point cloud is combined on the basis of image segmentation, when the outline of the target object is unclear, the coordinates of the point cloud can also be combined to obtain the specific coordinate information of the target object, thereby ensuring the accuracy of the determination of the crushing operation point.

[0052] In an optional implementation of an embodiment of the present invention, a target object point cloud corresponding to a segmentation result of the target object is extracted in a lidar point cloud, including: obtaining a joint calibration result of a camera and a lidar; and according to the joint calibration result, extracting a target object point cloud corresponding to a segmentation result of the target object in the lidar point cloud.

[0053] The laser radar point cloud corresponding to the target segmentation mask is extracted by the joint calibration results of the camera and the laser radar in advance, such as extracting the corresponding stone laser radar point cloud from the stone image segmentation mask, which can ensure the accuracy of target object recognition. Among them, the joint calibration results may include: camera intrinsic parameter matrix, camera laser radar extrinsic parameter matrix, and camera image distortion coefficient. The camera intrinsic parameter matrix describes the parameters of the internal geometric characteristics of the camera, and these parameters define how to convert the three-dimensional world coordinates to the two-dimensional image coordinates. Common parameters include focal length, principal point, and pixel scale coefficient. The extrinsic parameter matrix describes the relative position and direction relationship between the camera and the laser radar. The extrinsic parameter matrix consists of a rotation matrix and a translation vector, which can transform the coordinate system of one sensor into the coordinate system of another sensor. The laser radar point cloud data can be projected onto the camera image through the extrinsic parameter matrix, or the feature points in the image can be back-projected into the three-dimensional space to achieve sensor data fusion. The camera image distortion coefficient is used to quantify the nonlinear distortion of the image, correct the image, and improve the accuracy of vision-based measurement.

[0054] Step 130: extract point cloud features from the target object point cloud to obtain multiple point cloud tokens.

[0055] Among them, point cloud feature extraction can be implemented by a point cloud feature extraction backbone network. For example, the point cloud feature extraction backbone network can be a convolutional network (Sparsely Embedded Convolutional Detection, SECOND) that processes point cloud data. Figure 1d Schematic diagram of the structure of a point cloud feature extraction backbone network provided according to the first embodiment of the present invention. Figure 1dAs shown, in the present invention, only the network before the SECOND network detection head, that is, the backbone network of the SENCOND network is used to extract point cloud features. Specifically, the SENCOND network is composed of a feature voxelization module, a voxelized feature extraction module, a 3D sparse convolution module and a detection head RPN module. In the present invention, only the backbone network of the SENCOND network is used to extract point cloud features, that is, the RPN module is not used. The backbone network reasoning process includes voxel grouping of the original point cloud, and each voxelized grouping square includes the number of point clouds and their coordinates inside it, followed by voxel feature extraction (Vexel Feature Extrator, VFE). The VFE layer takes the data of all points in the same voxel as input, and uses a fully connected network (FCN) composed of a linear connection (Linear) layer, a batch normalization (BatchNorm) layer and an activation layer (ReLU) to extract voxel features. The output of the voxel feature extraction module is sent to the 3D sparse convolution layer for further feature extraction. The 3D sparse convolution designs the 3D sparse convolution input and output rules based on the hash table. It determines in advance which voxels are not empty and need to be calculated by 3D convolution and determines their output positions, so as to solve the useless calculation of empty voxel features and reduce the amount of calculation.

[0056] In an embodiment of the present invention, the voxel features after the 3D sparse convolution can be tokenized. Specifically, they can be flattened first, and then position encoded to obtain multiple point cloud tokens to prepare for decoding.

[0057] Step 140: Obtain vehicle information of the excavator, perform feature extraction on the vehicle information, and obtain a vehicle status token.

[0058] The vehicle information of the excavator may be information required when the excavator performs crushing operations. For example, the vehicle information may include vehicle positioning data and excavator boom joint angle position data. Feature extraction of the vehicle information to obtain a vehicle state token may be feature extraction and tokenization. The feature extraction of the vehicle information may be implemented by a multi-layer perceptron (MLP) network. Figure 1e Schematic diagram of an MLP structure provided according to Embodiment 1 of the present invention. MLP is generally composed of several linear layers.

[0059] In an optional implementation of an embodiment of the present invention, vehicle information of an excavator is obtained, and features of the vehicle information are extracted to obtain a vehicle status token, including: obtaining vehicle positioning data of the excavator and joint angle position data of the excavator's boom arm; performing feature extraction and tokenization on the vehicle positioning data and joint angle position data through a multi-layer perceptron network to obtain a vehicle status token.

[0060] Specifically, the vehicle positioning data of the excavator and the joint angle position data of the excavator boom can be feature extracted through MLP and tokenized to obtain a vehicle state token. The determination of the vehicle state token and the determination of the point cloud token can be performed simultaneously.

[0061] Step 150: Input the point cloud token and the vehicle status token into the decoder, and decode to obtain the crushing operation point position and the crushing operation angle prompt value.

[0062] In an embodiment of the present invention, the decoder may be an attention mechanism decoder (Transformer Decoder). Through the Transformer Decoder, the point cloud token of the target object such as a stone and the vehicle status token may be decoded and inferred to obtain the position of the crushing operation point and the crushing operation angle prompt value.

[0063] Specifically, in an optional implementation of an embodiment of the present invention, the point cloud token and the vehicle status token are input into a decoder, and decoded to obtain the crushing operation point position and the crushing operation angle prompt value, including: inputting the point cloud token and the vehicle status token into the decoder, and decoding to obtain the embedded vector of the crushing operation point position and the crushing operation angle prompt; inputting the embedded vector into a multi-layer perceptron detection head to obtain the crushing operation point position and the crushing operation angle prompt value.

[0064] Figure 1f 1 is a schematic diagram of a decoder structure provided according to the first embodiment of the present invention. Figure 1f As shown in the figure, the decoder Transformer Decoder consists of 4 transformer decoder blocks, each of which includes a self-attention layer, a normalization layer (LayerNorm), a cross attention layer, and a feed-forward neural network layer (FFN). Among them, the cloud tokens and vehicle status tokens can be Figure 1f The input is shown in the Transformer Decoder to obtain the embedded vector of the crushing operation point location and the crushing operation angle prompt.

[0065] Afterwards, you can Figure 1f As shown in the figure, the output of Transformer Decoder is passed through the heads of two MLPs to obtain the location of the crushing operation point and the crushing operation angle prompt value. Among them, one MLP can be used to output categories (classes), such as stones to be crushed, stones to be moved, and grid hole stuck areas; the other MLP can be used to output the corresponding operation coordinates and excavator operation angle, that is, the location of the crushing operation point and the crushing operation angle prompt value.

[0066] The technical solution of this embodiment is to segment the image to be crushed through a segmentation network such as yolov7 to obtain each stone to be crushed, the stone to be moved and the grid hole stuck material area; then through the joint calibration result of the camera and the laser radar, the target object point cloud corresponding to the segmentation result of the target object is extracted in the laser radar point cloud; the target object point cloud is sent to the point cloud data processing network such as the backbone network of SECOND, and tokenized to obtain multiple point cloud tokens; at the same time, the vehicle positioning data and joint angle position data are feature extracted and tokenized through a multi-layer perceptron network to obtain a vehicle status token; finally, the point cloud token and the vehicle status token are input into a decoder such as Transformer Decoder for decoding, and the two MLPs are combined to obtain the crushing operation point position and the crushing operation angle prompt value, which solves the problem of insufficient intelligence when determining the crushing operation point. The end-to-end crushing operation point determination method can improve the intelligence level of the crushing operation point determination, and the combination of point cloud and image data can avoid the problem of insufficient robustness in extreme cases.

[0067] In the embodiment of the present invention, from collecting the image to be crushed to obtaining the position of the crushing operation point and the crushing operation angle prompt value, an end-to-end approach is adopted without manual operation, thereby realizing fully intelligent determination of the crushing operation point, which can improve the accuracy, robustness, rationality and efficiency of the determination of the crushing operation point.

[0068] It should be noted that each network model used in the embodiment of the present invention can be used for data collection and training generation before specific application. Data needs to be annotated before training. The annotation process includes image annotation and point cloud annotation. Image annotation can be performed first and then point cloud annotation. Image annotation can mark the stones that need to be broken, the stones that need to be moved, and the grid hole stuck material area that needs to be poked in each image. The image annotation results can be stored in a txt file. Figure 1g The present invention provides an image annotation schematic diagram according to an embodiment of the present invention.

[0069] Point cloud annotation can mark the coordinates of the crushing point of each stone to be crushed, the coordinates of the moving point of the stone to be moved, and the coordinates of the material poking point of the grid hole, and combine the material status and vehicle status in the image to mark the execution prompt angle value of the operation action. The point cloud annotation results can be saved in a txt file. Figure 1h The present invention provides a point cloud annotation schematic diagram according to an embodiment of the present invention.

[0070] Based on Figure 1g and Figure 1h The annotation results can be used to train the model in the training server. For example, the training can be completed by the collaboration of the CPU and GPU in the training server, and the overall algorithm domain controller can take less than 50ms.

[0071] Embodiment 2

[0072] Figure 2a 1 is a flow chart of an end-to-end crushing operation point determination method provided according to the second embodiment of the present invention. This embodiment is a further addition to the above technical solution. The technical solution in this embodiment can be combined with each optional solution in one or more of the above embodiments. Figure 2a As shown, the method includes:

[0073] Step 210: Acquire the image to be crushed in the unmanned crushing operation area through a camera, and segment the image to be crushed to obtain a segmentation result of the target object.

[0074] In an optional implementation of the embodiment of the present invention, the image to be crushed is segmented to obtain the segmentation result of the target object, including: segmenting the image to be crushed through an image segmentation network to obtain the segmentation result of at least one of the following target objects in the image: each stone to be crushed, the stone to be moved, and the grid hole stuck material area.

[0075] Step 220: Acquire a laser radar point cloud through the laser radar, and extract a target object point cloud corresponding to the segmentation result of the target object in the laser radar point cloud.

[0076] In an optional implementation of an embodiment of the present invention, a target object point cloud corresponding to a segmentation result of the target object is extracted in a lidar point cloud, including: obtaining a joint calibration result of a camera and a lidar; and according to the joint calibration result, extracting a target object point cloud corresponding to a segmentation result of the target object in the lidar point cloud.

[0077] Step 230: extract point cloud features from the target object point cloud to obtain multiple point cloud tokens.

[0078] Step 240: Obtain vehicle information of the excavator, and perform feature extraction on the vehicle information to obtain a vehicle status token.

[0079] In an optional implementation of an embodiment of the present invention, vehicle information of an excavator is obtained, and features of the vehicle information are extracted to obtain a vehicle status token, including: obtaining vehicle positioning data of the excavator and joint angle position data of the excavator's boom arm; performing feature extraction and tokenization on the vehicle positioning data and joint angle position data through a multi-layer perceptron network to obtain a vehicle status token.

[0080] Step 250: Input the point cloud token and the vehicle status token into the decoder, and decode to obtain the crushing operation point position and the crushing operation angle prompt value.

[0081] In an optional implementation of an embodiment of the present invention, the point cloud token and the vehicle status token are input into a decoder, and decoded to obtain the crushing operation point position and the crushing operation angle prompt value, including: inputting the point cloud token and the vehicle status token into the decoder, and decoding to obtain the embedded vector of the crushing operation point position and the crushing operation angle prompt; inputting the embedded vector into a multi-layer perceptron detection head to obtain the crushing operation point position and the crushing operation angle prompt value.

[0082] Step 260: According to the location of the crushing operation point and the crushing operation angle prompt value, the crushing operation trajectory of the excavator is planned to obtain the crushing operation trajectory.

[0083] When planning the crushing operation trajectory, many factors can be considered. For example, when planning the crushing operation trajectory, wall collision safety, rationality of operation actions, and operation efficiency can be considered. Among them, wall collision safety can input the wall coordinates of the unmanned crushing operation area into the crushing operation trajectory planning. When planning the crushing operation trajectory, the wall coordinates are avoided to prevent the excavator from hitting the wall. The rationality of the operation action can be to optimize the sequence of the excavator's actions, consider the angle limit of the excavator's telescopic boom arm, and the smoothness of the action. When considering the operation efficiency, the intensity and efficiency of the crushing can be considered.

[0084] Step 270: Control the telescopic boom arm of the excavator to perform the crushing operation according to the crushing operation trajectory.

[0085] Figure 2b 1 is a schematic diagram of a crushing operation point determination result provided according to the second embodiment of the present invention. Figure 2b As shown, each stone can be segmented in the image, and the coordinates of the center point of each stone on the upper surface can be obtained through the point cloud, so as to control the excavator to crush the center point of the stone at the crushing operation angle. Figure 2c The diagram is a schematic diagram of an excavator performing autonomous crushing according to the second embodiment of the present invention. The telescopic boom arm of the excavator is controlled according to the crushing operation trajectory to crush, move and poke each stone in turn.

[0086] Figure 3FIG. 1 is a flow chart of another end-to-end method for determining a crushing operation point according to the second embodiment of the present invention. Figure 3 As shown in the figure, an application example of the end-to-end crushing operation point determination method can be: fully collect crushing operation data under various working conditions, select and annotate the data, use the annotated data for model training, and the trained model can be converted from the pytorch model format (pth file) to the tensorrt model format (engine file) and deployed to the GPU domain controller. The end-to-end crushing operation point determination system is powered on, and the node starts and self-checks. Among them, the end-to-end crushing operation point determination system can include multiple nodes, such as data acquisition nodes, perception nodes, planning nodes, control nodes, and cloud platform interaction nodes. Each node starts in turn and performs node status self-check to ensure that all software nodes work normally. The data acquisition node collects the image to be crushed, the laser radar point cloud, the vehicle positioning data, and the joint angle position data in real time and in parallel through cameras and laser radars. The data acquisition node can synchronize the timestamp based on the laser radar point cloud timestamp to ensure that all data are at the same time point. The synchronized images to be crushed, LiDAR point cloud, vehicle positioning data and joint angle position data can be sent to the perception node for end-to-end model reasoning to obtain the results of the crushing operation point determination, that is, the crushing operation point positions and crushing operation angle prompt values ​​of the stones to be crushed, the stones to be moved, and the grid hole stuck material area. The crushing operation point positions and crushing operation angle prompt values ​​can be sent to the planning node for unmanned crushing operation trajectory planning. Trajectory planning can simultaneously consider factors such as wall collision safety, operation action rationality and operation efficiency. The crushing operation trajectory generated by the planning node can be sent to the control node. The control node uses control algorithms such as proportional integral differential (PID) or model predictive control (MPC) based on the crushing operation trajectory to control the telescopic boom arm of the unmanned excavator to perform crushing operations. The operation perception results, vehicle body status, operation progress and other data during the unmanned crushing operation can be sent to the cloud platform for operation status display and manual monitoring. At the same time, the instructions such as start operation, pause operation and end operation issued by the cloud platform can control the operation status of the unmanned excavator in real time.

[0087] The technical solution of the embodiment of the present invention can be applied to the unmanned crushing operation of unmanned excavators, such as unmanned crushing operation at a fixed material port, or unmanned crushing operation in a fixed area autonomously. End-to-end detection of unmanned crushing operation points based on the frontier convolutional network and transformer network in the field of artificial intelligence can be performed without using any artificial design rules, with strong anthropomorphism and pure data drive, and can adapt to the detection of unmanned crushing operation points in various weather and various extreme environments; the end-to-end network architecture design is adopted, which is purely data driven, and the true value of the data annotated by artificial experts can give the network a high degree of anthropomorphism, so that the network detection results have extremely high rationality, thereby making the success rate of the operation action high and the effectiveness high; the end-to-end network architecture design of the fused image point cloud is adopted, in the case of strong light, backlight, weak light, and dirty camera lens, even if the outline of the target object such as stones in the image is not clear, the final output is the decoding result based on the point cloud token and the vehicle status token, which utilizes the high robustness of the point cloud, thereby making the entire algorithm detection process have higher robustness, which can ensure that the detection results are reasonable and maintain the effectiveness of the detection. It adopts an end-to-end network architecture design and is purely data-driven. The more data there is, the stronger the model performance will be, so it can adapt to various dynamic and complex operating environments.

[0088] Embodiment 3

[0089] Figure 4 Schematic diagram of the structure of an end-to-end crushing operation point determination device provided according to the third embodiment of the present invention. Figure 4 As shown, the device includes: an image segmentation module 410, a point cloud extraction module 420, a point cloud token determination module 430, a vehicle state token determination module 440 and a crushing operation point determination module 450. Among them:

[0090] The image segmentation module 410 is used to obtain the image to be crushed in the unmanned crushing operation area through a camera, and to segment the image to be crushed to obtain a segmentation result of the target object;

[0091] The point cloud extraction module 420 is used to obtain a laser radar point cloud through a laser radar, and extract a target object point cloud corresponding to a segmentation result of the target object in the laser radar point cloud;

[0092] The point cloud token determination module 430 is used to extract point cloud features from the target object point cloud to obtain a plurality of point cloud tokens;

[0093] The vehicle status token determination module 440 is used to obtain the vehicle information of the excavator and perform feature extraction on the vehicle information to obtain the vehicle status token;

[0094] The crushing operation point determination module 450 is used to input the point cloud token and the vehicle status token into the decoder, and decode to obtain the crushing operation point position and the crushing operation angle prompt value.

[0095] Optionally, the image segmentation module 410 includes:

[0096] The image segmentation unit is used to segment the image to be crushed through the image segmentation network to obtain the segmentation result of at least one of the following target objects in the image: each stone to be crushed, the stone to be moved, and the grid hole stuck material area.

[0097] Optionally, the point cloud extraction module 420 includes:

[0098] A joint calibration result acquisition unit, used to obtain the joint calibration result of the camera and the laser radar;

[0099] The point cloud extraction unit is used to extract the target object point cloud corresponding to the segmentation result of the target object in the laser radar point cloud according to the joint calibration result.

[0100] Optionally, the vehicle status token determination module 440 includes:

[0101] A vehicle data acquisition unit, used to acquire vehicle positioning data of the excavator and joint angle position data of the excavator boom;

[0102] The vehicle state token determination unit is used to extract features and tokenize the vehicle positioning data and joint angle position data through a multi-layer perceptron network to obtain a vehicle state token.

[0103] Optionally, the crushing operation point determination module 450 includes:

[0104] An embedding vector determination unit, used for inputting the point cloud token and the vehicle status token into a decoder, and decoding to obtain an embedding vector of a crushing operation point position and a crushing operation angle prompt;

[0105] The crushing operation point determination unit is used to input the embedded vector into the multi-layer perceptron detection head to obtain the crushing operation point position and the crushing operation angle prompt value.

[0106] Optionally, the device further includes:

[0107] A crushing operation trajectory determination module is used to plan the crushing operation trajectory of the excavator according to the crushing operation point position and the crushing operation angle prompt value after obtaining the crushing operation point position and the crushing operation angle prompt value to obtain the crushing operation trajectory;

[0108] The crushing operation control module is used to control the telescopic boom arm of the excavator to perform crushing operations according to the crushing operation trajectory.

[0109] The end-to-end crushing operation point determination device provided in the embodiment of the present invention can execute the end-to-end crushing operation point determination method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0110] Embodiment 4

[0111] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0112] like Figure 5 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0113] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0114] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an end-to-end crushing operation point determination method.

[0115] In some embodiments, the end-to-end crushing operation point determination method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the end-to-end crushing operation point determination method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the end-to-end crushing operation point determination method in any other appropriate manner (e.g., by means of firmware).

[0116] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0117] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0118] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0119] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0120] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0121] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0122] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0123] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An end-to-end method for determining a crushing operation point, characterized in that: include: Acquire the image to be crushed in the unmanned crushing operation area through the camera, and segment the image to be crushed to obtain the segmentation result of the target object; Acquire a laser radar point cloud through a laser radar, and extract a target object point cloud corresponding to the segmentation result of the target object from the laser radar point cloud; Extracting point cloud features from the target object point cloud to obtain a plurality of point cloud tokens; Acquire vehicle information of the excavator, and perform feature extraction on the vehicle information to obtain a vehicle status token; The point cloud token and the vehicle status token are input into a decoder, and the position of the crushing operation point and the crushing operation angle prompt value are obtained by decoding.

2. The method according to claim 1, characterized in that Segmenting the image to be fragmented to obtain a segmentation result of the target object includes: The image to be crushed is segmented through an image segmentation network to obtain a segmentation result of at least one of the following target objects in the image: each stone to be crushed, the stone to be moved, and the grid hole stuck material area.

3. The method according to claim 1, characterized in that Extracting a target object point cloud corresponding to the segmentation result of the target object from the laser radar point cloud includes: Get the joint calibration results of the camera and lidar; According to the joint calibration result, a target object point cloud corresponding to the segmentation result of the target object is extracted from the laser radar point cloud.

4. The method according to claim 1, characterized in that: Obtain the vehicle information of the excavator, and perform feature extraction on the vehicle information to obtain a vehicle status token, including: Obtaining vehicle positioning data of the excavator and joint angle position data of the excavator boom; The vehicle positioning data and the joint angle position data are subjected to feature extraction and tokenization through a multi-layer perceptron network to obtain a vehicle status token.

5. The method according to claim 1, characterized in that The point cloud token and the vehicle status token are input into a decoder, and the position of the crushing operation point and the crushing operation angle prompt value are obtained by decoding, including: Input the point cloud token and the vehicle status token into a decoder, and decode to obtain an embedded vector of a crushing operation point position and a crushing operation angle prompt; The embedding vector is input into the multi-layer perceptron detection head to obtain the crushing operation point position and crushing operation angle prompt value.

6. The method according to claim 1, characterized in that After obtaining the crushing operation point position and crushing operation angle prompt value, it also includes: According to the crushing operation point position and the crushing operation angle prompt value, the crushing operation trajectory of the excavator is planned to obtain the crushing operation trajectory; The telescopic boom arm of the excavator is controlled to perform the crushing operation according to the crushing operation trajectory.

7. An end-to-end crushing operation point determination device, characterized in that: include: An image segmentation module is used to obtain an image to be crushed in an unmanned crushing operation area through a camera, and to segment the image to be crushed to obtain a segmentation result of a target object; A point cloud extraction module, used for acquiring a laser radar point cloud through a laser radar, and extracting a target object point cloud corresponding to a segmentation result of the target object from the laser radar point cloud; A point cloud token determination module is used to extract point cloud features from the target object point cloud to obtain a plurality of point cloud tokens; A vehicle status token determination module is used to obtain vehicle information of the excavator and perform feature extraction on the vehicle information to obtain a vehicle status token; The crushing operation point determination module is used to input the point cloud token and the vehicle status token into a decoder, and decode to obtain the crushing operation point position and the crushing operation angle prompt value.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the end-to-end crushing operation point determination method described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the end-to-end crushing operation point determination method according to any one of claims 1 to 6 when executed.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the end-to-end crushing operation point determination method according to any one of claims 1 to 6.