A port timber intelligent counting method and device based on depth estimation
By combining monocular depth estimation and target detection models, and using depth information to assist target detection, the problem of low accuracy in counting timber inside the grab bucket was solved, resulting in more accurate timber counting and improved operational efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-04-10
AI Technical Summary
Existing deep learning-based target detection methods have difficulty distinguishing between timber inside the grab bucket and surrounding interfering timber in port timber counting, resulting in low counting accuracy and affecting operational efficiency.
By combining a monocular depth estimation model with a target detection model, depth information is used to assist target detection, and interference objects not inside the grasping structure are eliminated. The average depth is determined using depth values, and targets with deviations greater than a threshold are eliminated for accurate counting.
It improved the accuracy and automation of timber counting at the port, and increased operational efficiency.
Smart Images

Figure CN120997649B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification belong to the field of image object recognition, and particularly relate to a port timber intelligent counting method and device based on depth estimation. BACKGROUND
[0002] In the port timber tallying operation, accurately counting the number of timbers in the grab bucket of a portal crane or a forklift is the key to improving operation efficiency and reducing human error. However, the existing target detection method based on deep learning (such as YOLO11) has difficulty in accurately distinguishing the timbers in the grab bucket from the surrounding stacked interference timbers when facing the complex environment of the port terminal operation site, because there are other timbers in front and back of the grab bucket, and the timbers inside the grab bucket and the surrounding stacked timbers have similar features in the RGB image. This leads to low counting accuracy of the timbers in the grab bucket and affects the operation efficiency. SUMMARY
[0003] Embodiments of the present disclosure provide a port timber intelligent counting method and device based on depth estimation, aiming to solve one or more of the above problems and other potential problems.
[0004] According to a first aspect of the present disclosure, a port timber intelligent counting method based on depth estimation is provided. The method comprises: in response to a grabbing operation completion instruction, acquiring an original image of a target scene; determining a depth map corresponding to the original image based on a monocular depth estimation model, so as to splice the original image and the depth map into four-channel data; based on the four-channel data, outputting a recognition result by a trained target detection model, wherein the input channel of the first layer convolution of the target detection model is four, and all channels are read when the data is loaded, and the recognition result is labeled with a target box corresponding to a grabbing structure and a grabbing target in the target box; determining the average depth of the grabbing target in the current grabbing operation based on the depth value of each grabbing target, and removing the grabbing target whose deviation proportion from the average depth is greater than a proportion threshold, and then counting the remaining grabbing targets.
[0005] According to a second aspect of the present disclosure, a port timber intelligent counting device based on depth estimation is provided. The device comprises an image stitching module configured to acquire an original image of a target scene in response to a grabbing operation completion instruction, determine a depth map corresponding to the original image based on a monocular depth estimation model, and stitch the original image and the depth map into four-channel data; a model prediction module configured to output a recognition result based on the four-channel data from a trained target detection model, the input channel of the first layer convolution of the target detection model being four, and all channels being read when data is loaded, the recognition result being labeled with a target box corresponding to a grabbing structure and a grabbing target in the target box; and a target counting module configured to determine an average depth of the grabbing target in the current grabbing operation based on the depth value of each grabbing target, eliminate the grabbing target whose deviation from the average depth is greater than a proportion threshold, and count the remaining grabbing targets.
[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising one or more processors, and a memory associated with the one or more processors, the memory being configured to store program instructions, the program instructions being configured to perform the method provided by the first aspect when executed by the one or more processors.
[0007] According to a fourth aspect of the present disclosure, a computer program product is provided, comprising a computer program configured to implement the method provided by the first aspect when executed by a processor.
[0008] The method provided by the embodiments of the present disclosure can introduce depth information obtained by a monocular depth estimation model, embed the depth information as a fourth waveband into a target detection model, assist in enhancing the spatial positioning capability of the target detection model, better eliminate interference objects outside the grabbing structure, and further more accurately identify and count the number of target objects inside the grabbing structure, thereby improving the automation level and work efficiency in this scene. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by describing in detail the embodiments thereof with reference to the attached drawings. In the drawings:
[0010] Figure 1 A flowchart of a port timber intelligent counting method based on depth estimation of some embodiments of the present disclosure is shown;
[0011] Figure 2 A schematic diagram of the principle process of a target detection model in a training stage of some embodiments of the present disclosure is shown;
[0012] Figure 3 A structural schematic diagram of a port timber intelligent counting device based on depth estimation of some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic block diagram of an electronic device showing some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0014] For the purpose of clarity, technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings and embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0015] The terms "comprise", "comprising", "include", "including", "have", "having" and any variations thereof in the specification and claims and above mentioned attached drawings, are intended to cover a non-exclusive inclusion. For example, a process, method, system, product or apparatus that includes a list of steps or units is not limited to the listed steps or units, but can optionally further include other steps or units not listed or can optionally further include other steps or units inherent to such process, method, product or apparatus. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting".
[0016] Figure 1 A flowchart of a deep estimation-based port wood intelligent counting method 100 of some embodiments of the present disclosure is shown. The method 100 may, for example, be executed by a terminal, which can include but is not limited to a mobile phone, a tablet computer, a desktop computer, a server, etc. As shown in Figure 1 At block 102, the method 100 can acquire an original image of a target scene in response to a grabbing operation completion instruction, determine a depth map corresponding to the original image based on a monocular depth estimation model, and splice the original image and the depth map into four-channel data.
[0017] In this embodiment, the target scene may, for example, be a scene such as port terminal operation with a complex environment, and the scene includes a grabbing structure capable of grabbing a grabbing target. The grabbing structure may be a grab bucket of a portal crane, a grab bucket of a forklift, or the like, and the grabbing target may be a cylindrical material such as wood or steel with a cross section. After the grabbing structure completes a grabbing operation, the grabbing structure can send a grabbing operation completion instruction to a terminal to inform the terminal that the grabbing is completed and image acquisition can be performed. After receiving the grabbing operation completion instruction, the terminal can perform image acquisition on the target scene through a camera arranged in the target scene to obtain an original image. In order to facilitate subsequent better recognition, the camera can be arranged to capture an image corresponding to the cross section of the grabbing target in a certain direction, so as to facilitate counting according to the number of cross sections.
[0018] In addition, the original image collected can also be processed according to the monocular depth estimation model to obtain a depth map. The monocular depth estimation model may, for example, be Depth Anything V2, which is a monocular depth estimation model based on deep learning. Depth Anything V2 adopts large-scale unsupervised pre-training and an efficient feature extraction architecture, and can predict the depth information of a scene from a single RGB image. The model usually adopts multi-scale feature fusion and attention mechanisms to improve the accuracy of depth estimation, and is suitable for 3D perception tasks in complex scenes. Since the monocular depth estimation model used in the present application has not been adjusted, the existing trained model can be directly used, or the model can be trained according to the conventional model training process.
[0019] After obtaining the depth map corresponding to the original image, the original image and the depth map can be spliced to obtain four-channel data. The splicing method may be to normalize the depth map to 0-255, save it as the fourth band of the original image, and obtain four-channel data in the form of RGB+D.
[0020] At block 104, the method 100 can output a recognition result based on the four-channel data by using a trained target detection model. The input channel of the first layer convolution of the target detection model is four, and all channels are read when the data is loaded. The recognition result is marked with a target box corresponding to the grabbing structure and the grabbing target in the target box.
[0021] In this embodiment, the target detection model can be YOLO11, and the network of the model will be modified. The input channel of the first convolution kernel of the Backbone of YOLO11 is modified to 4 to support the processing of four-band data, so that the trained target detection model can process four-channel data to output the recognition result. The recognition result will label the target box of the area where the grabbing structure is located and label each grabbing target in the target box based on the four-channel data. Generally, in order to ensure the recognition effect, both sides of the grabbing claw in the image should be visible, not just one side, so as to avoid some grabbing targets being blocked by other grabbing targets in the image. The camera generally continuously collects multiple images for processing, and the target detection model can only process images in which both sides of the grabbing claw are visible.
[0022] At block 106, the method 100 can determine the average depth of the grabbing targets in the current grabbing operation based on the depth values of the grabbing targets, and count the remaining grabbing targets after excluding the grabbing targets whose deviation from the average depth is greater than the proportion threshold.
[0023] In this embodiment, due to the complexity of the target scene, other interference targets may be stored around the grabbing structure, so that the target box identified will also include a small part of the interference targets in the vicinity. Therefore, after determining the target box and filtering the targets outside the target box, further filtering of the grabbing targets in the target box is still needed. Specifically, the depth values of the grabbing targets that are actually grabbed should be close, and the interference targets in the surrounding environment will have a relatively obvious difference in depth values because they are not actually in the same position, and most of the grabbing targets in the target box should be the grabbing targets that are actually grabbed. Therefore, the average depth of the current grabbing operation can be directly determined according to the depth values of the grabbing targets. As an example, the average depth can be directly calculated as the average of all depth values, or the median of the depth values can be determined as the average depth, or the depth values can be divided according to a predetermined depth range, and the middle value of the depth range with the most number of divisions can be taken as the average depth, etc. After determining the average depth, the deviation proportion between the depth value of each grabbing target and the average depth is calculated. The deviation proportion can be calculated by dividing the difference between the depth value and the average depth by the average depth. A proportion threshold (for example, 20%) is set in advance, and the grabbing targets whose deviation proportion is greater than the proportion threshold are excluded, and the remaining grabbing targets are counted. Each identified grabbing target will be labeled in the recognition result, so the number of labels remaining at this time can be counted.
[0024] In one implementation, determining the average depth of the grabbing targets in the current grabbing operation based on the depth values of the grabbing targets comprises:
[0025] Based on the depth values of the grabbing targets, the depth positions corresponding to the grabbing targets are determined respectively, and each depth position is divided based on a preset gradient value.
[0026] The target depth position with the most divided grabbing targets is determined to determine the average depth of the grabbing targets in the current grabbing operation according to the depth values in the target depth position.
[0027] In this embodiment, the depth values can be divided into depth positions (for example, 1-50, 51-100, etc.) according to a preset gradient value (for example, 50). Then, the depth position to which the depth value of each grabbing target belongs is determined. After the division of all depth values is completed, the position with the most divided grabbing targets is determined as the target depth position, and the average depth is determined according to the depth values in the target depth position. As an example, the average depth can be the average value, median value, etc. of the depth values in the position.
[0028] In an implementable manner, the average depth of the grabbing targets in the current grabbing operation is determined according to the depth values in the target depth position, including:
[0029] The average value of the depth values in the target depth position is calculated, and the average value is taken as the average depth of the grabbing targets in the current grabbing operation.
[0030] In this embodiment, the average value of the depth values divided into the target depth position is calculated, which is taken as the average depth of the grabbing targets actually grabbed this time.
[0031] In an implementable manner, the method further includes:
[0032] Based on the historical image data, a training sample is determined, and the training sample includes a four-channel data sample and an identification result sample;
[0033] Based on the four-channel data sample, a predicted identification result is generated by an initial target detection model; and
[0034] The initial target detection model is trained at least once with the identification result sample as a supervision signal to obtain a target detection model.
[0035] In this embodiment, Figure 2An illustrative diagram of a principle process 200 of a target detection model of some embodiments of the present disclosure in a training stage is shown. Historical images data can store images of a number of historical captured targets captured by a grabbing device. A depth map of the historical image data is determined by a monocular depth estimation model to construct four-channel data samples 210-1, and after the four-channel data samples 210-1 are labeled to obtain recognition result samples 210-2, training samples 210 can be obtained, and the division ratio of the training set and the validation set in the training samples 210 can be 8:2. The recognition result samples 210-2 can be labeled using the X-AnyLabeling open source tool, and the labeled categories need to include the cross section of the target and the grabbing structure. The initial target detection model used for training can be a YOLO11 model, etc. In the process of training the target detection model 220 using the training samples 210, the generator 221 of the model can generate a predicted recognition result 222 based on the four-channel data samples 210-1. According to the comparison of the recognition result samples 210-2 and the predicted recognition result 222, a generator loss 223 can be obtained, and the predicted recognition result 222 is judged by using the generator loss 223, and the loss judgment can be judged by using the comparison loss function. The generator loss 223 can be a large value, and then the generator loss 223 can be back propagated to the generator 221 to guide the optimization of the parameters of the generator 221, and a round of supervised training of the generator 221 is realized. Such a training process can be iteratively executed round by round until the generator 221 can generate more accurate predicted recognition results 222, that is, the loss value calculated by the loss function is smaller. After the training is completed, the target detection model 220 can output a recognition result 230.
[0036] In an implementation manner, based on the historical image data, the training samples are determined, including:
[0037] Based on the historical image data, an original image sample is obtained, and a corresponding depth map sample of the original image sample is determined based on a monocular depth estimation model, and both sides of the grabbing claws of the grabbing structure in the original image sample are visible;
[0038] The normalized depth map sample is saved as the fourth waveband of the original image sample to obtain four-channel data samples;
[0039] The grabbing structure in the four-channel data samples is labeled with a target box, and the cross section of the target in the target box is labeled to obtain recognition result samples; and
[0040] Based on the four-channel data samples and the recognition result samples, the training samples are constructed.
[0041] In this embodiment, since the thickness, length, etc. of the targets such as wood and steel are different, and the grabbing process cannot guarantee that the cross sections of all targets are in the same plane, if the angle of the camera is not directly opposite the grabbing structure and can clearly see both sides of the grabbing structure, it is easy to cause the target to be blocked. Therefore, the selected original image samples need to have both sides of the grabbing claws visible. In addition, in order to reduce the influence of environmental factors, the labeled data should include as many different weather conditions as possible, such as daytime, nighttime, foggy day, rainy day, backlight, etc. At the same time, it is also necessary to ensure the balance of the sample quantity of different grabbing structures such as the grab bucket of the door machine and the grab bucket of the forklift, so that the trained samples are more robust. Finally, the four-channel data samples are constructed with these original image samples and their corresponding depth map samples, and the training samples are constructed with the four-channel data samples and the labeled recognition result samples.
[0042] In an implementation manner, after determining the training samples based on the historical image data, the method further includes:
[0043] The training samples are preprocessed, and the preprocessing includes Mosaic data enhancement, rotation processing, and color disturbance.
[0044] In this embodiment, the training samples can also be preprocessed to further enrich the sample data. For example, the Mosaic data enhancement method can be used. Four pictures are spliced into one picture through random scaling, cropping, and arrangement, which can greatly enrich the detection data set, increase many small targets, and also reduce the GPU memory. The image can also be randomly rotated horizontally, vertically, and by 90 degrees. The image can also be color disturbed to change the color brightness of the image, etc.
[0045] In an implementation manner, the initial target detection model is trained based on the stochastic gradient descent method, and the Mosaic data enhancement is canceled in the last preset round of training.
[0046] In this embodiment, the random gradient descent method can be used in the training process of the model, for example, 100 rounds of training, and the result is verified on the verification set after each round of training. The first 90 rounds can start Mosaic data enhancement to enhance small target detection, and the last 10 rounds can turn off Mosaic data enhancement to stabilize model training, help the model better adjust the weight, and adapt to the real scene. In addition, automatic mixed precision (Automatic Mixed Precision, AMP) can also be used in the training process to combine the use of single-precision floating-point numbers (FP32) and half-precision floating-point numbers (FP16) in the deep learning training process to improve the training speed and reduce the memory occupation, and improve the numerical stability.
[0047] Figure 3A structural schematic diagram of the port wood intelligent counting device 300 based on depth estimation of some embodiments of the present disclosure is shown. Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment focuses on the difference from other embodiments. Especially, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant part can be referred to the part of the method embodiment. As shown in Figure 3 The device 300 includes an image stitching module 301 configured to, in response to a grabbing operation completion instruction, acquire an original image of a target scene, determine a depth map corresponding to the original image based on a monocular depth estimation model, and stitch the original image and the depth map into four-channel data; a model prediction module 302 configured to, based on the four-channel data, output a recognition result from a trained target detection model, the input channel of the first layer convolution of the target detection model being four, and all channels being read when loading data, and the recognition result being labeled with a target box corresponding to a grabbing structure and a grabbing target in the target box; and a target counting module 303 configured to determine an average depth of the grabbing target in the current grabbing operation based on the depth values of each grabbing target, and count the remaining grabbing targets after excluding the grabbing targets with a deviation proportion from the average depth greater than a proportion threshold.
[0048] In an implementable manner, the target counting module 303 is further configured to determine a depth range corresponding to each grabbing target based on the depth value of the grabbing target, each depth range being divided based on a preset gradient value; and determine a target depth range with the most divided grabbing targets, so as to determine the average depth of the grabbing target in the current grabbing operation according to the depth values in the target depth range.
[0049] In an implementable manner, the target counting module 303 is further configured to calculate the average value of the depth values in the target depth range, and take the average value as the average depth of the grabbing target in the current grabbing operation.
[0050] In an implementable manner, the device further includes a model training module configured to determine training samples based on historical image data, the training samples including four-channel data samples and recognition result samples; generate a predicted recognition result from an initial target detection model based on the four-channel data samples; and perform at least one round of model training on the initial target detection model with the recognition result samples as a supervision signal, to obtain the target detection model.
[0051] In an implementation, the model training module is further configured to obtain an original image sample based on the historical image data, and determine a corresponding depth map sample of the original image sample based on the monocular depth estimation model, both sides of the grabbing structure in the original image sample are visible; save the normalized depth map sample as a fourth wave band of the original image sample to obtain a four-channel data sample; label a target frame of the grabbing structure in the four-channel data sample, and label a cross section of the grabbing target in the target frame to obtain a recognition result sample; and construct the training sample based on the four-channel data sample and the recognition result sample.
[0052] In an implementation, the model training module is further configured to pre-process the training sample, and the pre-processing includes Mosaic data enhancement, rotation processing, and color disturbance.
[0053] In an implementation, the initial target detection model is trained based on a stochastic gradient descent method, and the Mosaic data enhancement is cancelled in the last preset round of training.
[0054] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted by a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0055] Figure 4 A block diagram of an electronic device 400 is shown, which can implement various embodiments of the present disclosure. As Figure 4As shown, the electronic device 400 includes a processor 410, a disk drive 420, an input / output interface 430, a network interface 440, and a memory 450. The processor 410, the disk drive 420, the input / output interface 430, the network interface 440, and the memory 450 can be communicatively connected through a bus 460.
[0056] The processor 410 can be implemented in a general-purpose CPU, a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute related programs to implement the technical solutions provided in the present application.
[0057] The memory 450 can be implemented in a Read Only Memory (ROM), a Random Access Memory (RAM), a static storage, a dynamic storage device, etc. The memory 450 can store an operating system 451 for controlling the operation of the electronic device 400, a Basic Input Output System (BIOS) 452 for controlling the low-level operation of the electronic device 400. In addition, a web browser 453, a data storage management system 454, etc. can also be stored. In general, when the technical solutions provided in the present application are implemented by software or firmware, the related program codes are stored in the memory 450 and are executed by the processor 410.
[0058] The input / output interface 430 is configured to connect an input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, a prompt light, etc.
[0059] The network interface 440 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0060] The bus 460 includes a path for transmitting information between various components (such as the processor 410, the disk drive 420, the input / output interface 430, the network interface 440, and the memory 450) of the device.
[0061] It should be noted that although the above device only shows the processor 410, the disk drive 420, the input / output interface 430, the network interface 440, the memory 450, the bus 460, etc., in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can only contain the components necessary to implement the method of the present application, and does not necessarily contain all the components shown in the figure.
[0062] Program code for carrying out operations of the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, so that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. Program codes can be executed entirely on a machine, partially on a machine, partially on a machine as a separate software package, and partially on a remote machine, or entirely on a remote machine or server.
[0063] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of the machine-readable storage medium will include one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In addition, although the operations are depicted in a specific order, it should be understood that the operations are required to be performed in the specific order shown or in a sequential order, or that all the illustrated operations should be performed to achieve the desired results. In certain circumstances, multi-tasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single implementation. Conversely, various features described in the context of a single implementation can also be separated and implemented in multiple implementations.
[0064] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for intelligent counting of port timber based on depth estimation, characterized in that, The method includes: In response to the completion command of the grabbing operation, the original image of the target scene is acquired, and the depth map corresponding to the original image is determined based on the monocular depth estimation model. The original image and the depth map are then stitched together to form four-channel data. The original image is an image of the orientation corresponding to the cross-section of the grabbing target. Based on the four-channel data, the trained target detection model outputs the recognition result. The first convolutional layer of the target detection model has four input channels, and all channels are read when the data is loaded. The recognition result is marked with the target box corresponding to the grasping structure and the grasping target within the target box. The average depth of each grasping target in this grasping operation is determined based on the depth value of each grasping target. After removing grasping targets whose deviation from the average depth is greater than a threshold, the remaining grasping targets are counted.
2. The intelligent counting method for port timber based on depth estimation according to claim 1, characterized in that, Determining the average depth of the grasped target in this grasping operation based on the depth values of each grasped target includes: Based on the depth value of the grasping target, a depth level corresponding to each grasping target is determined, and each depth level is divided based on a preset gradient value; and The target depth level with the most grasped targets is determined, and the average depth of the grasped targets in this grasping operation is determined based on the depth values within the target depth level.
3. The intelligent counting method for port timber based on depth estimation according to claim 2, characterized in that, Determining the average depth of the target object in this grasping operation based on the depth values within the target depth range includes: Calculate the average value of each depth value within the target depth range, and use the average value as the average depth of the grasping target in this grasping operation.
4. The intelligent counting method for port timber based on depth estimation according to claim 1, characterized in that, The method further includes: Based on historical image data, training samples are determined, which include four-channel data samples and recognition result samples. Based on the four-channel data samples, the initial target detection model generates predicted recognition results; and Using the identified result samples as supervision signals, the initial target detection model is trained for at least one round to obtain the target detection model.
5. The intelligent counting method for port timber based on depth estimation according to claim 4, characterized in that, The process of determining training samples based on historical image data includes: Based on historical image data, original image samples are obtained, and the corresponding depth map samples of the original image samples are determined based on a monocular depth estimation model. Both sides of the grasping claws of the grasping structure are visible in the original image samples. The normalized depth map sample is saved as the fourth band of the original image sample to obtain a four-channel data sample. The target bounding boxes of the grasping structures in the four-channel data samples are labeled, and the cross-sections of the grasping targets within the target bounding boxes are labeled to obtain recognition result samples; and Training samples are constructed based on the four-channel data samples and the recognition result samples.
6. The intelligent counting method for port timber based on depth estimation according to claim 4, characterized in that, After determining the training samples based on historical image data, the process also includes: The training samples are preprocessed, including Mosaic data augmentation, rotation processing, and color perturbation.
7. The intelligent counting method for port timber based on depth estimation according to claim 4, characterized in that, The initial target detection model is trained based on stochastic gradient descent, and Mosaic data augmentation is removed in the final preset training rounds.
8. A port timber intelligent counting device based on depth estimation, characterized in that, The device includes: The image stitching module is configured to respond to the capture operation completion command, acquire the original image of the target scene, determine the depth map corresponding to the original image based on the monocular depth estimation model, and stitch the original image and the depth map into four-channel data. The original image is the image of the cross-section of the captured target corresponding to the orientation. The model prediction module is configured to output recognition results from the trained object detection model based on the four-channel data. The first convolutional layer of the object detection model has four input channels, and all channels are read when the data is loaded. The recognition results are marked with the target box corresponding to the grasping structure and the grasping target within the target box. The target counting module is configured to determine the average depth of the target in the current grasping operation based on the depth value of each target, and to remove targets whose deviation from the average depth is greater than a threshold value, and then count the remaining targets.
9. An electronic device, characterized in that, include: One or more processors, and A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the port timber intelligent counting method based on depth estimation as described in any one of claims 1-7.
10. A computer program product, characterized in that, The system includes a computer program that, when executed by a processor, implements a port timber intelligent counting method based on depth estimation according to any one of claims 1-7.
Citation Information
Patent Citations
Automatic counting method and system for silkworm cocoons
CN103246920A
Rice ear detection counting and rice yield estimation method based on unmanned aerial vehicle
CN117853961A