Green detection model applied to autonomous online training of vision separation system

By using RGBD cameras and multi-threaded parallel computing in the visual single-piece separation system, an autonomous online training green detection model is realized, solving the problem of insufficient recognition capabilities of new package types and high demand for hardware computing power, improving the recognition accuracy and efficiency, and meeting the requirements of low-carbon and environmental protection.

CN120107549APending Publication Date: 2025-06-06JIN HOUNG FUH (CHUZHOU) CONVEYING EQUIP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510169262.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

Smart Images

  • Figure CN120107549A_ABST
    Figure CN120107549A_ABST
Patent Text Reader

Abstract

The invention discloses a green detection model applied to autonomous online training of a visual separation system, and relates to the technical field of visual separation, and the model comprises the steps: a deployment module collects an RGB image and a depth image through deploying an RGBD camera; the calculation module calculates and obtains the number of frame target packages of the RGB image and the depth image; when the numbers are inconsistent, the comparison module records the images as abnormal images, and compares and outputs redundant target coordinates after matching; the uploading module uploads the abnormal RGB image and the coordinates to a cloud for storage; the labeling module automatically generates a labeling result when the cloud new samples are accumulated to a preset number; the training module regularly inputs the number of the new annotation files into a data input interface of the reinforcement model for retraining, and the cloud sends out a notification after training is completed; according to the application module, a visual separation terminal system automatically downloads a model file to complete full-automatic unmanned model updating. Automatic online training and self-adaptive optimization are achieved, and the recognition precision, the processing speed and the stability of the visual separation system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vision separation technology, and in particular to a green detection model applied to autonomous online training of a vision separation system. Background Art

[0002] Most of the current visual single-piece separation systems use a package recognition system based on a convolutional neural network. The pre-work of this method is to collect various package samples as comprehensively as possible to create a data set. Since there are many types of packages in the logistics industry, there are all kinds of sizes, outer packaging, and types, and the amount of data in this data set is very large. Taking the model involved in this invention as an example, the labeled sample data set exceeds 40,000. The data set includes 15 categories, including packages of various sizes, types, and various situations commonly seen in different scenarios.

[0003] Sample collection is a time-consuming and labor-intensive process. At present, due to the increasingly broad application prospects of AI, how to automatically collect sample annotation data sets is one of the key research directions. Taking the data set mentioned in the present invention as an example, the production process of the data set took up to 2 years, and the collection was as comprehensive as possible, but there are still problems. Since the system is not only used in the express logistics industry, but also in various warehousing and logistics industries such as medicine and food, new types of identification targets will still be encountered. Although the model has a certain learning and classification ability and can effectively identify targets outside the sample, there are still unexpected situations that cause the model to be unable to correctly identify, affecting the accuracy of the system. Only when the system is abnormal, the sample data set standards of the new variety are manually collected and the training is strengthened, and the model is replaced to update the system. This process is unacceptable to end users.

[0004] In addition, deep learning methods rely on the computing power of GPUs, and on the capabilities of hardware, especially graphics cards. In addition, with the endless increase in samples, the amount of model data is getting larger and larger, which affects the computing efficiency and requires higher and higher computing power and energy consumption, which is not in line with the development trend of low-carbon environmental protection. Therefore, it is urgent to design a new green and environmentally friendly model that can autonomously increase useful samples while forgetting samples with extremely low usage rates. This model can not only be used in visual separation systems, but also the green model that is continuously improved by the separation system can be used in other intelligent systems based on visual recognition. Summary of the invention

[0005] Based on the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a green detection model for autonomous online training of a visual separation system to solve the above-mentioned technical problems.

[0006] To achieve the above object, the present invention provides the following technical solution: a green detection model for autonomous online training of a visual separation system, comprising:

[0007] Deployment module: collect RGB images and depth images by deploying RGBD cameras, and the RGB images and depth images are located in the same coordinate system;

[0008] Computing module: Multi-threaded parallel computing is used to obtain the number of target packages N in each frame of the RGB image through a deep learning algorithm. 1 , calculate the number of target packages N in the corresponding depth map of the same frame through the height detection algorithm 2 ;

[0009] Comparison module: When N 1 =N 2 When N 1 <N 2 When , the RGB image and the depth image are abnormal images, save the two images, compare the coordinates of each target package in the same coordinate system, and output the redundant target coordinates after matching (X e ,Y e );

[0010] Upload module: The RGB image in the abnormal image and the abnormal coordinates (X e ,Y e ) is bound and uploaded to the cloud and saved as new sample information classification;

[0011] Labeling module: Through the automatic labeling tool based on X-AnyLabeling, Cloud-Auto-Labeling is designed to retrieve cloud database information at regular intervals. When the number of new samples in the cloud reaches the preset number, the system automatically generates labeling results using the true value, RGB information and coordinate information;

[0012] Training module: New annotated samples are automatically matched and classified into new types of package samples according to the characteristics of the existing model data set. The cloud database regularly inputs the reinforcement model data input interface for retraining according to the number of new annotated files, and clears the samples with forgetting call times less than the preset number. After the training is completed, the cloud sends a notification to each visual separation subsystem;

[0013] Application module: By adopting cloud contribution technology, the visual separation terminal system automatically downloads the model file in the BoxDetectionmodel.onne format from the cloud to the specified model folder to complete the model update in a fully automatic and unmanned manner.

[0014] The present invention is further configured such that the RGBD camera is deployed after the RBG camera, and the RBG camera realizes preliminary separation of the front-end crowded packages through a deep learning model.

[0015] The present invention is further configured to obtain the target package number N of each frame of the RGB image through a deep learning algorithm 1,include:

[0016] Preprocess the frame image of the RGB image;

[0017] Use the trained deep learning algorithm to perform target detection on the preprocessed frame image. The target detection decomposes the frame image into multiple regions and determines whether each region contains the target package. The output is the bounding box and confidence of each detected target.

[0018] Through the output of target detection, the number of valid target packages in each frame is counted, and the number of target packages is obtained by counting the number of bounding boxes with confidence greater than a preset threshold;

[0019] By comparing the center coordinates of each target in the target bounding box, it is verified whether the target is accurately positioned in the separation area. When the target position deviation is detected to be greater than the preset threshold, it is marked.

[0020] The present invention is further configured such that the center coordinates of each target in the target bounding box are: Among them, (x center (i),y center (i)) is the center coordinate of the i-th target, x min (i) and x max (i) is the minimum and maximum value of the horizontal coordinate of the i-th target bounding box, y min (i) and y max (i) is the minimum and maximum value of the ordinate of the i-th target bounding box;

[0021] The calculation logic of the target position deviation is: Among them, Δ target (i) is the target position deviation of the i-th target bounding box, x expected (i) and y expected (i) are the expected horizontal and vertical coordinates of the target bounding box in the separated region.

[0022] The present invention is further configured to calculate the number of target packages N in the corresponding depth map of the same frame by using a height detection algorithm. 2 ,include:

[0023] The depth data D(x,y) in the depth map is obtained by using a height detection algorithm, where D(x,y) represents the depth value of the position (x,y) in the image;

[0024] According to the set minimum depth threshold and maximum depth threshold, pixels with depth values ​​outside this range are removed to obtain the valid area;

[0025] Based on the extracted effective area, the target package area is identified and segmented;

[0026] Count the number of target packages in the corresponding depth map of the same frame.

[0027] The present invention is further configured such that the extraction logic of the effective area is: Among them, D valid (x,y) is the effective area, D min is the minimum depth threshold, D max is the maximum depth threshold;

[0028] Perform threshold segmentation on the effective area and set the height threshold H threshold , the effective area is divided into background and target packages, Among them, R target (x, y) is the effective area after segmentation, and the number of target packages in the current frame image is counted.

[0029] The present invention is further configured such that when N 1 <N 2 When , the RGB image and the depth image are abnormal images, save the two images, compare the coordinates of each target package in the same coordinate system, and output the redundant target coordinates after matching (X e ,Y e ),include:

[0030] Extract the center coordinates of each target package from the RGB image and depth image;

[0031] Calculate the distance metric between the center coordinates of the target package in any RGB image and the center coordinates of the target package in any depth image. When the distance metric is less than the set matching threshold, match the corresponding target packages and output all coordinates in the depth image that fail to match the target package in the RGB image as redundant target coordinates.

[0032] The present invention is further configured to mark the center coordinates of the target package extracted from the RGB image as Where i∈{1,2,...,N 1}, the center coordinates of the target package extracted from the depth map are marked as Where j∈{1,2,...,N 2};

[0033] The calculation logic of the distance metric is: Among them, D ij It is the distance metric between the i-th target package extracted from the RGB image and the j-th target package extracted from the depth image.

[0034] The present invention is further configured that the automatic labeling tool based on X-AnyLabeling automatically generates labeling results using the true value and RGB information and the coordinate information system, including:

[0035] Read the RGB image to be annotated and the abnormal coordinates (X e ,Y e );

[0036] Generate the label of the target package according to the set labeling rules, compare the true value with the generated label, and automatically generate the labeling result when they are consistent.

[0037] The present invention is further configured such that the set labeling rules include pre-trained deep learning models or rule matching.

[0038] The present invention provides a green detection model for autonomous online training of a visual separation system, including a deployment module: an RGBD camera is deployed to collect an RGB image and a depth image, wherein the RGB image and the depth image are located in the same coordinate system; a calculation module: multi-threaded parallel calculation is adopted to obtain the number N of target packages in each frame of the RGB image through a deep learning algorithm 1 , calculate the number of target packages N in the corresponding depth map of the same frame through the height detection algorithm 2 ; Comparison module: When N 1 =N 2 When N 1 <N 2 When , the RGB image and the depth image are abnormal images, save the two images, compare the coordinates of each target package in the same coordinate system, and output the redundant target coordinates after matching (X e ,Y e ); Upload module: The RGB image in the abnormal image and the abnormal coordinates (X e ,Y e) is bound and uploaded to the cloud, and saved as a new sample information classification; Labeling module: Through the automatic labeling tool based on X-AnyLabeling, Cloud-Auto-Labeling is designed to periodically retrieve cloud database information. When the cloud new samples accumulate to the preset number, the true value and RGB information and coordinate information system automatically generate labeling results; Training module: New labeled samples automatically match and classify new package sample types according to the characteristics of the existing model data set. The cloud database regularly inputs the enhanced model data input interface for retraining according to the number of new labeled files, and cleans up the samples with forgotten call times less than the preset number. After the training is completed, the cloud sends a notification to each visual separation subsystem; Application module: By adopting the cloud contribution technology, the visual separation terminal system automatically downloads the model file in the cloud BoxDetectionmodel.onne format to the specified model folder to complete the model update under full automatic unmanned operation, and the beneficial effects include:

[0039] 1. Efficient package detection and classification capabilities: Combining the RGB image and depth image collected by the RGBD camera, and performing multi-threaded parallel calculations through deep learning algorithms, the recognition accuracy and processing speed of the target package are effectively improved. The number of target packages under the depth image is calculated by the height detection algorithm, and compared with the RGB image, which can accurately determine the location of the target package, reduce false detection and missed detection, and improve detection accuracy;

[0040] 2. Autonomous online training and adaptive capabilities: Using cloud databases and automatic labeling tools, new samples are automatically labeled through X-AnyLabeling. Combining real label information with RGB image data, the labeling results can be automatically generated in the cloud and the new samples can be classified and saved. Whenever the number of new samples reaches the preset number, retraining is automatically performed to continuously optimize the model. This process does not require human intervention, greatly improving the adaptive ability of the visual separation system, and can dynamically adapt to new environments and new package types;

[0041] 3. Accurate data comparison and anomaly detection: When the number of target packages in the RGB image and the depth image is inconsistent, it is automatically judged as an abnormal image, and the RGB image and the depth image are saved for subsequent analysis. By comparing the target package coordinates in the same coordinate system, it can output the redundant target coordinates after matching, effectively identify and process potential errors or abnormal data, and improve the stability and reliability of the system.

[0042] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0044] Figure 1 A schematic diagram of an application scenario of a green detection model applied to autonomous online training of a visual separation system is shown as an exemplary embodiment of the present invention;

[0045] Figure 2 A structural diagram of a green detection model applied to autonomous online training of a visual separation system is shown as an exemplary embodiment of the present invention;

[0046] Figure 3 The flowchart of a green detection model applied to autonomous online training of a visual separation system is shown as an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.

[0048] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0049] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0050] The visual single-piece separation feature is to disperse dense packages into separate packages with spacing, such as Figure 1As shown, the blue area is the area where the trip has been separated. The present invention uses this system feature for design. There is no need to deploy one in each scene. It only needs to be deployed in a certain scene of different categories, and cloud-based shared synchronization is adopted. The RBG camera is used in the front-end crowded preliminary separation area to meet the package detection under the deep learning model. In the blue separation area, an RGBD camera is deployed, which can collect both RGB and depth images. b. Since the same RGBD camera collects two images at the same time, there is no data alignment problem. The image of the same frame can be collected at the same time, and the target position is also in the same coordinate system.

[0051] A green detection model for autonomous online training of visual separation systems, such as Figure 2 As shown, including:

[0052] Deployment module: collect RGB images and depth images by deploying RGBD cameras, and the RGB images and depth images are located in the same coordinate system;

[0053] Computing module: Multi-threaded parallel computing is used to obtain the number of target packages N in each frame of the RGB image through a deep learning algorithm. 1 , calculate the number of target packages N in the corresponding depth map of the same frame through the height detection algorithm 2 ;

[0054] Comparison module: When N 1 =N 2 When N 1 <N 2 When , the RGB image and the depth image are abnormal images, save the two images, compare the coordinates of each target package in the same coordinate system, and output the redundant target coordinates after matching (X e ,Y e );

[0055] Upload module: The RGB image in the abnormal image and the abnormal coordinates (X e ,Y e ) is bound and uploaded to the cloud and saved as new sample information classification;

[0056] Labeling module: Through the automatic labeling tool based on X-AnyLabeling, Cloud-Auto-Labeling is designed to retrieve cloud database information at regular intervals. When the number of new samples in the cloud reaches the preset number, the system automatically generates labeling results using the true value, RGB information and coordinate information;

[0057] Training module: New annotated samples are automatically matched and classified into new types of package samples according to the characteristics of the existing model data set. The cloud database regularly inputs the reinforcement model data input interface for retraining according to the number of new annotated files, and clears the samples with forgetting call times less than the preset number. After the training is completed, the cloud sends a notification to each visual separation subsystem;

[0058] Application module: By adopting cloud contribution technology, the visual separation terminal system automatically downloads the model file in the BoxDetectionmodel.onne format from the cloud to the specified model folder to complete the model update in a fully automatic and unmanned manner.

[0059] The present invention is further configured such that the RGBD camera is deployed after the RBG camera, and the RBG camera realizes preliminary separation of the front-end crowded packages through a deep learning model. Specifically, the RBG camera realizes preliminary separation of the front-end crowded packages by acquiring RGB images and processing the images using a deep learning model. This preliminary separation refers to using a deep learning model to perform target detection in the foreground or crowded area of ​​the image, identify possible packages, and perform certain separation.

[0060] The present invention is further configured to obtain the target package number N of each frame of the RGB image through a deep learning algorithm 1 ,include:

[0061] Preprocess the frame image of the RGB image; Preprocess the RGB image to prepare the image data so that it is suitable for the input of the deep learning model;

[0062] Use the trained deep learning algorithm to perform target detection on the preprocessed frame image, where target detection decomposes the frame image into multiple regions, and determines whether each region contains the target package, and outputs the bounding box and confidence of each detected target; specifically, the deep learning algorithm processes the preprocessed RGB image, including image decomposition, dividing the image into multiple regions to analyze whether each region contains the target package; region division and target detection, through the forward propagation of the deep learning model, the algorithm will detect the target in the image and generate a bounding box for each target. The bounding box is a rectangular area that indicates the boundary position of the target, and the model will assign a confidence to each target, that is, the possibility that the target belongs to the package category;

[0063] Through the output of target detection, the number of valid target packages in each frame is counted, and the number of target packages is obtained by counting the number of bounding boxes with confidence greater than a preset threshold. Specifically, each bounding box output by target detection has a confidence value. Usually a preset threshold is set, including 0.5, and only bounding boxes with confidence greater than this threshold are considered valid target packages. By counting the number of all valid bounding boxes, the number of target packages in the current frame image is obtained. This number is the output result of the target detection algorithm, which indicates the number of target packages detected in the frame image.

[0064] By comparing the center coordinates of each target in the target bounding box, it is verified whether the target is accurately positioned in the separation area, and when the target position deviation is detected to be greater than a preset threshold, it is marked. The present invention is further configured such that the center coordinates of each target in the target bounding box are: Among them, (x center (i),y center (i)) is the center coordinate of the i-th target, x min (i) and x max (i) is the minimum and maximum value of the horizontal coordinate of the i-th target bounding box, y min (i) and y max (i) is the minimum and maximum value of the ordinate of the i-th target bounding box; the calculation logic of the target position deviation is:

[0065] Among them, Δ target (i) is the target position deviation of the i-th target bounding box, x expected (i) and y expected (i) are the expected horizontal and vertical coordinates of the target bounding box in the separation area. Specifically, the above calculation logic mainly involves the bounding box calculation, center coordinate extraction and target position deviation calculation in target detection. In this context, the target detection algorithm first identifies each target and generates a bounding box for each target. Then, by calculating the center coordinates of the target bounding box and comparing the deviation between its position and the expected separation area, it is determined whether the target is accurately located in the separation area. The calculation of the target position deviation is used to determine whether the target is accurately located in the separation area. Assume that the separation area is a preset rectangular area, and the expected value of its horizontal coordinate is x expected , the expected value of the ordinate is y expected The deviation value is the difference between the center coordinates of the target bounding box and the expected position. If the deviation exceeds a preset threshold, it means that the target is not accurately located in the separation area. At this time, the target can be marked as an incorrect target or an abnormal target.

[0066] The present invention is further configured to calculate the number of target packages N in the corresponding depth map of the same frame by using a height detection algorithm. 2 ,include:

[0067] The depth data D(x,y) in the depth map is obtained through the height detection algorithm, where D(x,y) represents the depth value of the position (x,y) in the image; specifically, the depth map is an image generated by an RGBD camera, and each pixel corresponds to the depth information of an object or scene in three-dimensional space. The depth information is usually expressed in the form of grayscale values, where the size of the grayscale value represents the distance of each pixel, and the spatial layout of the object is described by the depth value of each pixel in the image. The height detection algorithm extracts the depth information in the depth map pixel by pixel, and obtains the depth value of each pixel to obtain the depth data D(x,y);

[0068] According to the set minimum depth threshold and maximum depth threshold, pixels whose depth values ​​exceed this range are removed to obtain a valid area; the present invention is further configured that the extraction logic of the valid area is: Among them, D valid (x,y) is the effective area, D min is the minimum depth threshold, D max is the maximum depth threshold; specifically, by setting the minimum depth threshold and the maximum depth threshold of the depth value, pixels beyond this range can be effectively filtered out, thereby obtaining a valid target area;

[0069] Based on the extracted effective area, the target package area is identified and segmented; the effective area is segmented by threshold, and the height threshold H is set threshold , the effective area is divided into background and target packages, Among them, R target (x, y) is the effective area after segmentation, and the number of target packages in the current frame image is counted; specifically, the height threshold H threshold It is used to distinguish the background and target package in the effective area. Specifically, the height threshold refers to the critical value of the depth value. The area below this value will be regarded as the background, and the area above this value will be regarded as the target package. The depth value of each pixel in the effective area is compared with the set threshold. If the depth value of the pixel is greater than the set height threshold, the area is considered to belong to the target package; otherwise, it is considered to belong to the background. According to the comparison result of the threshold, the effective area is divided into two parts: the target package area and the background area. The target package area is the area with a depth value greater than the threshold, and the background area is the area with a depth value less than or equal to the threshold;

[0070] Count the number of target packages in the depth map corresponding to the same frame. Specifically, after completing the threshold segmentation, it is necessary to identify the segmented target package area and count the number of target packages in the current frame image.

[0071] The present invention is further configured such that when N 1 <N 2 When , the RGB image and the depth image are abnormal images, save the two images, compare the coordinates of each target package in the same coordinate system, and output the redundant target coordinates after matching (X e ,Y e ),include:

[0072] The center coordinates of each target package are extracted from the RGB image and the depth image; the present invention is further configured to mark the center coordinates of the target package extracted from the RGB image as Where i∈{1,2,...,N 1}, the center coordinates of the target package extracted from the depth map are marked as Where j∈{1,2,...,N 2};

[0073] Calculate the distance metric between the center coordinates of the target package in any RGB image and the center coordinates of the target package in any depth image. When the distance metric is less than the set matching threshold, match the corresponding target packages and output all the coordinates in the depth image that cannot match the target package in the RGB image as redundant target coordinates. The calculation logic of the distance metric is: Among them, D ij It is the distance metric between the i-th target package extracted from the RGB image and the j-th target package extracted from the depth image. Specifically, the matching metric of the target packages in the RGB image and the depth image is calculated through the above logic. In order to determine whether the target packages in the RGB image and the depth image match, a matching threshold is defined. If the distance metric between the target packages is less than the matching threshold, the two target packages are considered to belong to the same physical package and are matched. For the target packages in the depth image that fail to match the target packages in the RGB image, they are considered to be redundant targets. The coordinates of these target packages are output.

[0074] The present invention is further configured that the automatic labeling tool based on X-AnyLabeling automatically generates labeling results using the true value and RGB information and the coordinate information system, including:

[0075] Read the RGB image to be annotated and the abnormal coordinates (X e ,Y e );

[0076] The label of the target package is generated by the set labeling rules, and the true value is compared with the generated label. When they are consistent, the labeling result is automatically generated. Specifically, the label of the target package is generated by the preset labeling rules. The labeling rules are set according to the characteristics and position coordinates of the target package, and are used to determine the type of the target package, the label content and the specific position of the label. The present invention is further configured that the set labeling rules include pre-trained deep learning models or rule matching, and common rule matching includes: the size of the target package, setting a size range, and only the area that meets the size will be marked as the target package; the position range, according to the coordinate information in the image, a position range is set to ensure that the labeling is limited to the target area; the target feature, through the color, shape, texture and other features, combined with the depth information for comprehensive judgment to ensure the accuracy of the labeling rules. Generate the target package label, and generate the target package label according to the RGB image and coordinate information according to the set labeling rules. Each label will contain the position coordinates and type of the target package, and by comparing with the true value, the true value is the manual label or the known correct label, check whether the automatically generated label is accurate, and by comparing the consistency between the automatically generated label and the true value, judge whether the labeling requirements are met. Common comparison methods include: precision, the degree of match between the generated label and the true value; recall, whether all target packages can be identified. When the automatically generated label is consistent with the true value, the system will confirm the labeling result and store or output it. This process does not require human intervention, thereby improving the efficiency and consistency of labeling. When the labels are consistent, the system considers the labeling to be correct and automatically saves it. If not, the system may issue an alarm or require manual review to ensure the quality of the labeling.

[0077] The overall flow chart of this patent is as follows Figure 3 As shown in the figure, the proposed model can effectively reduce the model's demand for hardware computing power, which meets the strategic needs of low-carbon, environmentally friendly, and green lightweight models (green model based on DeepLearning).

[0078] The detection model sets automatic annotation when the abnormal RGB image is greater than 1. The model is automatically updated when the number of annotated files is greater than 50, and the old samples are cleared and forgotten when the utilization rate is less than 0.03%. After verification by multiple sets of equipment in multiple application sites, the fully automatic system has a high initial iteration frequency. As time goes by, the iteration cycle begins to increase, the system recognition efficiency is significantly improved, and the accuracy rate has been maintained at more than 99.95% for a long time. At the same time, the model's requirements for computing power are decreasing. Compared with version 1.0 and version 3.37 with 27 iterations, the number of samples has decreased by 3.8% from 44,125 to 42,448. At the same time, the model can be directly used in a series of visual inspection-based equipment and projects such as visual package tracking and visual palletizers, with high economic benefits.

[0079] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0080] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0081] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0082] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0083] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0084] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0085] In the several embodiments provided in the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0086] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0087] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0088] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.

[0089] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A green detection model for autonomous online training of a visual separation system, characterized in that: include: Deployment module: collect RGB images and depth images by deploying RGBD cameras, and the RGB images and depth images are located in the same coordinate system; Computing module: adopts multi-threaded parallel computing, obtains the number of target packages N1 in each frame of the RGB image through the deep learning algorithm, and calculates the number of target packages N2 in the corresponding depth image of the same frame through the height detection algorithm; Comparison module: When N1 = N2, continue to perform cyclic comparison on the frame images. When N1 < N2, the RGB image and the depth image are abnormal images. Save the two images, compare the coordinates of each target package in the same coordinate system, and output the redundant target coordinates (X e , Y e ); Upload module: The RGB image in the abnormal image and the abnormal coordinates (X e ,Y e ) is bound and uploaded to the cloud and saved as new sample information classification; Labeling module: Through the automatic labeling tool based on X-AnyLabeling, Cloud-Auto-Labeling is designed to retrieve cloud database information at regular intervals. When the number of new samples in the cloud reaches the preset number, the system automatically generates labeling results using the true value, RGB information and coordinate information; Training module: New annotated samples are automatically matched and classified into new types of package samples according to the characteristics of the existing model data set. The cloud database regularly inputs the reinforcement model data input interface for retraining according to the number of new annotated files, and clears the samples with forgetting call times less than the preset number. After the training is completed, the cloud sends a notification to each visual separation subsystem; Application module: By adopting cloud contribution technology, the visual separation terminal system automatically downloads the model file in the BoxDetectionmodel.onne format from the cloud to the specified model folder to complete the model update in a fully automatic and unmanned manner.

2. A green detection model for autonomous online training of a visual separation system according to claim 1, characterized in that: The RGBD camera is deployed after the RBG camera, and the RBG camera achieves preliminary separation of the front-end crowded packages through a deep learning model.

3. The green detection model for autonomous online training of a visual separation system according to claim 1, characterized in that: The target package number N1 of each frame of the RGB image is obtained through a deep learning algorithm, including: Preprocess the frame image of the RGB image; Use the trained deep learning algorithm to perform target detection on the preprocessed frame image. The target detection decomposes the frame image into multiple regions and determines whether each region contains the target package. The output is the bounding box and confidence of each detected target. Through the output of target detection, the number of valid target packages in each frame is counted, and the number of target packages is obtained by counting the number of bounding boxes with confidence greater than a preset threshold; By comparing the center coordinates of each target in the target bounding box, it is verified whether the target is accurately positioned in the separation area. When the target position deviation is detected to be greater than the preset threshold, it is marked.

4. The green detection model for autonomous online training of a visual separation system according to claim 3, characterized in that: The center coordinates of each object in the object bounding box are: Among them, (x center (i),y center (i)) is the center coordinate of the i-th target, x min (i) and x max (i) is the minimum and maximum value of the horizontal coordinate of the i-th target bounding box, y min (i) and y max (i) is the minimum and maximum value of the ordinate of the i-th target bounding box; The calculation logic of the target position deviation is: Among them, Δ target (i) is the target position deviation of the i-th target bounding box, x expected (i) and y expected (i) are the expected horizontal and vertical coordinates of the target bounding box in the separated region.

5. The green detection model for autonomous online training of a visual separation system according to claim 1, characterized in that: The number of target packages N2 in the corresponding depth map of the same frame is calculated through the height detection algorithm, including: The depth data D(x,y) in the depth map is obtained by using a height detection algorithm, where D(x,y) represents the depth value of the position (x,y) in the image; According to the set minimum depth threshold and maximum depth threshold, pixels with depth values ​​outside this range are removed to obtain the valid area; Based on the extracted effective area, the target package area is identified and segmented; Count the number of target packages in the corresponding depth map of the same frame.

6. The green detection model for autonomous online training of a visual separation system according to claim 5, characterized in that: The extraction logic of the valid area is: Among them, D valid (x,y) is the effective area, D min is the minimum depth threshold, D max is the maximum depth threshold; Perform threshold segmentation on the effective area and set the height threshold H threshold , the effective area is divided into background and target packages, Among them, R target (x, y) is the effective area after segmentation, and the number of target packages in the current frame image is counted.

7. The green detection model for autonomous online training of a visual separation system according to claim 1, characterized in that: When N1 < N2, the RGB image and the depth image are abnormal images. Save the two images, compare the coordinates of each target package in the same coordinate system, and output the coordinates (X e , Y e ) of the redundant targets after matching, including: Extract the center coordinates of each target package from the RGB image and depth image; Calculate the distance metric between the center coordinates of the target package in any RGB image and the center coordinates of the target package in any depth image. When the distance metric is less than the set matching threshold, match the corresponding target packages and output all coordinates in the depth image that fail to match the target package in the RGB image as redundant target coordinates.

8. The green detection model for autonomous online training of a visual separation system according to claim 7, characterized in that: The center coordinates of the target package extracted from the RGB image are marked as Among them, i∈{1,2,...,N1}, the center coordinates of the target package extracted from the depth map are marked as Where j∈{1,2,...,N2}; The calculation logic of the distance metric is: Among them, D ij It is the distance metric between the i-th target package extracted from the RGB image and the j-th target package extracted from the depth image.

9. The green detection model for autonomous online training of a visual separation system according to claim 1, characterized in that: The automatic labeling tool based on X-AnyLabeling uses the true value and RGB information and coordinate information system to automatically generate labeling results, including: Read the RGB image to be annotated and the abnormal coordinates (X e ,Y e ); Generate the label of the target package according to the set labeling rules, compare the true value with the generated label, and automatically generate the labeling result when they are consistent.

10. The green detection model for autonomous online training of a visual separation system according to claim 9, characterized in that: The set labeling rules include pre-trained deep learning models or rule matching.