Container positioning method, device, container positioning equipment and storage medium

Through the method of image segmentation and feature extraction, combined with the information of the container itself and the surrounding environment, the problem of low container positioning accuracy in the prior art is solved, and higher positioning accuracy and reliability are achieved.

CN114092408BActive Publication Date: 2025-06-06HEFEI JIZHIJIA ROBOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111249625.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-06-06
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

In the prior art, the container positioning device is positioned through the image features of the container itself, resulting in low positioning accuracy, which may cause positioning errors and missed detection.

Method used

A container positioning method is adopted to obtain image areas and semantic information through image segmentation, determine the position relationship between image areas, and determine the candidate area of ​​the target container based on the preset position relationship. Then, image features are extracted to determine the pixel candidate position, and finally the positioning result of the target container is determined based on the candidate region and the pixel candidate position.

Benefits of technology

Improve the accuracy of container positioning, reduce positioning errors and missed detection, and enhance the reliability of positioning by fusing the image characteristics of the container itself and the position information of surrounding objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092408B_ABST
    Figure CN114092408B_ABST
Patent Text Reader

Abstract

The present invention provides a container positioning method, device, container positioning equipment and storage medium, wherein the container positioning method includes: acquiring an image captured by an image sensor, performing image segmentation on the image, obtaining at least two image regions and semantic information corresponding to each image region; determining the positional relationship between each image region according to the semantic information corresponding to each image region, and determining the candidate region of the target container according to the positional relationship and the preset positional relationship; extracting the image features of the image, and determining the pixel candidate positions in the image that have the same features as the target container according to the image features; and determining the positioning result of the target container based on the candidate regions and the pixel candidate positions. The image features of the container itself and the position information of the objects around the container are integrated to improve the positioning accuracy and avoid positioning errors, missed detection and other problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technology, and in particular to a container positioning method, device, container positioning equipment and storage medium. Background Art

[0002] In recent years, with the rapid development of e-commerce, the number of user orders has increased exponentially. A warehouse needs to store a huge amount of goods. How to automatically sort these goods has become the key to improving storage efficiency. At present, warehouses mostly use container positioning equipment to realize automatic sorting of goods. The container positioning equipment takes out the target container containing the goods from the shelf to be sorted and then moves it to the designated location.

[0003] In order to realize the automatic pick-and-place function of the container positioning device, the container positioning device needs to identify and locate the position of the container. In the prior art, the container image is often captured by an image acquisition device set on the container positioning device, and then the container image is analyzed for features, and the position of the container is located by the image features of the container itself. However, the image features of the container itself are fixed, such as the shape and edge of the container, which have nothing to do with its location and environment. Therefore, the positioning accuracy of the container is not high when the container is identified and located by the image features of the container itself, which may cause positioning errors, missed detection and other problems. Summary of the invention

[0004] In view of this, embodiments of the present invention provide a container positioning method, an apparatus, a container positioning device, and a storage medium to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of the present invention, a container positioning method is provided, comprising:

[0006] Acquire an image captured by an image sensor, perform image segmentation on the image, and obtain at least two image regions and semantic information corresponding to each image region;

[0007] Determine the positional relationship between the image regions according to the semantic information corresponding to the image regions, and determine the candidate region of the target container according to the positional relationship and the preset positional relationship;

[0008] Extracting image features of the image, and determining candidate pixel positions in the image that have the same features as the target container based on the image features;

[0009] Based on the candidate regions and the candidate pixel positions, the positioning result of the target container is determined.

[0010] Optionally, performing image segmentation on the image to obtain at least two image regions and semantic information corresponding to each image region includes:

[0011] Input the image into the image segmentation model to obtain the category identification of the pixels in the image;

[0012] Clustering pixels in the image according to the category identifier to obtain at least two image regions included in the image;

[0013] According to the category identification of the pixels in each image region, the semantic information of each image region is determined.

[0014] Optionally, determining the positional relationship between the image regions according to the semantic information corresponding to the image regions includes:

[0015] Treat each image region as a node;

[0016] Determine a second image region adjacent to the first image region according to semantic information of the first image region, connect nodes of the first image region and nodes of the second image region, and construct a topological relationship graph;

[0017] The topological relationship diagram is used as the positional relationship between various image regions.

[0018] Optionally, determining a candidate area of ​​the target container according to the position relationship and the preset position relationship includes:

[0019] Determining the confidence of the target image area according to the positional relationship and the preset positional relationship, wherein the target image area is an image area including the target container;

[0020] When the confidence is greater than the confidence threshold, the target image region is determined as a candidate region of the target container.

[0021] Optionally, determining the confidence level of the target image area according to the position relationship and the preset position relationship includes:

[0022] Determine the pixels to be verified that are adjacent to the target container according to the positional relationship;

[0023] Determining reference pixels adjacent to the target container according to a preset position relationship;

[0024] Determine the similarity between the pixel to be verified and the reference pixel;

[0025] According to a preset confidence rule, the confidence corresponding to the similarity is determined, and the confidence corresponding to the similarity is used as the confidence of the target image area.

[0026] Optionally, the image feature is a geometric feature of the object; extracting the image feature of the image includes:

[0027] Extracting geometric features of objects in the image, the geometric features including at least one of the following: point features, line features, and angle features;

[0028] Taking geometric features as image features of images;

[0029] Accordingly, according to the image features, the candidate pixel positions in the image having the same features as the target container are determined, including:

[0030] determining whether the geometric features constitute a target shape, the target shape being a shape corresponding to the target container;

[0031] In the case where the geometric features constitute a target shape, the position of the target shape in the image is determined as a pixel candidate position.

[0032] Optionally, the image feature is a color feature of the object; extracting the image feature of the image includes:

[0033] Extract the color features of objects in the image and use the color features as image features of the image;

[0034] Accordingly, according to the image features, the candidate pixel positions in the image having the same features as the target container are determined, including:

[0035] Determine a target color feature among the extracted color features that is identical to a color feature of the target container;

[0036] The position of the target color feature in the image is determined as the pixel candidate position.

[0037] Optionally, determining a positioning result of the target container based on the candidate area and the candidate pixel position includes:

[0038] Determine whether there is a target reference position located in the candidate area among the pixel candidate positions;

[0039] If so, the candidate region is determined as the region where the target container is located, and the target reference position is determined as the position where the target container is located.

[0040] Optionally, the target container is a cargo box, and the target shape is a rectangular frame.

[0041] According to a second aspect of an embodiment of the present invention, there is provided a container positioning device, comprising:

[0042] An image segmentation module is configured to obtain an image captured by an image sensor, perform image segmentation on the image, and obtain at least two image regions and semantic information corresponding to each image region;

[0043] A candidate region determination module is configured to determine the positional relationship between the image regions according to the semantic information corresponding to the image regions, and determine the candidate region of the target container according to the positional relationship and the preset positional relationship;

[0044] A candidate position determination module is configured to extract image features of the image and determine, based on the image features, a candidate position of a pixel in the image that has the same features as the target container;

[0045] The positioning result determination module is configured to determine the positioning result of the target container based on the candidate area and the candidate pixel position.

[0046] Optionally, the image segmentation module includes a category identification obtaining unit, an image region obtaining unit and a semantic information determining unit;

[0047] A category identification obtaining unit is configured to input the image into the image segmentation model to obtain category identifications of pixels in the image;

[0048] An image region obtaining unit is configured to cluster pixels in the image according to the category identifier to obtain at least two image regions included in the image;

[0049] The semantic information determination unit is configured to determine the semantic information of each image region according to the category identification of the pixels in each image region.

[0050] Optionally, the candidate region determination module includes a position relationship determination subunit;

[0051] The position relationship determination subunit is configured to take each image area as a node; determine the second image area adjacent to the first image area based on the semantic information of the first image area, connect the nodes of the first image area and the nodes of the second image area, and construct a topological relationship graph; and use the topological relationship graph as the position relationship between each image area.

[0052] Optionally, the candidate region determination module further includes a confidence determination unit and a candidate region determination unit;

[0053] a confidence determination unit configured to determine the confidence of a target image region according to the positional relationship and a preset positional relationship, wherein the target image region is an image region including a target container;

[0054] The candidate region determining unit is configured to determine the target image region as a candidate region of the target container when the confidence level is greater than a confidence level threshold.

[0055] Optionally, the confidence determination unit includes a to-be-verified pixel determination subunit, a reference pixel determination subunit, a similarity determination subunit and a confidence determination subunit;

[0056] The to-be-verified pixel determination subunit is configured to determine the to-be-verified pixel adjacent to the target container according to the position relationship;

[0057] A reference pixel determination subunit is configured to determine reference pixels adjacent to the target container according to a preset position relationship;

[0058] A similarity determination subunit is configured to determine the similarity between the pixel to be verified and the reference pixel;

[0059] The confidence determination subunit is configured to determine the confidence corresponding to the similarity according to a preset confidence rule, and use the confidence corresponding to the similarity as the confidence of the target image area.

[0060] Optionally, the image feature is a geometric feature of the object; the candidate position determination module is further configured to:

[0061] Extracting geometric features of objects in the image, the geometric features including at least one of the following: point features, line features, and angle features;

[0062] Taking geometric features as image features of images;

[0063] determining whether the geometric features constitute a target shape, the target shape being a shape corresponding to the target container;

[0064] In the case where the geometric features constitute a target shape, the position of the target shape in the image is determined as a pixel candidate position.

[0065] Optionally, the image feature is a color feature of the object; and the candidate position determination module is further configured to:

[0066] Extract the color features of objects in the image and use the color features as image features of the image;

[0067] Determine a target color feature among the extracted color features that is identical to a color feature of the target container;

[0068] The position of the target color feature in the image is determined as the pixel candidate position.

[0069] Optionally, the positioning result determination module is further configured to:

[0070] Determine whether there is a target reference position located in the candidate area among the pixel candidate positions;

[0071] If so, the candidate region is determined as the region where the target container is located, and the target reference position is determined as the position where the target container is located.

[0072] Optionally, the target container is a cargo box, and the target shape is a rectangular frame.

[0073] According to a third aspect of an embodiment of the present invention, there is provided a container positioning device, comprising: an image sensor, a memory, and a processor;

[0074] The image sensor is used to capture images and transmit them to the processor;

[0075] The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the container positioning method are implemented.

[0076] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned container positioning method are implemented.

[0077] The present embodiment provides a container positioning method, an apparatus, a container positioning device and a storage medium. The container positioning method includes: acquiring an image captured by an image sensor, performing image segmentation on the image, obtaining at least two image areas and semantic information corresponding to each image area; determining the positional relationship between the image areas according to the semantic information corresponding to each image area, and determining a candidate area of ​​a target container according to the positional relationship and a preset positional relationship; extracting image features of the image, and determining, according to the image features, a candidate position of a pixel in the image having the same features as the target container; and determining a positioning result of the target container based on the candidate area and the candidate pixel position.

[0078] In addition to the target container, the image captured by the image sensor may also include other objects. The target container and other objects are in different image areas in the image. Therefore, the image captured by the image sensor can be segmented to obtain at least two image areas and semantic information corresponding to each image area. The semantic information of each image area can represent the type of object it includes and the positional relationship with other image areas. Then, the positional relationship between each image area can be determined based on the semantic information corresponding to each image area. Subsequently, the positional relationship between each image area and the preset positional relationship can be combined to determine the candidate area of ​​the target container in the captured image. After that, the captured image can be analyzed to extract image features. Based on the image features, the candidate pixel positions in the captured image that have the same features as the target container are determined. Subsequently, the candidate areas and pixel candidate positions are combined to comprehensively determine the positioning result of the target container. In this way, with the help of pre-acquired preset position information, the positional relationship between each image area in the collected image can be auxiliary verified to determine the candidate area of ​​the target container, and then the target container can be located in combination with the image features of the target container itself. The image features of the container itself and the position information of objects around the container are integrated to improve the positioning accuracy and avoid positioning errors, missed detections and other problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 is a flow chart of a container positioning method provided by an embodiment of the present invention;

[0080] Figure 2 is a schematic diagram of a container positioning scenario provided by an embodiment of the present invention;

[0081] Figure 3 is a schematic diagram of segmenting an image region provided by an embodiment of the present invention;

[0082] Figure 4 A topological relationship diagram provided by an embodiment of the present invention;

[0083] Figure 5 is a flow chart of another container positioning method provided by an embodiment of the present invention;

[0084] Figure 6 is a structural schematic diagram of a container positioning device provided by an embodiment of the present invention;

[0085] Figure 7 It is a structural block diagram of a container positioning device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0086] Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present invention, so the present invention is not limited to the specific implementation disclosed below.

[0087] The terms used in one or more embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present invention. The singular forms of "a", "said" and "the" used in one or more embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present invention refers to and includes any or all possible combinations of one or more associated listed items.

[0088] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present invention, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0089] First, the terms involved in one or more embodiments of the present invention are explained.

[0090] Container: Also generally called a material box or cargo box, it is a physical entity that holds goods / materials, including plastic boxes, cartons, paper boxes, plastic baskets, etc.

[0091] Container positioning equipment: also known as C robot (Carry handling robot), is an automated device that can pick up and transport containers from shelves.

[0092] Geometric features: features such as point features, line features and / or angle features that can represent the geometric shape of an object.

[0093] Color features: various data features representing pixel color space, such as RGB and HSV.

[0094] At present, there are mainly two types of container positioning solutions, one based on the image features of the container itself, such as the edge of the container, and the other using deep learning. The container positioning solution based on the image features of the container itself ignores the environmental information around the container, and it is easy to cause misidentification and missed detection by relying solely on the image features of the container itself. The positioning accuracy of the container positioning solution using deep learning needs to rely on annotation, and the positioning accuracy is not high, so it is mostly used as a rough positioning for selecting candidate areas.

[0095] In order to address the above problems, the present invention provides a container positioning method, a container positioning device, a container positioning equipment, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0096] The container positioning method provided in the embodiment of the present invention may be implemented by a computing device, a server, etc. for positioning a container, or a container positioning device for automatically positioning a container. The container positioning method provided in the embodiment of the present invention may be implemented by at least one of software, hardware circuits, and logic circuits set in the execution subject.

[0097] Figure 1 A flow chart of a container positioning method provided according to an embodiment of the present invention is shown, which specifically includes the following steps:

[0098] Step 102: Acquire an image captured by an image sensor, perform image segmentation on the image, and obtain at least two image regions and semantic information corresponding to each image region.

[0099] In the embodiment of the present invention, the image sensor can capture the scene within the field of view, and transmit the captured image to the execution subject of the embodiment of the present invention. After the execution subject obtains the image captured by the image sensor, it can perform image segmentation on the acquired image to obtain at least two image areas and semantic information corresponding to each image area. The image sensor includes but is not limited to an RGB camera, a grayscale camera, an RGB-D camera, an infrared camera, etc.

[0100] In addition to the target container, the image captured by the image sensor may also include other objects. The target container and other objects may be in different image areas in the image, and each image area may include an object. Therefore, the image captured by the image sensor may be segmented to obtain at least two image areas and semantic information corresponding to each image area. The semantic information of each image area may indicate the type of object it includes and the positional relationship with other image areas. The target container may be a cargo box, and the other objects except the target container may be other cargo boxes, shelves, signs on shelves, etc.

[0101] In the embodiment of the present invention, the scene within the field of view of the image sensor is a scene in which a container is placed and other objects exist around the container, such as Figure 2 A schematic diagram of a container positioning scenario provided by an embodiment of the present invention is shown.

[0102] In one implementation of the embodiment of the present invention, performing image segmentation on an image to obtain at least two image regions and semantic information corresponding to each image region can be implemented in the following manner:

[0103] Input the image into the image segmentation model to obtain the category identification of the pixels in the image;

[0104] Clustering pixels in the image according to the category identifier to obtain at least two image regions included in the image;

[0105] According to the category identification of the pixels in each image region, the semantic information of each image region is determined.

[0106] In an embodiment of the present invention, the image segmentation model may refer to a pre-trained model that can identify the category identification of pixels in an input image. In specific implementation, the above-mentioned image segmentation model can be trained by the following method: obtaining a sample image, the sample image includes a sample label, wherein the sample label is the sample identification of the pixel in the sample image; inputting the sample image into the initial model to obtain the predicted identification of the pixel; determining the loss value based on the predicted identification and sample identification of the pixel, training the initial model based on the loss value, and continuing to obtain sample images until the training stop condition is reached to obtain the image segmentation model.

[0107] It should be noted that the sample identification of a pixel in a sample image refers to the identification of a pixel to be obtained by inputting the sample image into the image segmentation model, and the predicted identification of a pixel refers to the identification of a pixel output by the initial model after the sample image is input. In addition, the training stop condition may be that the loss value is less than a preset threshold. Specifically, after training the initial model based on the loss value, it can be determined whether the loss value is less than the preset threshold. If not, return to the above step of obtaining the sample image and continue training. If so, it is determined that the training stop condition is reached.

[0108] In practical applications, the preset threshold is a critical value of the loss value. When the loss value is greater than or equal to the preset threshold, it means that there is still a certain deviation between the prediction result of the initial model and the actual result, and the parameters of the initial model still need to be adjusted, and sample images continue to be acquired to continue training the model; when the loss value is less than the preset threshold, it means that the prediction result of the initial model is close enough to the actual result, and training can be stopped. The value of the preset threshold can be determined according to actual conditions, and the embodiment of the present invention does not limit this.

[0109] In an embodiment of the present invention, a cross entropy loss function can be calculated based on the predicted identifier of the pixel and the sample identifier of the pixel included in the sample label to generate a loss value. The sample label refers to the result that the image segmentation model actually wants to output, that is, the sample identifier of the pixel included in the sample label is the real result, and the sample image is input into the initial model, and the predicted identifier of the output pixel is the predicted result. When the difference between the predicted result and the real result is small enough, it means that the predicted result is close enough to the real result. At this time, the initial model training is completed and the image segmentation model is obtained. The obtained image segmentation model can identify the category identifier of the pixel in the input image.

[0110] By calculating the loss value, the difference between the model's predicted results and the actual results can be intuitively shown. The initial model can then be trained specifically and the parameters adjusted. The specific training situation of the initial model can be judged based on the loss value, and training can be continued if the training is not qualified to improve the model's analytical ability, thereby effectively improving the rate and effect of model training.

[0111] It should be noted that the trained image segmentation model has the ability to identify the category identification of pixels in the image. When the image captured by the image sensor is input into the trained image segmentation model, the image segmentation model can output the category identification of the pixels in the image. The category identification can be used to indicate the type of object corresponding to the pixel. Since the image can include different objects, and the category identification of the pixels corresponding to different objects in the image is different, the same pixels in the image can be clustered into one image region according to the obtained category identification, thereby obtaining at least two image regions included in the image.

[0112] In practical applications, after clustering the pixels on the image and determining at least two image regions included in the image, the semantic information of each image region can be determined based on the category identification of the pixels in each image region. The semantic information of each image region may include the type of object included in the image region, and the positional relationship with other image regions (i.e., the positional relationship between the object in the image region and other objects).

[0113] For example, Figure 3 FIG. 4 shows a schematic diagram of segmenting an image region provided by an embodiment of the present invention. Figure 3 As shown, the image captured by the image sensor is input into the image segmentation model to obtain the category identification of the pixels in the image, and the pixels in the image are clustered according to the category identification to obtain Figure 3 In the five image areas shown, the category of the pixels in image area 1 is identified as cargo box 1, the category of the pixels in image area 2 is identified as the sign on the shelf, the category of the pixels in image area 3 is identified as shelf 1, the category of the pixels in image area 4 is identified as shelf 2, and the category of the pixels in image area 5 is identified as cargo box 2.

[0114] According to the category identification of the pixels in each image area, the semantic information of each image area can be determined. The semantic information of image area 1 is: cargo box 1, located in the middle of image area 2, image area 3, image area 4 and image area 5, and adjacent to image area 2, image area 3, image area 4 and image area 5; the semantic information of image area 2 is: the sign on the shelf, located above image area 1; the semantic information of image area 3 is: shelf 1, located to the right of image area 1; the semantic information of image area 4 is: shelf 2, located below image area 1; the semantic information of image area 5 is: cargo box 2, located to the left of image area 1.

[0115] In the embodiment of the present invention, an image segmentation model can be used to segment an image captured by an image sensor into different image regions, and semantic information of each image region can also be obtained. The semantic information of the image region indicates what objects are included in the image region and the positional relationship between the objects and other objects, thereby facilitating the subsequent determination of the positional relationship between the image regions, determining the candidate region of the target container, introducing the position information of objects around the target container, and improving the positioning accuracy.

[0116] Furthermore, after acquiring the image captured by the image sensor, the image is segmented, and before obtaining at least two image regions and semantic information corresponding to each image region, the image captured by the image sensor can also be denoised to remove noise in the captured image, making the image smoother and improving the accuracy of subsequent segmentation or extraction of image features. In specific implementation, the method of removing noise in the captured image includes but is not limited to using filters such as Gaussian filtering and mean filtering to denoise the image.

[0117] Step 104 : determining the positional relationship between the image regions according to the semantic information corresponding to the image regions, and determining the candidate region of the target container according to the positional relationship and a preset positional relationship.

[0118] In the embodiment of the present invention, by segmenting the image captured by the image sensor, at least two image regions and semantic information corresponding to each image region can be obtained. The semantic information corresponding to each image region can represent the type of object included in itself and the positional relationship with other image regions. Therefore, according to the semantic information corresponding to each image region, the positional relationship between each image region can be determined, and thus the candidate region of the target container can be determined according to the positional relationship between each image region and the preset positional relationship. Among them, the candidate region of the target container can refer to the image region where the target container may be located in the captured image.

[0119] In a specific implementation, the preset position relationship may refer to the position of the target container and the position relationship between the target container and other objects acquired in advance, and the position relationship between each image region refers to the position of the target container and the position relationship between the target container and other objects acquired by processing and analyzing the acquired image. Therefore, after the acquired image is processed and analyzed to determine the position relationship between each image region, the position relationship determined by processing and analysis may be verified based on the preset position relationship acquired in advance to determine the candidate region of the target container.

[0120] In an implementation of the embodiment of the present invention, the positional relationship between the image regions may be represented by a topological relationship diagram, that is, the positional relationship between the image regions may be determined according to the semantic information corresponding to the image regions, which may be implemented in the following manner:

[0121] Treat each image region as a node;

[0122] Determine a second image region adjacent to the first image region according to semantic information of the first image region, connect nodes of the first image region and nodes of the second image region, and construct a topological relationship graph;

[0123] The topological relationship diagram is used as the positional relationship between various image regions.

[0124] It should be noted that the semantic information corresponding to each image area can represent the type of object it includes, as well as its positional relationship with other image areas, that is, its positional relationship with other objects. Therefore, each image area can be used as a node in a topological relationship graph. Based on the semantic information corresponding to each image area, the positional relationship between each image area can be determined, and adjacent nodes can be connected to construct a corresponding topological relationship graph, which can represent the positional relationship between each image area.

[0125] In one possible implementation, the first image region may be an image region located at a center position among the at least two image regions obtained by segmentation, and the second image region may be an image region that may be adjacent to the first image region. That is, the image region located at a center position among the at least two image regions obtained by segmentation may be taken as a starting point, and other adjacent image regions may be connected in sequence to construct a corresponding topological relationship graph.

[0126] In another possible implementation, the first image area may refer to an image area containing a target container, and the second image area may refer to an image area containing other objects outside the target container. The container may refer to a cargo box, and the target container may refer to any cargo box in the image that needs to be located, that is, each cargo box in the image can be used as a target container for location. Specifically, starting from the image area containing the target container in the image, the image areas around the target container are sequentially connected to construct a corresponding topological relationship diagram.

[0127] Using the above example, assuming that the target container is cargo box 1, according to the semantic information of image area 1, it is determined that image area 1 is located in the middle of image area 2, image area 3, image area 4 and image area 5. According to the semantic information of image area 2, it can be determined that image area 2 is located above image area 1. According to the semantic information of image area 3, it can be determined that image area 3 is located to the right of image area 1. According to the semantic information of image area 4, it can be determined that image area 4 is located below image area 1. According to the semantic information of image area 5, it can be determined that image area 5 is located to the left of image area 1. Therefore, the nodes of image area 2, image area 3, image area 4 and image area 5 can be set above, below, to the right and left of the node of image area 1 in sequence, and the nodes of image area 1 can be connected in sequence with the nodes of image area 2, image area 3, image area 4 and image area 5 to construct the following. Figure 4 A topological relationship diagram provided by an embodiment of the present invention is shown.

[0128] Furthermore, when the objects around the target container are of different object types from the target container, they have relatively easy-to-identify environmental features, and it is relatively easy to locate the target container. However, if the objects around the target container are of the same object type as the target container, they do not have relatively easy-to-identify environmental features, and it is relatively difficult to locate the target container. Therefore, after connecting the nodes of the first image area and the nodes of the second image area, the weight of the connecting edge between the nodes of the first image area and the nodes of the second image area can also be set according to the object types included in the first image area and the second image area.

[0129] In practical applications, since two adjacent objects are of different types, they have more object features and are easier to locate the target container. Therefore, when the first image area and the second image area include different object types, the weight of the connecting edge between the nodes of the first image area and the second image area can be set to be larger. When the first image area and the second image area include the same object type, the weight of the connecting edge between the nodes of the first image area and the second image area can be set to be smaller, thereby distinguishing whether the object types corresponding to the two connected nodes are the same based on the weight.

[0130] Using the above example, Figure 4 As shown, the object types of image region 1 and image region 2 are different, so the weight of the connecting edge between the nodes of image region 1 and the nodes of image region 2 can be 0.3; the object types of image region 1 and image region 3 are different, so the weight of the connecting edge between the nodes of image region 1 and the nodes of image region 3 can also be 0.3; the object types of image region 1 and image region 4 are different, so the weight of the connecting edge between the nodes of image region 1 and the nodes of image region 4 can also be 0.3; the object types of image region 1 and image region 5 are the same, so the weight of the connecting edge between the nodes of image region 1 and the nodes of image region 5 can be 0.1.

[0131] In an embodiment of the present invention, each image area can be used as a node, and adjacent nodes can be connected to construct a topological relationship graph. The topological relationship graph is used as the positional relationship between each image area. The positional relationship between each image area is represented by the topological relationship graph, making the positional relationship between each image area simple and clear.

[0132] In an implementation of the embodiment of the present invention, the confidence of the target image region may be determined according to the positional relationship between the image regions and the preset positional relationship, thereby determining the candidate region of the target container. That is, determining the candidate region of the target container according to the positional relationship and the preset positional relationship may be implemented as follows:

[0133] Determining the confidence of the target image area according to the positional relationship and the preset positional relationship, wherein the target image area is an image area including the target container;

[0134] When the confidence is greater than the confidence threshold, the target image region is determined as a candidate region of the target container.

[0135] Specifically, the confidence level may refer to the probability that the target image region determined by the recognition processing of the captured image is a candidate region containing the target container, that is, the reliability of the target image region being the candidate region. In addition, the confidence level threshold may refer to a preset value used to determine the similarity between the objects around the target image region in the captured image and the objects around the target container known in advance, that is, whether the objects around the target image region in the captured image are substantially the same as the objects around the target container known in advance.

[0136] In practical applications, when the positional relationship between each image region is a topological relationship graph, the preset positional relationship may also be a topological relationship graph, and the two topological relationship graphs may be compared later to determine the confidence of the target image region. If the connection edge between two adjacent nodes in the topological relationship graph corresponding to the positional relationship between each image region carries a weight, then the connection edge between two adjacent nodes in the topological relationship graph corresponding to the preset positional relationship may also carry a weight, and the two topological relationship graphs may be compared later to determine the confidence of the subsequent region.

[0137] It should be noted that based on the preset position relationship, it is possible to know in advance what objects are placed around the target container, and then based on the position relationship between each image area, it is possible to determine what objects exist around the target image area in the captured image, thereby determining whether the objects around the target image area in the captured image are the same as the objects around the target container known in advance, thereby determining the confidence level of the target image area.

[0138] If the confidence is greater than the confidence threshold, it means that the objects around the target image area in the captured image are roughly the same as the objects around the target container known in advance, and the reliability that the target container is contained in the target image area is high. Therefore, the target image area can be determined as a candidate area for the target container at this time; if the confidence is not greater than the confidence threshold, it means that the objects around the target image area in the captured image are relatively different from the objects around the target container known in advance, and the reliability that the target container is contained in the target image area is not high. Therefore, the target image area is not a candidate area for the target container at this time, that is, the candidate area for the target container cannot be determined, the target container positioning fails, and there is no need to perform subsequent image feature extraction. The staff can be notified to take other measures.

[0139] In an implementation of the embodiment of the present invention, the confidence of the target image area can be determined based on pixel similarity, that is, the confidence of the target image area is determined according to the position relationship and the preset position relationship, which can be implemented as follows:

[0140] Determine the pixels to be verified that are adjacent to the target container according to the positional relationship;

[0141] Determining reference pixels adjacent to the target container according to a preset position relationship;

[0142] Determine the similarity between the pixel to be verified and the reference pixel;

[0143] According to a preset confidence rule, the confidence corresponding to the similarity is determined, and the confidence corresponding to the similarity is used as the confidence of the target image area.

[0144] Specifically, the pixel to be verified may refer to an object included in an image area adjacent to the target image area in the captured image, that is, the pixel to be verified may refer to a pixel adjacent to the target container in the captured image; the reference pixel may refer to an object adjacent to the target container in a preset position relationship acquired in advance. In addition, the similarity between the pixel to be verified and the reference pixel may refer to the repetition of the pixel to be verified and the reference pixel, that is, how many of the pixels to be verified are the same as the reference pixels. The pixels to be verified and the reference pixels may be the same in position and object type. Furthermore, the confidence rule may refer to a pre-set rule for determining the confidence corresponding to the similarity. For example, the confidence rule may be a pre-set correspondence between the similarity and the confidence. After the similarity is determined, the corresponding confidence may be determined based on the correspondence.

[0145] It should be noted that, based on the positional relationship between each image area, it is possible to determine what objects exist around the target image area in the captured image, that is, the pixels to be verified adjacent to the target container; based on the preset positional relationship, it is possible to determine what objects are placed around the target container in the pre-known information, that is, the reference pixels adjacent to the target container. Then, it is possible to determine the percentage of the target objects in the pixels to be verified that are the same as the reference pixels in the total number of pixels to be verified, and this percentage is the similarity between the pixels to be verified and the reference pixels. Subsequently, based on the preset confidence rules, the confidence corresponding to the similarity, that is, the confidence of the target image area, can be determined.

[0146] Continuing with the above example, assuming that the target container is still box 1, the preset position relationship is that box 1 is located between the mark on the shelf, shelf 1, shelf 2 and box 2, the mark on the shelf is located above box 1, shelf 1 is located to the right of box 1, shelf 2 is located below box 1, and box 2 is located to the left of box 1. That is, the reference pixels adjacent to the target container box 1 are: the mark on the shelf, shelf 1, shelf 2 and box 2, and the positions of the reference pixels are: the mark on the shelf-above, shelf 1-right, shelf 2-below, box 2-left. From the above example content, it can be seen that the pixels to be verified adjacent to the target container box 1 are: the mark on the shelf, shelf 1, shelf 2 and box 2, and the positions of the pixels to be verified are: the mark on the shelf-above, shelf 1-right, shelf 2-below, box 2-left. From the above, it can be seen that the object type and position of the pixel to be verified are exactly the same as those of the reference pixel, that is, the similarity between the pixel to be verified and the reference pixel is 100% at this time. Assuming that the confidence level corresponding to the similarity between 100% and 70% is 0.9, the confidence level of the target image area is greater than the confidence threshold 0.7 (pre-set), and the image area 1 corresponding to the target container box 1 can be determined as the candidate area of ​​the target container.

[0147] In the embodiment of the present invention, based on the positional relationship between each image area and the preset positional relationship, it can be determined whether the objects around the target image area in the captured image are substantially the same as the objects around the target container known in advance, thereby determining the confidence of the target image area, and subsequently determining the candidate area of ​​the target container based on the confidence, that is, the candidate area of ​​the target container is determined based on the position information of the surrounding objects of the target container, which facilitates the subsequent fusion of the image features of the container itself and the position information of the objects around the container, so as to locate the target container, improve the positioning accuracy, and avoid positioning errors, missed detections, and other problems.

[0148] Step 106 , extracting image features of the image, and determining candidate pixel positions in the image that have the same features as the target container based on the image features.

[0149] It should be noted that the image feature may refer to the feature of the target container itself in the image. Based on the image feature, the position of the object having the same feature as the target container in the image, ie, the candidate pixel position, may be accurately determined.

[0150] In one implementation of the embodiment of the present invention, since the target container is a cargo box, and the cargo box often has a specific geometric shape, the geometric features in the collected image can be extracted to locate the candidate pixel positions, that is, the image features are the geometric features of the object. At this time, the image features of the image can be extracted in the following manner:

[0151] Extracting geometric features of objects in the image, the geometric features including at least one of the following: point features, line features, and angle features;

[0152] Taking geometric features as image features of images;

[0153] Accordingly, according to the image features, the candidate pixel positions having the same features as the target container in the image can be determined by:

[0154] determining whether the geometric features constitute a target shape, the target shape being a shape corresponding to the target container;

[0155] In the case where the geometric features constitute a target shape, the position of the target shape in the image is determined as a pixel candidate position.

[0156] Specifically, the geometric features of an object may refer to features related to geometric shapes, such as point features, line features, and / or corner features, etc. Of course, they may also be directly shape features, etc., and the embodiments of the present invention do not limit this. In addition, the target shape may be a shape corresponding to the target container. In a possible implementation, the target shape may be a rectangular frame. Of course, in actual applications, the target shape may also be other shapes, such as a circle, a polygon, etc. In actual implementation, the captured image may be analyzed to extract geometric features such as point features, line features, and / or corner features of the object in the image, and the geometric features such as point features, line features, and / or corner features may be used as image features of the image.

[0157] It should be noted that after extracting geometric features such as point features, line features and / or corner features of each object in the image, the shape of the object composed of geometric features such as point features, line features and / or corner features can be determined, and whether there is a shape corresponding to the target container in the object shape. If so, it means that the object of the target shape composed of geometric features such as point features, line features and / or corner features has the same features as the target container and may be the target container. At this time, the position of the target shape in the image can be determined as the pixel candidate position. If not, it means that the object composed of geometric features such as point features, line features and / or corner features does not have the same features as the target container, and at this time, it can be determined that there is no pixel candidate position.

[0158] In the embodiment of the present invention, the geometric features of the target container itself can be used to determine the candidate positions of pixels in the collected image that have the same geometric features as the target container, so that based on the geometric features of the target container itself, the positions of objects with the same geometric features as the target container in the image can be accurately determined. Subsequently, the position information of objects surrounding the target container and the positions of objects with the same geometric features as the target container in the image can be combined to comprehensively determine the positioning result of the target container, thereby improving the positioning accuracy.

[0159] In one implementation of the embodiment of the present invention, since the target container is a cargo box, and the cargo box may have a color different from other objects, the color features in the collected image may be extracted to locate the candidate pixel positions, that is, the image features are the color features of the object. At this time, the image features of the image may be extracted in the following manner:

[0160] Extract the color features of objects in the image and use the color features as image features of the image;

[0161] Accordingly, according to the image features, the candidate pixel positions having the same features as the target container in the image can be determined by:

[0162] Determine a target color feature among the extracted color features that is identical to a color feature of the target container;

[0163] The position of the target color feature in the image is determined as the pixel candidate position.

[0164] In actual implementation, the collected image can be analyzed to extract the color features of the object in the image, such as red, yellow or green, and the color features can be used as the image features of the image.

[0165] It should be noted that after extracting the color features of each object in the image, the target color feature that is the same as the color feature of the target container can be determined among the extracted color features. The object corresponding to the target color feature has the same features as the target container and may be the target container. Therefore, the position of the target color feature in the image can be determined as the pixel candidate position.

[0166] In the embodiment of the present invention, the color features of the target container itself can be used to determine the candidate positions of pixels in the collected image that have the same color features as the target container, thereby accurately determining the positions of objects with the same color features as the target container in the image based on the color features of the target container itself. Subsequently, the position information of objects surrounding the target container and the positions of objects with the same color features as the target container in the image can be combined to comprehensively determine the positioning result of the target container, thereby improving the positioning accuracy.

[0167] Step 108: Determine the positioning result of the target container based on the candidate regions and the candidate pixel positions.

[0168] It should be noted that the candidate area may refer to the image area where the target container may be located in the captured image, and the pixel candidate position may refer to the position of an object having the same characteristics as the target container in the captured image, that is, the position of the object that may be the target container in the image. Therefore, by fusing the candidate area and the pixel candidate position, the positioning result of the target container can be accurately determined.

[0169] In an implementation of the embodiment of the present invention, determining the positioning result of the target container based on the candidate area and the candidate pixel position can be implemented in the following manner:

[0170] Determine whether there is a target reference position located in the candidate area among the pixel candidate positions;

[0171] If so, the candidate region is determined as the region where the target container is located, and the target reference position is determined as the position where the target container is located.

[0172] It should be noted that the pixel candidate position is the position of an object having the same characteristics as the target container, that is, the position of the object that may be the target container in the image, that is, the object at each pixel candidate position may be the target container, and the candidate area is the image area where the target container may be located.

[0173] In practical applications, the candidate area and the pixel candidate position information can be integrated. If there is a target reference position located in the candidate area in the pixel candidate position, it means that there is an object with the same characteristics as the target container in the image area where the target container may be located. Therefore, the object in the candidate area is most likely the target container to be located. At this time, the candidate area is the area where the target container is located, and the target reference position located in the candidate area is the exact position of the target container. That is, the positioning result of the target container is the target reference position of the target container in the candidate area.

[0174] In addition, if the target reference position located in the candidate area does not exist in the pixel candidate position, it means that there is no object with the same characteristics as the target container in the image area where the target container may be located. At this time, it can be determined that the candidate area most likely does not include the target container. At this time, it can be determined that the target container positioning has failed, and the staff can be reminded to take other measures.

[0175] Continuing with the above example, the target container is cargo box 1, and the candidate area of ​​the target container is image area 1. Assuming that the determined pixel candidate positions are position 1, position 2, position 3, and position 4, among which position 3 is in image area 1 in the image, it can be determined that image area 1 is the area where the target container is located, and position 3 is the position where the target container is located, that is, the positioning result of the target container is that the target container is at position 3 in image area 1.

[0176] By applying the embodiment of the present invention, the image captured by the image sensor may include other objects in addition to the target container. The target container and other objects are located in different image regions in the image. Therefore, the image captured by the image sensor may be segmented to obtain at least two image regions and semantic information corresponding to each image region. The semantic information of each image region may indicate the type of object included in it and the positional relationship with other image regions. Then, the positional relationship between each image region may be determined based on the semantic information corresponding to each image region. Subsequently, the positional relationship between each image region and the preset positional relationship may be combined to determine the candidate region of the target container in the captured image. Afterwards, the captured image may be analyzed to extract image features. Based on the image features, the candidate pixel positions in the captured image having the same features as the target container may be determined. Subsequently, the candidate regions and the candidate pixel positions may be combined to comprehensively determine the positioning result of the target container. In this way, with the help of pre-acquired preset position information, the positional relationship between each image area in the collected image can be auxiliary verified to determine the candidate area of ​​the target container, and then the target container can be located in combination with the image features of the target container itself. The image features of the container itself and the position information of objects around the container are integrated to improve the positioning accuracy and avoid positioning errors, missed detections and other problems.

[0177] The container in the above embodiment may be a cargo box. For ease of understanding, the container positioning method provided by the embodiment of the present invention is introduced below in combination with the application scenario of cargo box positioning. In the embodiment of the present invention, the container positioning device includes a positioning device for positioning the cargo box, an image sensor is installed at the bottom of the positioning device, and the container positioning device also includes a processor, which executes the following steps: Figure 5 The container positioning method shown, Figure 5 A flow chart of another container positioning method provided by an embodiment of the present invention is shown, which specifically includes the following steps.

[0178] Step 1: The image sensor captures the image.

[0179] Specifically, the image sensor includes but is not limited to an RGB camera, a grayscale camera, an RGB-D camera, an infrared camera, etc.

[0180] Step 2: De-noise the acquired image to obtain the target image.

[0181] Specifically, the method of removing noise includes but is not limited to using filters such as Gaussian filtering and mean filtering. Removing noise from an image can make the image smoother and reduce the impact of noise on subsequent steps.

[0182] Step 3: Segment the target image to obtain at least two image regions and the semantic information corresponding to each image region.

[0183] The specific image segmentation method is: input the image into the image segmentation model to obtain the category identification of the pixels in the image; cluster the pixels in the image according to the category identification to obtain at least two image regions included in the image; determine the semantic information of each image region according to the category identification of the pixels in each image region.

[0184] In the embodiment of the present invention, an image segmentation model can be used to segment an image captured by an image sensor into different image regions, and semantic information of each image region can also be obtained. The semantic information of the image region indicates what objects are included in the image region and the positional relationship between the objects and other objects, thereby facilitating the subsequent determination of the positional relationship between the image regions, determining the candidate region of the target container, introducing the position information of objects around the target container, and improving the positioning accuracy.

[0185] Step 4: Determine the positional relationship between the image regions according to the semantic information corresponding to the image regions, and determine the candidate region of the target container according to the positional relationship between the image regions and the preset positional relationship.

[0186] The specific method for determining the positional relationship is: taking each image area as a node; determining the second image area adjacent to the first image area based on the semantic information of the first image area, connecting the nodes of the first image area and the nodes of the second image area to construct a topological relationship graph; and taking the topological relationship graph as the positional relationship between each image area.

[0187] In an embodiment of the present invention, each image area can be used as a node, and adjacent nodes can be connected to construct a topological relationship graph. The topological relationship graph is used as the positional relationship between each image area. The positional relationship between each image area is represented by the topological relationship graph, making the positional relationship between each image area simple and clear.

[0188] The specific candidate region determination method is as follows: according to the positional relationship and the preset positional relationship, the confidence of the target image region is determined, and the target image region is the image region containing the target container; when the confidence is greater than the confidence threshold, the target image region is determined as the candidate region of the target container. The confidence of the target image region is determined according to the positional relationship and the preset positional relationship, including: according to the positional relationship, determining the pixel to be verified adjacent to the target container; according to the preset positional relationship, determining the reference pixel adjacent to the target container; determining the similarity between the pixel to be verified and the reference pixel; according to the preset confidence rule, determining the confidence corresponding to the similarity, and using the confidence corresponding to the similarity as the confidence of the target image region.

[0189] In the embodiment of the present invention, based on the positional relationship between each image area and the preset positional relationship, it can be determined whether the objects around the target image area in the captured image are substantially the same as the objects around the target container known in advance, thereby determining the confidence of the target image area, and subsequently determining the candidate area of ​​the target container based on the confidence, that is, the candidate area of ​​the target container is determined based on the position information of the surrounding objects of the target container, which facilitates the subsequent fusion of the image features of the container itself and the position information of the objects around the container, so as to locate the target container, improve the positioning accuracy, and avoid positioning errors, missed detections, and other problems.

[0190] Step 5: Extract the image features of the target image, and determine the candidate pixel positions in the target image that have the same features as the target container based on the image features.

[0191] The specific image feature extraction method is: extracting geometric features of objects in the image, the geometric features include at least one of the following: point features, line features and angle features; using the geometric features as image features of the image. Accordingly, the specific pixel candidate position determination method can be: determining whether the geometric features constitute a target shape, the target shape is the shape corresponding to the target container; if the geometric features constitute the target shape, determining the position of the target shape in the image as the pixel candidate position.

[0192] A specific method for extracting image features may also be: extracting color features of objects in an image, and using the color features as image features of the image. Accordingly, a specific method for determining candidate pixel positions may be: determining a target color feature that is the same as the color feature of the target container among the extracted color features; and determining the position of the target color feature in the image as the candidate pixel position.

[0193] In the embodiments of the present invention, the candidate positions of pixels having the same image features as the target container in the collected image can be determined by using the image features of the target container itself, thereby accurately determining the positions of objects having the same image features as the target container in the image based on the image features of the target container itself. Subsequently, the positioning result of the target container can be comprehensively determined by combining the position information of objects surrounding the target container and the positions of objects having the same image features as the target container in the image, thereby improving the positioning accuracy.

[0194] Step 6: Based on the candidate regions and pixel candidate positions, determine the positioning result of the target container.

[0195] The specific method for determining the positioning result is: determine whether there is a target reference position located in the candidate area among the pixel candidate positions; if so, determine the candidate area as the area where the target container is located, and determine the target reference position as the position where the target container is located.

[0196] In the container positioning solution provided by the embodiment of the present invention, the positional relationship between each image area in the collected image can be auxiliary verified with the help of pre-acquired preset position information, and the candidate area of ​​the target container can be determined. Then, the target container is positioned in combination with the image features of the target container itself, which integrates the image features of the container itself and the position information of objects around the container, improves the positioning accuracy, and avoids problems such as positioning errors and missed detections.

[0197] Corresponding to the above method embodiment, the present invention also provides a container positioning device embodiment, Figure 6 FIG. 2 shows a schematic diagram of the structure of a container positioning device provided by an embodiment of the present invention. Figure 6 As shown, the device comprises:

[0198] The image segmentation module 620 is configured to obtain an image captured by an image sensor, perform image segmentation on the image, and obtain at least two image regions and semantic information corresponding to each image region;

[0199] The candidate region determination module 640 is configured to determine the positional relationship between the image regions according to the semantic information corresponding to the image regions, and determine the candidate region of the target container according to the positional relationship and the preset positional relationship;

[0200] The candidate position determination module 660 is configured to extract image features of the image and determine the candidate positions of pixels in the image that have the same features as the target container based on the image features;

[0201] The positioning result determination module 680 is configured to determine the positioning result of the target container based on the candidate area and the candidate pixel position.

[0202] By applying the embodiment of the present invention, the image captured by the image sensor may include other objects in addition to the target container. The target container and other objects are located in different image regions in the image. Therefore, the image captured by the image sensor may be segmented to obtain at least two image regions and semantic information corresponding to each image region. The semantic information of each image region may indicate the type of object included in it and the positional relationship with other image regions. Then, the positional relationship between each image region may be determined based on the semantic information corresponding to each image region. Subsequently, the positional relationship between each image region and the preset positional relationship may be combined to determine the candidate region of the target container in the captured image. Afterwards, the captured image may be analyzed to extract image features. Based on the image features, the candidate pixel positions in the captured image having the same features as the target container may be determined. Subsequently, the candidate regions and the candidate pixel positions may be combined to comprehensively determine the positioning result of the target container. In this way, with the help of pre-acquired preset position information, the positional relationship between each image area in the collected image can be auxiliary verified to determine the candidate area of ​​the target container, and then the target container can be located in combination with the image features of the target container itself. The image features of the container itself and the position information of objects around the container are integrated to improve the positioning accuracy and avoid positioning errors, missed detections and other problems.

[0203] Optionally, the image segmentation module 620 includes a category identification obtaining unit, an image region obtaining unit and a semantic information determining unit;

[0204] A category identification obtaining unit is configured to input the image into the image segmentation model to obtain category identifications of pixels in the image;

[0205] An image region obtaining unit is configured to cluster pixels in the image according to the category identifier to obtain at least two image regions included in the image;

[0206] The semantic information determination unit is configured to determine the semantic information of each image region according to the category identification of the pixels in each image region.

[0207] Optionally, the candidate region determination module 640 includes a position relationship determination subunit;

[0208] The position relationship determination subunit is configured to take each image area as a node; determine the second image area adjacent to the first image area based on the semantic information of the first image area, connect the nodes of the first image area and the nodes of the second image area, and construct a topological relationship graph; and use the topological relationship graph as the position relationship between each image area.

[0209] Optionally, the candidate region determination module 640 further includes a confidence determination unit and a candidate region determination unit;

[0210] a confidence determination unit configured to determine the confidence of a target image region according to the positional relationship and a preset positional relationship, wherein the target image region is an image region including a target container;

[0211] The candidate region determining unit is configured to determine the target image region as a candidate region of the target container when the confidence level is greater than a confidence level threshold.

[0212] Optionally, the confidence determination unit includes a to-be-verified pixel determination subunit, a reference pixel determination subunit, a similarity determination subunit and a confidence determination subunit;

[0213] The to-be-verified pixel determination subunit is configured to determine the to-be-verified pixel adjacent to the target container according to the position relationship;

[0214] A reference pixel determination subunit is configured to determine reference pixels adjacent to the target container according to a preset position relationship;

[0215] A similarity determination subunit is configured to determine the similarity between the pixel to be verified and the reference pixel;

[0216] The confidence determination subunit is configured to determine the confidence corresponding to the similarity according to a preset confidence rule, and use the confidence corresponding to the similarity as the confidence of the target image area.

[0217] Optionally, the image feature is a geometric feature of the object; the candidate position determination module 660 is further configured to:

[0218] Extracting geometric features of objects in the image, the geometric features including at least one of the following: point features, line features, and angle features;

[0219] Taking geometric features as image features of images;

[0220] determining whether the geometric features constitute a target shape, the target shape being a shape corresponding to the target container;

[0221] In the case where the geometric features constitute a target shape, the position of the target shape in the image is determined as a pixel candidate position.

[0222] Optionally, the image feature is a color feature of the object; the candidate position determination module 660 is further configured to:

[0223] Extract the color features of objects in the image and use the color features as image features of the image;

[0224] Determine a target color feature among the extracted color features that is identical to a color feature of the target container;

[0225] The position of the target color feature in the image is determined as the pixel candidate position.

[0226] Optionally, the positioning result determination module 680 is further configured to:

[0227] Determine whether there is a target reference position located in the candidate area among the pixel candidate positions;

[0228] If so, the candidate region is determined as the region where the target container is located, and the target reference position is determined as the position where the target container is located.

[0229] Optionally, the target container is a cargo box, and the target shape is a rectangular frame.

[0230] The above is a schematic scheme of a container positioning device of this embodiment. It should be noted that the technical scheme of the container positioning device and the technical scheme of the container positioning method described above are of the same concept, and the details not described in detail in the technical scheme of the container positioning device can be referred to the description of the technical scheme of the container positioning method described above.

[0231] Figure 7 The structural block diagram of a container positioning device provided by an embodiment of the present invention is shown. The components of the container positioning device 700 include but are not limited to an image sensor 710, a memory 720 and a processor 730. The processor 730 is connected to the image sensor 710 and the memory 720 via a bus 740, and the database 760 is used to store data.

[0232] The container positioning device 700 also includes an access device 750, which enables the container positioning device 700 to communicate via one or more networks 770. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 750 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world microwave interconnection access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0233] In one embodiment of the present invention, the above components of the container positioning device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 7 The container positioning device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of the present invention. Those skilled in the art can add or replace other components as needed.

[0234] The image sensor 710 is used to capture images and transmit the images to the processor 730; the processor 730 is used to execute the following computer executable instructions, which, when executed by the processor, implement any step of the above-mentioned container positioning method.

[0235] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the container positioning method described above are of the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the container positioning method described above.

[0236] An embodiment of the present invention further provides a computer-readable storage medium storing computer instructions, which are used to implement any of the steps of the above-mentioned container positioning method when executed by a processor.

[0237] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the container positioning method described above are of the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the container positioning method described above.

[0238] The above describes specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0239] Computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electric carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of computer readable media may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer readable media do not include electric carrier signals and telecommunication signals.

[0240] It should be noted that, for the convenience of description, the aforementioned method embodiments are all described as a series of action combinations, but those skilled in the art should be aware that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0241] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0242] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the present invention. The present invention selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A container positioning method, It is characterized in that The method comprises: Acquire an image captured by an image sensor, perform image segmentation on the image, and obtain at least two image regions and semantic information corresponding to each image region, wherein the semantic information of each image region indicates the type of object included in the image region and the positional relationship with other image regions; Determining, according to semantic information corresponding to each image region, a positional relationship between the image regions, and determining, according to the positional relationship and a preset positional relationship, a candidate region of a target container, wherein the preset positional relationship is a pre-acquired position of the target container and a positional relationship between the target container and other objects, and determining, according to the positional relationship and the preset positional relationship, the candidate region of the target container comprises: determining, according to the positional relationship between the image regions, objects surrounding the target container in the image, and determining, based on the objects surrounding the target container determined according to the preset positional relationship and the objects surrounding the target container in the image, the candidate region of the target container; Extracting image features of the image, and determining, based on the image features, candidate pixel positions in the image that have the same features as the target container; Based on the candidate area and the candidate pixel position, a positioning result of the target container is determined.

2. The container positioning method according to claim 1, It is characterized in that The performing image segmentation on the image to obtain at least two image regions and semantic information corresponding to each image region includes: Inputting the image into an image segmentation model to obtain category identifications of pixels in the image; Clustering pixels in the image according to the category identifier to obtain at least two image regions included in the image; According to the category identification of the pixels in each image region, the semantic information of each image region is determined.

3. The container positioning method according to claim 1, It is characterized in that The determining the positional relationship between the image regions according to the semantic information corresponding to the image regions includes: Taking each of the image regions as a node; Determine, according to the semantic information of the first image region, a second image region adjacent to the first image region, connect the nodes of the first image region and the nodes of the second image region, and construct a topological relationship graph; The topological relationship diagram is used as the positional relationship between the various image regions.

4. The container positioning method according to claim 1, It is characterized in that The determining the candidate area of ​​the target container according to the position relationship and the preset position relationship includes: Determining the confidence of a target image region according to the positional relationship and the preset positional relationship, the target image region being an image region including the target container; When the confidence is greater than a confidence threshold, the target image region is determined as a candidate region of the target container.

5. The container positioning method according to claim 4, It is characterized in that The step of determining the confidence level of the target image area according to the position relationship and the preset position relationship includes: Determining pixels to be verified that are adjacent to the target container according to the positional relationship; Determining reference pixels adjacent to the target container according to the preset position relationship; Determining the similarity between the pixel to be verified and the reference pixel; According to a preset confidence rule, the confidence corresponding to the similarity is determined, and the confidence corresponding to the similarity is used as the confidence of the target image area.

6. The container positioning method according to any one of claims 1 to 5, It is characterized in that The image feature is a geometric feature of an object; and extracting the image feature of the image includes: Extracting geometric features of the object in the image, wherein the geometric features include at least one of the following: point features, line features, and angle features; using the geometric features as image features of the image; Accordingly, determining, according to the image features, candidate pixel positions in the image having the same features as the target container includes: determining whether the geometric features constitute a target shape, the target shape being a shape corresponding to the target container; In the case where the geometric features constitute a target shape, the position of the target shape in the image is determined as the candidate pixel position.

7. The container positioning method according to any one of claims 1 to 5, It is characterized in that The image feature is a color feature of the object; and extracting the image feature of the image includes: Extracting color features of objects in the image, and using the color features as image features of the image; Accordingly, determining, according to the image features, candidate pixel positions in the image having the same features as the target container includes: Determine a target color feature among the extracted color features that is identical to a color feature of the target container; The position of the target color feature in the image is determined as the pixel candidate position.

8. The container positioning method according to any one of claims 1 to 5, It is characterized in that The determining the positioning result of the target container based on the candidate area and the pixel candidate position includes: Determine whether there is a target reference position located in the candidate area among the pixel candidate positions; If so, the candidate area is determined as the area where the target container is located, and the target reference position is determined as the position where the target container is located.

9. The container positioning method according to claim 6, It is characterized in that The target container is a cargo box, and the target shape is a rectangular frame.

10. A container positioning device, It is characterized in that The device comprises: An image segmentation module is configured to obtain an image captured by an image sensor, perform image segmentation on the image, and obtain at least two image regions and semantic information corresponding to each image region, wherein the semantic information of each image region indicates the type of object included in the image region and the positional relationship with other image regions; The candidate region determination module is configured to determine the positional relationship between each image region according to semantic information corresponding to each image region, and determine the candidate region of the target container according to the positional relationship and a preset positional relationship, wherein the preset positional relationship is a pre-acquired position of the target container and a positional relationship between the target container and other objects, and the determining the candidate region of the target container according to the positional relationship and the preset positional relationship comprises: determining objects around the target container in the image according to the positional relationship between each image region, and determining the candidate region of the target container based on the objects around the target container determined according to the preset positional relationship and the objects around the target container in the image; a candidate position determination module configured to extract image features of the image, and determine, based on the image features, a candidate position of a pixel in the image that has the same features as the target container; The positioning result determination module is configured to determine the positioning result of the target container based on the candidate area and the candidate pixel position.

11. A container positioning device, include: Image sensors, memory and processors; The image sensor is used to capture images and transmit the images to the processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the container positioning method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium storing computer instructions, wherein when the instructions are executed by a processor, the steps of the container positioning method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Vertebral positioning recognition and segmentation method based on FCN neural network and antagonistic learning

    CN109523523A

  • Visual SLAM method based on instance segmentation

    CN110738673A