Detection method and device, object monitoring system, computing equipment and storage medium

By using multi-objective detection model and image processing technology, the difficulty of identifying small or moving QR codes in complex environments is solved, and efficient and accurate QR code recognition and tracking is achieved.

CN120124651APending Publication Date: 2025-06-10WESTLAKE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311686500.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently identify and track small or moving QR codes, especially in complex image environments, resulting in slow recognition speed and low accuracy.

Method used

The pre-trained multi-object detection model is used to input images to process them, detect multiple types of detection targets, including QR codes, and crop the image area through the detection box, and identify the QR codes based on interpolation, adaptive threshold binarization and quadrilateral fitting techniques.

Benefits of technology

It realizes efficient identification and tracking of small or sports QR codes, improves detection speed and accuracy, and enables real-time detection and tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124651A_ABST
    Figure CN120124651A_ABST
Patent Text Reader

Abstract

The invention relates to a detection method and device, an object monitoring system, computing equipment and a storage medium. A detection method includes acquiring an image including a two-dimensional code for attaching to an object and recording information about the object. The method further includes inputting the image into a pre-trained multi-target detection model to process the image. The multi-target detection model is configured to detect multiple types of detection targets including two-dimensional codes. The method further comprises at least one of the following steps: based on a detection frame of a two-dimensional code in an image obtained from a multi-target detection model, cutting out an image area of the determined detection frame of the two-dimensional code from the image, and identifying the two-dimensional code from the cut image area to read information about an object recorded by the two-dimensional code; or, based on the actual size of the two-dimensional code, the size of the two-dimensional code in the image and the position of the two-dimensional code in the image obtained from the multi-target detection model, the actual position of the two-dimensional code relative to the imaging device used for shooting the image is determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of detection, and more particularly to detection methods, detection devices, object monitoring systems, and associated computing devices and non-transitory storage media, where the object can be, for example but not limited to, an insect such as a bumblebee. Background Art

[0002] A two-dimensional code is a form of information encoding that typically uses black and white patterns to represent the concept of a "0" and "1" bitstream underlying computer logic. A two-dimensional code recording information about an object can be attached to the object for object detection. For example, in a common shopping scenario, a two-dimensional code recording product information can be attached to a product, and a cashier can scan the two-dimensional code with a barcode scanner to read the product information. Summary of the Invention

[0003] According to a first aspect of the present disclosure, a detection method is provided. The method includes obtaining an image that includes a two-dimensional code for attaching to an object and recording information about the object. The method further includes inputting the image into a pre-trained multi-object detection model to process the image, the multi-object detection model being configured to detect multiple classes of detection targets, the multiple classes of detection targets including two-dimensional codes. The method further includes at least one of the following: cropping an image region of the determined detection box of the two-dimensional code from the image based on the detection box of the two-dimensional code in the image obtained from the multi-object detection model, and identifying the two-dimensional code from the cropped image region to read the information about the object recorded by the two-dimensional code; or determining the actual position of the two-dimensional code relative to an imaging device used to capture the image based on the actual size of the two-dimensional code, the size of the two-dimensional code in the image, and the position of the two-dimensional code in the image obtained from the multi-object detection model.

[0004] According to a second aspect of the present disclosure, a detection device is provided. The device includes: an image acquisition module configured to acquire an image, the image including a two-dimensional code for attaching to an object and recording information about the object; a target detection module configured to input the image into a pre-trained multi-target detection model to process the image, the multi-target detection model being configured to detect multiple types of detection targets, the multiple types of detection targets including the two-dimensional code; and at least one of a two-dimensional code recognition module and a position determination module. The two-dimensional code recognition module is configured to crop an image region of the determined detection frame of the two-dimensional code from the image based on the detection frame of the two-dimensional code in the image obtained from the multi-target detection model, and recognize the two-dimensional code from the cropped image region to read the information about the object recorded by the two-dimensional code. The position determination module is configured to determine the actual position of the two-dimensional code relative to an imaging device for capturing the image based on the actual size of the two-dimensional code, the size of the two-dimensional code in the image, and the position of the two-dimensional code in the image obtained from the multi-target detection model.

[0005] According to a third aspect of the present disclosure, an object monitoring system is provided. The system includes: an imaging device configured to capture an image of an object; and a processing device configured to execute the detection method according to the first aspect of the present disclosure to monitor the object based on the image captured by the imaging device.

[0006] According to a fourth aspect of the present disclosure, a computing device is provided. The computing device includes: one or more processors; and a memory storing computer-executable instructions, the computer-executable instructions, when executed by the one or more processors, causing the one or more processors to execute the detection method according to the first aspect of the present disclosure.

[0007] According to a fifth aspect of the present disclosure, a non-transitory storage medium storing computer-executable instructions is provided, the computer-executable instructions, when executed by a computer, causing the computer to execute the detection method according to the first aspect of the present disclosure.

[0008] Other features and advantages of the present disclosure will become clearer through the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. Description of the Drawings

[0009] From the following description of the embodiments of the present disclosure shown in conjunction with the accompanying drawings, the foregoing and other features and advantages of the present disclosure will become clear. The drawings are incorporated herein and form a part of the specification, further for explaining the principles of the present disclosure and enabling those skilled in the art to make and use the present disclosure. Among them:

[0010] Figure 1Shows a flowchart of a detection method according to some embodiments of the present disclosure;

[0011] Figure 2A and Figure 2B Exemplarily shows a bear peak image captured by an imaging device, with a two-dimensional code attached to the back of each bear peak.

[0012] Figure 3 Exemplarily shows an image of a bear peak nest captured by an imaging device.

[0013] Figure 4 Exemplarily shows a training image for a multi-object detection model according to some embodiments of the present disclosure;

[0014] Figure 5 Exemplarily shows the detection result of a multi-object detection model according to some embodiments of the present disclosure;

[0015] Figure 6 Exemplarily shows the interpolation result of the image region of the detection frame of the determined two-dimensional code cropped from the image according to some embodiments of the present disclosure;

[0016] Figure 7 Exemplarily shows according to some embodiments of the present disclosure Figure 6 the binarization result of each interpolation result;

[0017] Figure 8 Exemplarily shows a contour extraction method by quadrilateral fitting according to some embodiments of the present disclosure, and a contour extraction method by rectangle fitting according to a comparative example of the present disclosure;

[0018] Figure 9 Schematically shows a grid used to read information about an object recorded by a two-dimensional code from the extracted contour region according to some embodiments of the present disclosure;

[0019] Figures 10 to 13 Schematically shows a process for determining the actual position of a two-dimensional code relative to an imaging device according to some embodiments of the present disclosure;

[0020] Figure 14 Shows a non-limiting example process of applying the detection method according to some embodiments of the present disclosure;

[0021] Figure 15 Shows a schematic block diagram of a detection device according to some embodiments of the present disclosure;

[0022] Figure 16 Shows a schematic block diagram of an object monitoring system according to some embodiments of the present disclosure;

[0023] Figure 17 FIG. 1 shows a schematic block diagram of a computing device in accordance with some embodiments of the present disclosure.

[0024] Note that in the embodiments described below, sometimes the same reference numerals are used commonly between different drawings to denote the same or functionally identical parts, and their repeated description is omitted. In some cases, similar reference numerals and letters are used to denote similar items, and thus once an item is defined in one drawing, it is not necessary to discuss it further in subsequent drawings.

[0025] For ease of understanding, the positions, sizes, and ranges, etc. of the respective structures shown in the drawings and the like sometimes do not represent the actual positions, sizes, and ranges, etc. Therefore, the present disclosure is not limited to the positions, sizes, and ranges, etc. disclosed in the drawings and the like. DETAILED DESCRIPTION

[0026] Various exemplary embodiments of the present disclosure will be described in detail below with reference to the drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0027] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present disclosure, its application, or uses. That is, the structures and methods herein are shown in an exemplary manner to illustrate different embodiments of the structures and methods of the present disclosure. However, those skilled in the art will understand that they merely illustrate exemplary ways in which the present disclosure can be implemented, rather than exhaustive ways. In addition, the drawings are not necessarily drawn to scale, and some features may be enlarged to show details of specific components.

[0028] In addition, techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0029] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary, and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0030] In some scenarios, the object to which the QR code is attached may be at a relatively far distance, or the QR code may be very small (e.g., because the object to which the QR code is attached is small, or the QR code is required to be unobtrusive, etc.), or the object to which the QR code is attached may be in a moving state, resulting in various orientations of the QR code, making it difficult to recognize the QR code. For example, in an insect detection scenario, a QR code recording relevant information of the insect (e.g., a unique number of the insect, etc.) can be set on the insect (e.g., a bumblebee, etc.) to mark the insect for tracking and detecting its activities. By capturing an image of the insect with an imaging device (e.g., a camera), the insect can be identified by recognizing the QR code in the image. Since the body size of the insect is small and the attachment of the QR code preferably does not affect the activities of the insect itself, the QR code used to mark the insect often has a very small size. For example, for a bumblebee, the size of its QR code is usually about 3 millimeters × 3 millimeters, and such a QR code may only occupy approximately 30 pixels × 30 pixels in the captured image. In addition, due to the activities of the insect, the QR code often fails to be in the optimal detection orientation relative to the imaging device (the optimal detection orientation is usually facing the imaging device directly), so such a QR code may be distorted in the captured image. For the above reasons, it is very difficult to recognize small moving QR codes.

[0031] Taking a 5×5 QR code as an example, the QR code in the image can be recognized through the following process: Binarize the image into a black-and-white image; Search for all closed contours in the black-and-white image and select the closed contours that are approximately rectangular from them; Fit the selected closed contours with a rectangle and draw a 7×7 grid based on the outer edge of the fitted rectangle, so as to use the central 5×5 matrix to decode the QR code. However, the step of "searching for all closed contours in the black-and-white image and selecting the closed contours that are approximately rectangular from them" is very inefficient because in a complex image, there are a large number of regions with approximately rectangular shapes, and a large part of them are not real QR codes. Therefore, the above steps waste a lot of time and computing power on non-expected approximately rectangular regions, and the detection speed is very slow, resulting in difficulty in achieving real-time detection and tracking of the QR code (and the object to which it is attached).

[0032] Therefore, the present disclosure provides a detection method that can achieve real-time detection and tracking of the QR code (and the object to which it is attached) with high efficiency and high accuracy. The detection methods according to various embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the actual method may also include other additional steps, but in order to avoid obscuring the key points of the present disclosure, these other additional steps are not discussed herein and are not shown in the drawings.

[0033] It should also be understood that although many examples in this text are described with insects or more specifically bumblebees as the object, this is not restrictive. The detection method disclosed herein can be used to detect any two-dimensional code, is particularly useful for detecting small and / or moving two-dimensional codes, and can further be used to detect the object to which the two-dimensional code is attached, as well as the same-kind objects and associated objects of that object.

[0034] The same-kind objects described herein refer to objects belonging to the same category. The associated objects described herein refer to objects having a certain association. For example, when the object to which the two-dimensional code is attached is an insect, the same-kind objects can be insects without the attached two-dimensional code. Insects without the attached two-dimensional code can include at least one of the following: insects that originally had a two-dimensional code attached but the code was soiled, damaged, or fell off, offspring insects reproduced from the original insects, and foreign insects that were added to the insect community after the two-dimensional code was attached to the original insects. In addition, the associated objects of the insects include at least one of the following: insect forms at various stages of the growth period (e.g., insect eggs, insect larvae, insect cocoons, or pupae, etc.), insect nests, and / or their constituent structures.

[0035] Figure 1 There is shown a detection method 100 (which may be simply referred to as method 100 herein) according to some embodiments of the present disclosure, which includes steps S102 and S104, and may further include at least one of steps S106 and S108.

[0036] At step S102, an image is acquired, the image including a two-dimensional code for attachment to an object and recording information about the object.

[0037] For example, the image can be captured by an imaging device such as a camera. The acquired image can include a single image or multiple images, or can include one or more frames in an image frame sequence (e.g., from a video). For non-limiting illustrative purposes, Figure 2A and Figure 2B There are shown example images, each including multiple bumblebees, with a two-dimensional code attached to the back of each bumblebee, the two-dimensional code recording information about the bumblebee (e.g., the unique number of the bumblebee, etc.). From Figure 2A It can be seen that the proportion of the image size of the two-dimensional code in the entire image is very small and difficult to identify with a general barcode scanner. Figure 2B Compared with Figure 2A It can be a bumblebee image captured at a position closer to the bumblebee. Although Figure 2B the proportion of the image size of the two-dimensional code in the entire image in Figure 2A has increased relative to Figure 2BIt can be seen that since the bumblebees are in different postures, the QR code is in different orientations relative to the imaging device, which causes the shape of the QR code in the captured image to deviate from its actual shape. Specifically, it can be manifested that the angle between adjacent edges of the shape of the QR code with an actual rectangular shape in the captured image is not equal to 90°.

[0038] At step S104, the image is input into a pre-trained multi-object detection model to process the image. The multi-object detection model is configured to detect multiple classes of detection targets, and the multiple classes of detection targets include QR codes.

[0039] If the above method of selecting rectangular closed contours is used to identify the QR code, it is very likely to fail in the case where the QR code is displayed unclearly, incompletely, and / or deformed in the image. In contrast, if a multi-object detection model is used to detect the QR code, then in the training stage of the multi-object detection model, images containing QR codes with various degrees of clarity, integrity, and deformation can be used as training samples, so that the trained model can have better detection performance even for unclearly, incompletely, and / or deformed QR codes in the image. Moreover, the multi-object detection model no longer needs to "search all closed contours in the black-and-white image and select the closed contours approximated to rectangles", and will not waste time on incorrect rectangular areas, so it has an improved detection speed, which is beneficial to realizing real-time detection and tracking.

[0040] In particular, in some embodiments, the multiple classes of detection targets may further include at least one of an object and an associated object. For example, for a bumblebee as an object, refer to Figure 3 , the associated objects of the bumblebee may include cocoons, wax pots, etc. If the above method of selecting rectangular closed contours is used to identify the QR code and then identify the bumblebee to which it is attached, the bumblebee that originally had a QR code attached but the QR code is dirty, damaged, or fallen off will become unidentifiable. Moreover, in an actual bumblebee colony, offspring bumblebees are constantly being born, and there may also be foreign bumblebees joining the colony, and these bumblebees are not marked with QR codes, so they are also unidentifiable. Therefore, the amount of data that can be obtained only by using the above method of selecting rectangular closed contours to identify the QR code is very limited. In addition, in addition to individual bumblebees, as Figure 3 shows, in a bumblebee colony, there are also many other important visual information, such as the number of cocoons, wax pots, or other honeycomb structures in the honeycomb. These important information reflect the evolution of the bumblebee colony, but these associated objects are often not marked with QR codes and cannot be obtained by identifying the QR codes attached to the objects. Therefore, when the multiple classes of detection targets further include at least one of an object (bumblebee) and an associated object (cocoon and wax pot), more important information can be obtained in addition to the information contained in the QR code.

[0041] The multi-object detection model can be configured to output one or more of the categories, positions, and detection boxes of the objects detected in the image. To improve the detection accuracy, multi-task learning that performs category regression, position regression, and detection box regression on all detected objects can be required during the training phase. To improve the detection speed and save computing power, in the inference phase, it can be selected not to calculate the detection boxes of a certain or certain detected objects. For example, the detection box of the QR code can be calculated without calculating the detection boxes of the object and the associated object. Compared with the detection box, the calculation of the category and position of the object consumes less time and computing power. Therefore, it can be selected to calculate the categories and positions of all detected objects in the inference phase.

[0042] In some embodiments, method 100 includes determining the position of the QR code in the image obtained from the multi-object detection model as the position of the object in the image when the image further includes an object attached with a QR code. In this case, the multi-object detection model only needs to calculate the category of the object without calculating its position in the inference phase. In some embodiments, when the multi-class detected objects further include an object, method 100 may also include obtaining the position of the object in the image from the multi-object detection model when the image further includes an object attached with a QR code. In some embodiments, when the multi-class detected objects further include an object, method 100 may include obtaining the position of the object without a QR code attached in the image from the multi-object detection model when the image further includes an object without a QR code attached. In these cases, the multi-object detection model calculates the category and position of the object in the inference phase. In some embodiments, when the multi-class detected objects further include an associated object of the object and the associated object is not attached with a QR code, method 100 may include obtaining the position of the associated object in the image from the multi-object detection model when the image further includes the associated object. In this case, the multi-object detection model calculates the category and position of the associated object in the inference phase.

[0043] Detecting a QR code with a multi-object detection model can achieve improved detection accuracy compared to detecting a QR code with a single-object detection model. On the one hand, the training sample size of the multi-object detection model is large, and the problem of insufficient training data for a certain task (for example, the QR code detection task) can also be trained with the help of other tasks, and improved generalization performance can be achieved. On the other hand, especially when there is a correlation between the multi-class detected objects (for example, QR code, object, associated object), different detection tasks can be better jointly learned to mine the shared information of different detection tasks. For a certain task (for example, the QR code detection task), it can be determined whether the learned features are really effective through other tasks, difficult-to-learn features of this task can be learned through other tasks, or this task can be made to focus on the information expressions that other tasks also focus on.

[0044] For the purpose of non-limiting illustration,Figure 4 Exemplarily shown is a training image for training a multi-object detection model, which is labeled by bounding boxes around QR codes (tags), bumblebees, broods, and pots in the training image. These labeled training data can be used to train, for example, the MMDetection object detection model developed by OpenMMLab or an object detection model modified based on this model. It can be understood that other suitable multi-object detection models are also feasible. Accordingly, Figure 5 Exemplarily shown are the detection results of the multi-object detection model obtained by training for 2000 generations with a learning rate of 0.002, including the detection boxes and confidence levels of QR codes, bumblebees, broods, and pots. Figure 5 Targets of different categories can already be clearly distinguished, and by further optimizing the model training parameters and the number of training generations, the detection ability of the model can be further improved.

[0045] Return reference Figure 1 , at step S106, based on the detection box of the QR code in the image obtained from the multi-object detection model, the image region of the determined detection box of the QR code is cropped from the image, and the QR code is recognized from the cropped image region to read the information about the object recorded by the QR code. For example, refer to Figure 5 , a part of the image region can be cropped according to the detection box calculated by the multi-object detection model for the QR code, and this part of the image region is used to recognize the QR code.

[0046] In some embodiments, recognizing the QR code from the cropped image region to read the information about the object recorded by the QR code includes: performing binarization on the cropped image region; extracting the contour region of the binarized cropped image region; and reading the information about the object recorded by the QR code from the extracted contour region.

[0047] The image region cropped according to the detection box of the QR code may have a very low resolution. As mentioned above, for example, for a bumblebee, the size of its QR code is usually about 3 mm × 3 mm, and such a QR code may only occupy approximately 30 pixels × 30 pixels in the captured image. The low resolution may reduce the recognition accuracy of the QR code. Therefore, in some examples, before performing binarization on the cropped image region, interpolation can be performed on the cropped image region to improve the resolution of the cropped image region, and then binarization is performed on the interpolated cropped image region. The interpolation can be configured to increase the resolution of the cropped image region by 1 to 7 times, or by 2 to 5 times, or by 2 to 3 times. For example, when the cropped image region is 30 pixels × 30 pixels, it may be appropriate to enlarge it to 100 pixels × 100 pixels through interpolation.

[0048] For example, any suitable interpolation algorithm in the OpenCV library can be used for interpolation. For the purpose of non-limiting illustration, reference is made to Figure 6 , which exemplarily shows the original version (RAW) of the image region cropped according to the detection box calculated for the two-dimensional code by the multi-object detection model, the interpolated version (INTER_NEAREST) using the nearest neighbor interpolation method, the interpolated version (INTER_LINEAR) using the bilinear interpolation method, the interpolated version (INTER_AREA) using resampling using the pixel region relationship, the interpolated version (INTER_CUBIC) using the cubic interpolation method based on the 4×4 pixel neighborhood, and the interpolated version (INTER_LANCZOS4) using the Lanczos interpolation method based on the 8×8 pixel neighborhood. The inventors have found that, in the case of using the interpolation algorithm alone, INTER_CUBIC and INTER_LANCZOS4 among the above five interpolation algorithms can each have relatively high recognition accuracy, but also consume a certain time cost; in the case of using the interpolation algorithms in combination, the combination of INTER_LANCZOS4 and INTER_AREA in the combination of the above five interpolation algorithms can have relatively high recognition accuracy and low time cost.

[0049] In many cases, for example, when the lighting conditions (such as lighting angle and / or lighting intensity) during image capture are complex, the brightness at different positions in the image varies. Therefore, it may be inappropriate to binarize all pixels in the image region with a single threshold. For example, after binarization, the brighter part may contain too many white pixels and the darker part may contain too many black pixels. Therefore, an adaptive threshold can be used to binarize the image region. Specifically, in some embodiments, performing binarization on the cropped image region may include: for each pixel in the cropped image region, determining a threshold for binarizing each pixel based on the values of the pixels within a predetermined range around each pixel, and binarizing each pixel based on the determined threshold. For example, the pixels within a predetermined range around each pixel may include 7 to 15 pixels, or include 9 to 13 pixels, such as 11 pixels. In some examples, the threshold for binarizing each pixel may be determined based on the average value or weighted average value of the pixels within a predetermined range around each pixel. For example, a constant may be determined according to the lighting intensity during image capture, and then the threshold may be determined by subtracting this constant from the average value or weighted average value of the pixels within a predetermined range around each pixel. Exemplarily, when the lighting intensity during image capture is greater than the preset threshold lighting intensity, the constant may be positive; when the lighting intensity during image capture is less than the preset threshold lighting intensity, the constant may be negative; when the lighting intensity during image capture is equal to the preset threshold lighting intensity, the constant may be zero. Also, for example, the weights of the pixels at different angular directions relative to each pixel within a predetermined range around each pixel may be determined according to the lighting angle during image capture, and then the threshold may be determined as the weighted average value of the pixels within a predetermined range around each pixel. For the purpose of non-limiting illustration, reference Figure 7 , exemplarily shows the results of binarizing the original version of the image region cropped according to the detection frame of the two-dimensional code shown in Figure 6 and the interpolated versions obtained by processing with the above five interpolation algorithms respectively.

[0050] Conventionally, the contour region of the binarized cropped image region can be extracted by rectangle fitting. However, since the two-dimensional code is often not directly facing the field of view of the imaging device, the shape of the two-dimensional code in the image is often not a strictly rectangle, that is, the angle between adjacent edges may not be equal to 90°. Reference Figure 8The rectangular fitting process shown in the upper half. If the minimum bounding shape used is a rectangle, the read grid may deviate from the actual grid of the QR code, where the shaded cells indicate incorrect reads. For this reason, in some embodiments, extracting the contour region of the cropped image region that has been binarized includes: fitting the contour region of the cropped image region that has been binarized with a quadrilateral; using the obtained quadrilateral as the contour region extracted for the cropped image region that has been binarized. In particular, the angles of the respective corners of the quadrilateral used to fit the contour region may not be fixed. In particular, the distance between each side of the quadrilateral used to fit the contour region and the contour region can be restricted within a predetermined distance range. For example, the predetermined distance range can include 5 to 15 pixels, such as 10 pixels, and within this range, the contour region can be more accurately fitted with a quadrilateral. Figure 8 The lower half shows the quadrilateral fitting process, and the read grid more accurately corresponds to the actual grid of the QR code.

[0051] Furthermore, reading the information about the object recorded by the QR code from the extracted contour region may include: according to the m×n matrix corresponding to the QR code, applying an m×n grid to the obtained quadrilateral; for each cell in the m×n grid, determining the binarized value of each cell based on the value of the pixels falling into each cell; determining the m×n matrix corresponding to the QR code based on the binarized values of each cell in the m×n grid; decoding the m×n matrix to read the information about the object recorded by the QR code, where m and n are positive integers. In some embodiments, the m×n grid applied to the obtained quadrilateral can be obtained by connecting the m equal division points of the first pair of opposite sides of the quadrilateral and the n equal division points of the second pair of opposite sides of the quadrilateral. For the purpose of non-limiting illustration, referring to Figure 9 , assuming that the QR code corresponds to a 5×5 matrix, a 5×5 grid can be applied to the obtained quadrilateral by connecting the 5 equal division points of the two horizontal sides and the 5 equal division points of the two vertical sides. For example, the value of the central pixel of each cell can be determined as the binarized value of each cell, or the mode of all the pixel values of each cell can be determined as the binarized value of each cell, or other suitable methods can also be adopted to determine the binarized value of each cell.

[0052] Returning to the reference Figure 1 , at step S108, based on the actual size of the QR code, the size of the QR code in the image, and the position of the QR code in the image obtained from the multi-object detection model, determine the actual position of the QR code relative to the imaging device used to capture the image.

[0053] The position of the QR code in the image obtained from the multi-object detection model can be, for example but not limited to, the center position of the QR code in the image. It should be understood that for an object such as a QR code, the position of the object in the image obtained from the multi-object detection model may depend on how the position of the object is labeled when training the multi-object detection model. In this document, the position of the object in the image can be referred to as the image position of the object, and the size of the object in the image can be referred to as the image size of the object.

[0054] The size of the QR code in the field of view of the imaging device, i.e., the image size of the QR code, roughly reflects the actual distance between the QR code and the imaging device. The farther the QR code is from the imaging device, the smaller the QR code appears in the image captured by the imaging device. Additionally, since the QR code is usually not directly facing the imaging device, the angle between the adjacent edges of the QR code in the image is usually not equal to 90°, so the difference in the angles of the four corners of the QR code in the image can reflect the angle between the normal direction of the QR code and the direction of the line connecting the QR code and the imaging device. The size of the QR code in the image can be directly measured, or reflected by the size of the detection box calculated by the multi-object detection model for the QR code, or determined by the quadrilateral obtained by fitting. However, the angles of the four corners of the QR code in the image are difficult to determine through the detection box and instead need to be directly measured or determined by the quadrilateral obtained by fitting. Using the quadrilateral obtained by fitting to determine these angles may be more advantageous, but this results in having to determine the actual position of the QR code in space only after the QR code is recognized (at least after the quadrilateral fitting).

[0055] Therefore, based on the Euclidean geometric space relationship and the pinhole imaging theory, and assuming that the angular extent of the QR code relative to the imaging device in space is always a relatively small quantity, the present disclosure proposes a method for estimating the actual position of a rectangular object such as a QR code relative to the imaging device. This method can be implemented based on the actual size of the QR code, the size of the QR code in the image, and the position of the QR code in the image obtained from the multi-object detection model. Given the above assumption, this estimation method is particularly advantageous for determining the actual position of a distant or small object relative to the imaging device.

[0056] Specifically, in some embodiments, determining the actual position of the two-dimensional code relative to the imaging device includes: using a reference image containing a reference object captured by the imaging device, based on the image size and image position of the reference object in the reference image, and the included angle of the reference object relative to the imaging device determined by the actual size of the reference object and the actual position of the reference object relative to the imaging device, obtaining, through fitting, a proportionality coefficient between the included angle of the reference object relative to the imaging device and the image size of the reference object as a first function of the image position of the reference object, and based on the image position of the reference object and the actual position of the reference object relative to the imaging device, obtaining, through fitting, the elevation angle of the line connecting the reference object and the imaging device relative to the projection of the line on the plane where the imaging device is located as a second function of the image position of the reference object; determining the actual distance (L) of the two-dimensional code relative to the imaging device based on the actual size of the two-dimensional code and the included angle of the two-dimensional code relative to the imaging device determined by the first function, the position of the two-dimensional code in the image, and the size of the two-dimensional code in the image, determining the elevation angle (θ) of the line connecting the two-dimensional code and the imaging device relative to the projection of the line on the plane where the imaging device is located based on the second function and the position of the two-dimensional code in the image, and determining the azimuth angle (φ) of the projection of the line on the plane where the imaging device is located based on the position of the two-dimensional code in the image, so as to obtain the spherical coordinate representation (L, θ, φ) of the actual position of the two-dimensional code relative to the imaging device. In some examples, the fitting of the first function and the second function can be performed using a binary polynomial fit and using the polar coordinate representation of the image position of the reference object, and the first function and the second function obtained through fitting can be further expressed as functions of the rectangular coordinate representation of the image position of the reference object. Of course, other suitable fitting methods are also feasible. For example, the spherical coordinate representation (L, θ, φ) of the actual position of the two-dimensional code relative to the imaging device can also be converted into a rectangular coordinate representation (X, Y, Z). In addition, the size of the two-dimensional code in the image can be obtained by direct measurement, or determined by the size of the quadrilateral obtained through fitting, or approximately determined by the size of the detection frame of the two-dimensional code in the image obtained from the multi-object detection model. These sizes and the actual size of the two-dimensional code can all be known. The size information can include, for example, at least one of length, width, and area.

[0057] For the purpose of non-limiting illustration, the following will be combined with Figures 10 to 13 Describe the principle and process for determining the actual position of the two-dimensional code relative to the imaging device according to the present disclosure. As Figure 10As shown in the figure, taking the location of the imaging device as the origin O, a rectangular coordinate system (X, Y, Z) and a spherical coordinate system (L, θ, φ) are constructed. The plane where the imaging device is located is the XY plane. The actual distance between the QR code and the imaging device is L. The elevation angle of the line connecting the QR code and the imaging device with respect to the projection of this line on the plane where the imaging device is located is θ, and the azimuth angle of the projection of the line connecting the QR code and the imaging device on the plane where the imaging device is located is φ. φ can be directly determined from the coordinates (x, y) of the position of the QR code in the image, that is, φ = arctan(x / y). The solution of L and θ is relatively more complex.

[0058] Reference Figure 11 , the QR code can be represented as a rectangle ABCD with a width of d 1 and a length of d 2 in the real space. A square A'B'C'D' can be defined in the rectangle ABCD, whose center coincides with the center of the rectangle ABCD and whose side length is equal to d 2 . The rectangle ABCD appears as a quadrilateral rather than necessarily a rectangle in the field of view of the imaging device, and the square A'B'C'D' will be deformed in the same proportion as the rectangle ABCD. The scaling ratio of the quadrilateral (referred to as the first quadrilateral) corresponding to the rectangle ABCD in the field of view of the imaging device and the quadrilateral (referred to as the second quadrilateral) corresponding to the square A'B'C'D' in the field of view of the imaging device in the direction of the side CD of the original rectangle ABCD is fixed, that is, d 1 :d 2 . According to this scaling ratio, the second quadrilateral can be drawn from the first quadrilateral (corresponding to the QR code in the image) in the field of view of the imaging device. Subsequent estimations can be carried out according to the projection of the square A'B'C'D' into the second quadrilateral in the field of view of the imaging device.

[0059] An inscribed circle is made in the square A'B'C'D', and this inscribed circle is projected into an inscribed ellipse of the second quadrilateral in the field of view of the imaging device. Reference Figure 12 , the major axis of this inscribed ellipse is a and the minor axis is b. The opening angle of the inscribed ellipse with respect to the imaging device is γ (γ << 1), and the angle between the normal direction of the inscribed ellipse and the direction of the line connecting the center of the inscribed ellipse and the imaging device is δ. According to the spatial geometric relationship, there is L×γ = a, so L = a / γ. Therefore, as long as a and γ are known, the distance L between the center of the inscribed ellipse and the imaging device can be obtained. Since this inscribed ellipse is projected from the inscribed circle, a = d 2 . Thus, only by calculating the opening angle γ can the distance L be known. However, in practical applications, it is often difficult to directly measure the opening angle γ between the QR code and the imaging device, so it is necessary to indirectly calculate through the information in the image captured by the imaging device.

[0060] When the position of the target remains unchanged, the included angle γ of the target relative to the imaging device is proportional to the size (e.g., length or width) of the area occupied by the target in the image of the imaging device. Taking the width w of the area occupied by the target in the image of the imaging device as an example, there is γ = k × w. Thus, when the proportionality coefficient k is known, the included angle γ can be calculated by reading the width of the target in the image. Note that it can be understood that the included angle γ mentioned here as an example is a planar included angle, and this method can be similarly applied to the solid included angle ΔΩ of the inscribed ellipse relative to the imaging device, which is also proportional to the size (e.g., area) of the area occupied by the target in the image of the imaging device, where ΔΩ = L 2 ×A, A = πab, so there is L = (ΔΩ / πab) 1 / 2 , which will not be elaborated here.

[0061] However, due to the existence of image distortion, the proportionality coefficient k is different at different positions in the image, that is, the proportionality coefficient k is a function of the image position (x, y), k = g(x, y). Different imaging optical systems of the imaging device (e.g., different lens models) result in different degrees of image distortion, and it is difficult to directly obtain the analytical solution of the function k = g(x, y) based on optical principles. Therefore, the present disclosure proposes to use a calibration and fitting method to solve the function k = g(x, y) for the imaging device. Additionally, from Figure 12 it can be clear that the elevation angle θ is also a function of the image position (x, y), θ = h(x, y). The function θ = h(x, y) can be solved together during the process of solving the function k = g(x, y).

[0062] As a non-limiting specific example, a calibration paper can be printed (as shown in the left half of Figure 13 ), which includes an array of dots with a diameter D of 3 mm and a spacing B of 100 mm. Each dot can be used as a reference object. Here, the dots, D, and B are all exemplary and not restrictive. The dots can be replaced with any other suitable shape, for example, a shape that is convenient for measurement or a shape that is the same as or similar to the target to be measured (such as a QR code). The value of D can be determined based on the actual size of the target (such as a QR code), for example, D can be equal to the side length of the QR code. The value of B can be determined considering the fitting accuracy. The smaller B is, the more reference objects can be arranged on the calibration paper, and the more data points are available for fitting. As Figure 13As shown in the right half, adjust the focal length of the imaging device and place the calibration paper at the focal plane directly in front of the imaging device, such that the center point of the calibration paper is located at the center of the field of view of the imaging device. Measure the actual distance H from the calibration paper to the imaging device in the space. In this example, the actual distance from the center point of the calibration paper to the imaging device is equal to H. Each point on the calibration paper can be represented as the point (i, j) in the i-th row and j-th column starting from the center point of the calibration paper. For example, as Figure 13 shown, the center point of the calibration paper is (i = 0, j = 0), and the four adjacent points above, below, left, and right of it are (i = +1, j = 0), (i = -1, j = 0), (i = 0, j = -1), and (i = 0, j = +1) respectively. Capture an image of the calibration paper using the imaging device as a reference image, and measure the number of pixels w of the major axis of each point (i, j) in the reference image ij (due to distortion, the points may have changed from circular points to elliptical points). Since the actual size (D = 3 mm) of each point (i, j) on the calibration paper and its actual position in the space are known, the angular aperture of the point (i, j) relative to the imaging device can be calculated as the ratio of the actual size of the point (i, j) to the actual distance from the point (i, j) to the imaging device, specifically as

[0063]

[0064] Therefore, at the image coordinates (x ij , y ij ) where the point (i, j) is located, the scale factor can be calculated as the ratio of the angular aperture of the point (i, j) relative to the imaging device to the number of pixels w of the major axis of the point (i, j), specifically as ij the ratio, specifically as

[0065]

[0066] In addition, the elevation angle can be calculated as the arctangent of the ratio of the actual distance from the center point of the calibration paper to the imaging device to the actual distance from the center point of the calibration paper to the point (i, j), specifically as

[0067]

[0068] Of course, since the actual distance from the center point of the calibration paper to the imaging device, the actual distance from the point (i, j) to the imaging device, and the actual distance from the center point of the calibration paper to the point (i, j) are all known, the elevation angle θ can also be calculated using arcsin and arccos ij .

[0069] Calculate the scale factor k ij and the elevation angle θ ij for each point (i, j) in the reference image, and then respectively with the image coordinates (x ij , yij ) Perform fitting to obtain the fitting results of the functions k = g(x, y) and θ = h(x, y). For example, considering that the image distortion may be centrosymmetric, the image coordinates (x ij , y ij ) can be converted to polar coordinates (r ij , α ij ), where r ij is the distance from the point (i, j) in the reference image to the image center as the origin, and α ij is the rotation angle of the line connecting the point (i, j) in the reference image and the image center. Then, a binary polynomial fitting is performed with [r, r -1 , cosα, sinα] as the dependent variables.

[0070] After obtaining the functions k = g(x, y) and θ = h(x, y) by fitting, (L, θ, φ) can be obtained as follows.

[0071]

[0072] θ = h(x, y)

[0073]

[0074] where a is the length of the major axis of the inscribed ellipse, and w a is the number of pixels of the major axis of the inscribed ellipse. The scaling ratio relationship between the rectangle ABCD and the inscribed ellipse corresponds to the scaling ratio relationship between the image size of the rectangle ABCD and the image size of the inscribed ellipse. Therefore,

[0075]

[0076] where c is the actual size of the rectangle ABCD, i.e., the QR code, w c is the image size of the rectangle ABCD, i.e., the QR code, and (x, y) is the image position of the rectangle ABCD, i.e., the QR code. As mentioned above, the image size of the QR code can be directly measured, or determined by the size of the quadrilateral obtained by fitting, or approximated by the size of the detection frame of the QR code.

[0077] For ease of representation and subsequent analysis, the spherical coordinates (L, θ, φ) can be converted to Cartesian coordinates as follows.

[0078] X = Lcosθsinφ

[0079] Y = Lcosθcosφ

[0080] Z = Lsinθ

[0081] The average error between the spatial coordinates of the QR code relative to the imaging device obtained by the above estimation method and the manually measured spatial coordinates is very small, and experiments have found that it can be within ±1.648 cm.

[0082] After determining the actual position of the QR code, the actual positions of the object and the associated object can be inferred based on the actual position of the QR code. For example, in the case where the image further includes an object attached with a QR code, the actual position of the determined QR code relative to the imaging device can be used as the actual position of the object attached with the QR code relative to the imaging device. In the case where the multi-class detection target further includes an object, method 100 may also include determining the actual position of the object relative to the imaging device based on the actual position of the determined QR code relative to the imaging device, the position of the QR code in the image obtained from the multi-object detection model, and the position of the object in the image. In the case where the multi-class detection target further includes an object, method 100 may further include, in the case where the image further includes an object not attached with a QR code, determining the actual position of the object not attached with a QR code relative to the imaging device based on the actual position of the determined QR code relative to the imaging device, the position of the QR code in the image obtained from the multi-object detection model, and the position of the object not attached with a QR code in the image. In the case where the multi-class detection target further includes an associated object of the object, method 100 may further include, in the case where the image further includes an associated object, determining the actual position of the associated object relative to the imaging device based on the actual position of the determined QR code relative to the imaging device, the position of the QR code in the image obtained from the multi-object detection model, and the position of the associated object in the image.

[0083] The actual sizes of the object and the associated object can also be inferred based on the actual size of the QR code. Specifically, in the case where the multi-class detection target further includes an object, method 100 may include, in the case where the image further includes an object attached with a QR code, determining the actual size of the object based on the actual size of the QR code, the size of the QR code in the image, and the size of the object in the image. Method 100 may also include, in the case where the image further includes an object not attached with a QR code, determining the actual size of the object not attached with a QR code based on the actual size of the QR code, the size of the QR code in the image, and the size of the object not attached with a QR code in the image. Similarly, the size of the object in the image can be directly measured or approximated by the size of the detection box of the object in the image obtained from the multi-object detection model.

[0084] In addition, when the multiple types of detection targets further include associated objects of the object, the method may further include, when the image further includes associated objects, determining the actual size of the associated objects based on the actual size of the two-dimensional code, the size of the two-dimensional code in the image, and the size of the associated objects in the image. Similarly, the size of the associated objects in the image can be directly measured, or can be approximately determined by the size of the detection frames of the associated objects in the image obtained from the multi-target detection model.

[0085] One or more of the steps S106 for identifying the two-dimensional code, the step S108 for determining the actual position of the two-dimensional code, and the steps for determining the actual sizes of the object and the associated objects can be executed according to the detection requirements. When executing multiple ones of the steps S106 for identifying the two-dimensional code, the step S108 for determining the actual position of the two-dimensional code, and the steps for determining the actual sizes of the object and the associated objects, the multiple ones can be executed at least partially in parallel. In some embodiments, after successfully identifying the two-dimensional code from the cropped image region, one or more of the step S108 for determining the actual position of the two-dimensional code and the steps for determining the actual sizes of the object and the associated objects can be executed. This can ensure that the basis for calculating the actual position of the two-dimensional code and the actual sizes of the object and the associated objects is valid.

[0086] In addition, when the image includes multiple two-dimensional codes, the recognition of each two-dimensional code in the multiple two-dimensional codes in the image obtained from the multi-target detection model can be processed in parallel using multi-threading; or the determination of the actual position of each two-dimensional code in the multiple two-dimensional codes relative to the imaging device can be processed in parallel using multi-threading based on the actual size of each two-dimensional code in the multiple two-dimensional codes, the size of each two-dimensional code in the image, and the position of each two-dimensional code in the multiple two-dimensional codes in the image obtained from the multi-target detection model. When the image includes multiple objects or associated objects, after obtaining the size of at least one two-dimensional code in the image, the determination of the actual size of each object or associated object in the multiple objects or associated objects can be processed in parallel using multi-threading. In addition, when the image includes multiple objects or associated objects, after obtaining the position and actual position of at least one two-dimensional code in the image, the determination of the actual position of each object or associated object in the multiple objects or associated objects can be processed in parallel using multi-threading.

[0087] For non-limiting illustrative purposes, Figure 14An example process 200 of application method 100 is shown. First, an imaging device captures an image, and then multi-object detection is performed on the image to obtain the categories of the detected objects, their positions in the image, and the detection frames. As described above, the category, position, and detection frame can be calculated for the two-dimensional code, while for the object and the associated object, only the category and position can be calculated without calculating the detection frame, thereby improving the model inference speed. Search for the detected objects. When the category of the detected object is a two-dimensional code, the image region of the detection frame of the determined two-dimensional code is cropped from the image, and then interpolation, binarization using an adaptive threshold, quadrilateral fitting to extract the contour region, and two-dimensional code recognition are performed in sequence. In the case of successful recognition, the actual position of the two-dimensional code in space can be further determined, and then the record information and the actual position of the two-dimensional code are stored, and then the next detected object is searched. In the case of failed recognition, the next detected object is directly searched. In addition, when the category of the detected object is not a two-dimensional code, the category of the detected object and its position in the image are stored, and then the next detected object is searched. When the search ends, all information is saved and it is determined whether to continue capturing images. If images are to be captured continuously, the above process can be repeated for the newly captured images. If images are no longer to be captured, process 200 ends. Since method 100 has a high detection efficiency, the detection results can be output substantially in real time with the image captured by the imaging device, facilitating the real-time detection and tracking of two-dimensional codes, objects, and associated objects. Thus, method 100 can be advantageously applied to fields such as plant protection, product testing, and research in ecology, entomology, zoology, or environmental science.

[0088] The present disclosure also provides a detection device 300 (which may also be simply referred to as device 300 in this article) in another aspect. As Figure 15As shown, the apparatus 300 includes an image acquisition module 302 and an object detection module 304. The image acquisition module 302 is configured to acquire an image that includes a two-dimensional code for attaching to an object and recording information about the object. For example, various embodiments of step S102 of method 100 may be performed by the image acquisition module 302. The object detection module 304 is configured to input the image into a pre-trained multi-object detection model to process the image, and the multi-object detection model is configured to detect multiple classes of detection targets, and the multiple classes of detection targets include two-dimensional codes. For example, various embodiments of step S104 of method 100 may be performed by the object detection module 302. In some embodiments, the apparatus 300 may further include a two-dimensional code recognition module 306. The two-dimensional code recognition module 306 is configured to crop an image region of the determined two-dimensional code detection frame from the image based on the detection frame of the two-dimensional code in the image obtained from the multi-object detection model, and recognize the two-dimensional code from the cropped image region to read the information about the object recorded by the two-dimensional code. For example, various embodiments of step S106 of method 100 may be performed by the two-dimensional code recognition module 306. Additionally or alternatively, in some embodiments, the apparatus 300 may further include a position determination module 308. The position determination module 308 is configured to determine the actual position of the two-dimensional code relative to the imaging device used to capture the image based on the actual size of the two-dimensional code, the size of the two-dimensional code in the image, and the position of the two-dimensional code in the image obtained from the multi-object detection model. For example, various embodiments of step S108 of method 100 may be performed by the position determination module 308. In some embodiments, the apparatus 300 may further include a size determination module (not shown). For example, various embodiments of the steps of method 100 for determining the actual sizes of the object and the associated object may be performed by the size determination module.

[0089] Embodiments of the apparatus 300 are substantially similar to the foregoing embodiments of method 100, and thus will not be described in detail herein. For related parts, reference may be made to the description of the foregoing embodiments of method 100.

[0090] For example, the apparatus 300 may be implemented as a detection device including a processor for executing software modules stored in a memory. These software modules may include one or more of the above-mentioned modules such as the image acquisition module 302, the object detection module 304, the two-dimensional code recognition module 306, the position determination module 308, and the size determination module.

[0091] The present disclosure also provides an object monitoring system. Refer to Figure 16, the object monitoring system 400 includes an imaging device 402 and a processing device 404. The imaging device 402 is configured to capture an image of an object. The processing device 404 is configured to execute the method 100 according to any one of the embodiments of the present disclosure to monitor the object based on the image captured by the imaging device 402. The object monitoring system 400 can be applied to fields such as plant protection, product testing, and research in ecology, entomology, zoology, or environmental science.

[0092] The present disclosure also provides a computing device, which may include one or more processors and a memory storing computer-executable instructions. When the computer-executable instructions are executed by the one or more processors, the one or more processors are caused to execute the method 100 according to any one of the embodiments of the present disclosure. As Figure 17 shown, the computing device 500 may include one or more processors 502 and a memory 504 storing computer-executable instructions. When the computer-executable instructions are executed by the one or more processors 502, the one or more processors 502 are caused to execute the method 100 according to any one of the foregoing embodiments of the present disclosure. The one or more processors 502 may be, for example, the central processing unit (CPU) of the computing device 500. The one or more processors 502 may be any type of general-purpose processor, or may be a processor specifically designed to execute the method 100 according to any one of the embodiments of the present disclosure, such as an application-specific integrated circuit (“ASIC”). The memory 504 may include various computer-readable media accessible by the one or more processors 502. In various embodiments, the memory 504 described herein may include volatile and non-volatile media, removable and non-removable media. For example, the memory 504 may include any combination of the following: random access memory (“RAM”), dynamic RAM (“DRAM”), static RAM (“SRAM”), read-only memory (“ROM”), flash memory, cache memory, and / or any other type of non-transitory computer-readable media. The memory 504 may store instructions that, when executed by the processor 502, cause the processor 502 to execute the method 100 according to any one of the embodiments of the present disclosure.

[0093] The present disclosure also provides a non-transitory storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a computer, the computer is caused to execute the method 100 according to any one of the foregoing embodiments of the present disclosure.

[0094] The present disclosure also provides a computer program product, which may include instructions that, when executed by a processor, can implement the method 100 according to any one of the foregoing embodiments of the present disclosure. The instructions may be any set of instructions that are directly executed by one or more processors, such as machine code, or any set of instructions that are indirectly executed, such as scripts. The instructions may be stored in a target code format for direct processing by one or more processors, or stored in any other computer language, including scripts or collections of independent source code modules that are interpreted on demand or pre-compiled.

[0095] By applying the multi-target detection model to QR code recognition, the present disclosure improves the speed and accuracy of QR code recognition, especially for small and / or moving QR codes, enabling real-time detection for imaging devices with a frame rate of more than 1 fps. At the same time, it also allows obtaining more information other than QR codes, such as information about objects with or without QR codes attached and their associated objects. The present disclosure improves the accuracy and robustness of QR code recognition, especially for small and / or moving QR codes, especially under complex lighting conditions, through interpolation, binarization with an adaptive threshold, and quadrilateral fitting. The present disclosure also accurately estimates the three-dimensional spatial coordinates of a QR code relative to an imaging device from an image of a QR code with a known actual size, especially a small QR code, through an estimation method based on Euclidean geometric space relationships, the small hole imaging theory, and multivariate binomial fitting.

[0096] As used herein, the word "exemplary" means "serving as an example, instance, or illustration", rather than as a "model" to be precisely replicated. Any implementation described herein exemplarily is not necessarily to be construed as preferred or advantageous over other implementations. Moreover, the present disclosure is not limited by any theory expressed or implied in the technical field, background art, summary of the invention, or detailed description.

[0097] In addition, for reference purposes only, terms such as "first", "second", etc. may also be used herein, and thus are not intended to be limiting. For example, unless the context clearly indicates otherwise, words such as "first", "second", and other such numerical words referring to structures or elements do not imply an order or sequence. It should also be understood that when the term "comprising / including" is used herein, it indicates the presence of the stated features, wholes, steps, operations, units, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, units, and / or components and / or combinations thereof.

[0098] In the present disclosure, terms such as "acquire" and similar terms are used in a broad sense to cover all ways of obtaining an object. Therefore, "acquiring an object" includes, but is not limited to, "purchasing / ordering", "preparing / manufacturing", "arranging / setting up", "installing / assembling", "designing / building", "measuring / detecting" the object, etc.

[0099] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0100] In addition, when used in this application, words such as "herein", "above", "below", "hereinafter", "above-mentioned" and words with similar meanings shall refer to the whole of this application rather than any particular part of this application. Furthermore, unless otherwise clearly stated or otherwise understood in the context in which it is used, the conditional language used herein, such as "may", "might", "for example", "such as", etc., generally intends to indicate that certain embodiments include, while other embodiments do not include certain features, elements and / or states. Therefore, such conditional language generally does not intend to imply that one or more embodiments require in any way the features, elements and / or states, or whether to include these features, elements and / or states or to perform these features, elements and / or states in any particular embodiment.

[0101] Those skilled in the art should be aware that the boundaries between the above operations are merely illustrative. Multiple operations can be combined into a single operation, a single operation can be distributed among additional operations, and operations can be performed at least partially overlapping in time. Moreover, alternative embodiments can include multiple instances of a particular operation, and the order of operations can be changed in various other embodiments. However, other modifications, variations and substitutions are also possible. Aspects and elements of all the embodiments disclosed above can be combined in any manner and / or in combination with aspects or elements of other embodiments to provide multiple additional embodiments. Therefore, this specification and the drawings should be regarded as illustrative rather than restrictive.

[0102] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustration only and not for limiting the scope of the present disclosure. The embodiments disclosed herein can be combined arbitrarily without departing from the spirit and scope of the present disclosure. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A detection method, comprising: obtaining an image, the image including a two-dimensional code for attaching to an object and recording information about the object; inputting the image into a pre-trained multi-object detection model to process the image, the multi-object detection model being configured to detect multiple classes of detection targets, the multiple classes of detection targets including two-dimensional codes; and at least one of the following: Based on the detection box of the two-dimensional code in the image obtained from the multi-object detection model, cropping out the image region of the determined detection box of the two-dimensional code from the image, and identifying the two-dimensional code from the cropped image region to read the information about the object recorded by the two-dimensional code; or Based on the actual size of the two-dimensional code, the size of the two-dimensional code in the image, and the position of the two-dimensional code in the image obtained from the multi-object detection model, determining the actual position of the two-dimensional code relative to the imaging device used to capture the image.

2. The detection method according to claim 1, further comprising, when the image further includes the object to which the two-dimensional code is attached: determining the position of the two-dimensional code in the image obtained from the multi-object detection model as the position of the object in the image; or the multiple classes of detection targets further include an object, and wherein the detection method further includes obtaining the position of the object in the image from the multi-object detection model.

3. The detection method according to claim 1, further comprising at least one of the following: the multiple classes of detection targets further include an object, and wherein the detection method includes, when the image further includes the object not attached with the two-dimensional code, further obtaining the position of the object not attached with the two-dimensional code in the image from the multi-object detection model; or the multiple classes of detection targets further include an associated object of the object, the associated object not being attached with a two-dimensional code, and wherein the detection method includes, when the image further includes the associated object, further obtaining the position of the associated object in the image from the multi-object detection model.

4. The detection method according to claim 1, further comprising, when the image includes multiple two-dimensional codes, performing at least one of the following: Based on the detection box of each two-dimensional code in the multiple two-dimensional codes in the image obtained from the multi-object detection model, using multi-thread parallel processing to identify each two-dimensional code in the multiple two-dimensional codes; or Based on the actual size of each two-dimensional code in the multiple two-dimensional codes, the size of each two-dimensional code in the image, and the position of each two-dimensional code in the multiple two-dimensional codes in the image obtained from the multi-object detection model, using multi-thread parallel processing to determine the actual position of each two-dimensional code in the multiple two-dimensional codes relative to the imaging device.

5. The detection method according to claim 1, wherein identifying the two-dimensional code from the cropped image region to read the information about the object recorded by the two-dimensional code includes: performing binarization on the cropped image region; Extract the contour region of the cropped image region after binarization; and Read the information about the object recorded by the two-dimensional code from the extracted contour region.

6. The detection method according to claim 5, wherein, Identifying the two-dimensional code from the cropped image region to read the information about the object recorded by the two-dimensional code includes: Before performing binarization on the cropped image region, interpolate the cropped image region to improve the resolution of the cropped image region; Perform binarization on the interpolated cropped image region.

7. The detection method according to claim 5, wherein, Performing binarization on the cropped image region includes: For each pixel in the cropped image region, determine a threshold for binarizing each pixel based on the values of the pixels within a predetermined range around each pixel, and binarize each pixel based on the determined threshold.

8. The detection method according to claim 5, wherein, Extracting the contour region of the binarized cropped image region includes: Fitting the contour region of the binarized cropped image region with a quadrilateral, wherein the angles of the respective corners of the quadrilateral used for fitting the contour region are not fixed, and wherein the distances between the respective sides of the quadrilateral used for fitting the contour region and the contour region are limited within a predetermined distance range; Take the fitted quadrilateral as the contour region extracted for the binarized cropped image region.

9. The detection method according to claim 8, wherein, Reading the information about the object recorded by the two-dimensional code from the extracted contour region includes: According to the m×n matrix corresponding to the two-dimensional code, apply an m×n grid to the fitted quadrilateral; For each grid in the m×n grid, determine the binarized value of each grid based on the values of the pixels falling into each grid; Determine the m×n matrix corresponding to the two-dimensional code based on the binarized values of each grid in the m×n grid; Decode the m×n matrix to read the information about the object recorded by the two-dimensional code, where m and n are positive integers.

10. The detection method according to claim 9, wherein, The m×n grid applied to the fitted quadrilateral is obtained by connecting the m equal division points of the first pair of opposite sides of the quadrilateral and the n equal division points of the second pair of opposite sides of the quadrilateral.

11. The detection method according to claim 1, further includes, when the image further includes the object with the two-dimensional code attached thereto, performing at least one of the following: Taking the actual position of the determined two-dimensional code relative to the imaging device as the actual position of the object relative to the imaging device; or The multi-class detection targets further include an object, and wherein, The detection method includes determining the actual position of the object relative to the imaging device based on the actual position of the determined two-dimensional code relative to the imaging device, and the position of the two-dimensional code in the image and the position of the object in the image obtained from the multi-target detection model.

12. The detection method according to claim 1 further includes at least one of the following: The multiple types of detection targets further include an object, and wherein, the detection method includes, when the image further includes the object not attached with the two-dimensional code, determining the actual position of the object not attached with the two-dimensional code relative to the imaging device based on the actual position of the two-dimensional code relative to the imaging device determined, and the position of the two-dimensional code in the image and the position of the object not attached with the two-dimensional code in the image obtained from the multi-object detection model; or the multiple types of detection targets further include an associated object of the object, the associated object not being attached with a two-dimensional code, and wherein the detection method includes, when the image further includes the associated object, determining the actual position of the associated object relative to the imaging device based on the actual position of the two-dimensional code relative to the imaging device determined, and the position of the two-dimensional code in the image and the position of the associated object in the image obtained from the multi-object detection model.

13. The detection method according to claim 1, wherein, determining the actual position of the two-dimensional code relative to the imaging device for capturing the image includes: using a reference image including a reference object captured by the imaging device, based on the image size and image position of the reference object in the reference image, and the angle subtended by the reference object relative to the imaging device determined by the actual size of the reference object and the actual position of the reference object relative to the imaging device, obtaining, by fitting, a proportionality coefficient between the angle subtended by the reference object relative to the imaging device and the image size of the reference object as a first function of the image position of the reference object, and based on the image position of the reference object and the actual position of the reference object relative to the imaging device, obtaining, by fitting, the elevation angle of the line connecting the reference object and the imaging device relative to the projection of the line on the plane where the imaging device is located as a second function of the image position of the reference object; based on the actual size of the two-dimensional code and the angle subtended by the two-dimensional code relative to the imaging device determined by the first function, the position of the two-dimensional code in the image, and the size of the two-dimensional code in the image, determining the actual distance of the two-dimensional code relative to the imaging device, based on the second function and the position of the two-dimensional code in the image determining the elevation angle of the line connecting the two-dimensional code and the imaging device relative to the projection of the line on the plane where the imaging device is located, and based on the position of the two-dimensional code in the image determining the azimuth angle of the projection of the line on the plane where the imaging device is located, thereby obtaining a spherical coordinate representation of the actual position of the two-dimensional code relative to the imaging device.

14. The detection method according to claim 13, wherein, The fitting of the first function and the second function is a binary polynomial fitting and is performed using the polar coordinate representation of the image position of the reference object, and wherein the obtained first function and second function are represented as functions of the rectangular coordinate representation of the image position of the reference object.

15. The detection method according to claim 1, wherein, the multi-class detection targets further include an object, and wherein the detection method includes, in the case where the image further includes the object attached with the two-dimensional code: determining the actual size of the object based on the actual size of the two-dimensional code, the size of the detection frame of the two-dimensional code in the image obtained from the multi-object detection model, and the size of the detection frame of the object in the image obtained from the multi-object detection model.

16. The detection method according to any one of claims 1 to 15, wherein, after successfully recognizing the two-dimensional code from the cropped image region, determining the actual position of the two-dimensional code relative to the imaging device.

17. A detection device, comprising: an image acquisition module configured to acquire an image, the image including a two-dimensional code for attaching to an object and recording information about the object; a target detection module configured to input the image into a pre-trained multi-object detection model to process the image, the multi-object detection model being configured to detect multi-class detection targets, the multi-class detection targets including two-dimensional codes; and at least one of a two-dimensional code recognition module and a position determination module, wherein the two-dimensional code recognition module is configured to crop, from the image, the image region of the determined detection frame of the two-dimensional code based on the detection frame of the two-dimensional code in the image obtained from the multi-object detection model, and recognize the two-dimensional code from the cropped image region to read the information about the object recorded by the two-dimensional code, wherein the position determination module is configured to determine the actual position of the two-dimensional code relative to the imaging device for capturing the image based on the actual size of the two-dimensional code, the size of the two-dimensional code in the image, and the position of the two-dimensional code in the image obtained from the multi-object detection model.

18. An object monitoring system, comprising: an imaging device configured to capture an image of an object; a processing device configured to execute the detection method according to any one of claims 1 to 16 to monitor the object based on the image captured by the imaging device.

19. A computing device, comprising: one or more processors; and a memory storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to execute the detection method according to any one of claims 1 to 16.

20. A non-transitory storage medium storing computer-executable instructions that, when executed by a computer, cause the computer to execute the detection method according to any one of claims 1 to 16.