A localization method and decoding method for time-domain visual landmarks based on RGB
Through RGB-based time-domain visual landmarks, using red, green and blue coded blocks and image processing algorithms, the problem of visual landmark recognition in low resolution cameras and using dark environments is solved, and precise positioning and dynamic message delivery of robots and other devices is realized.
Patent Information
- Application Number
- CN202310079504.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-02-02
AI Technical Summary
The existing visual landmark system has limited recognition accuracy under low-resolution cameras and cannot be used in dark environments, the information can not be changed, and the scope of application is limited.
The RGB-based time-domain visual landmarks are used, and the red, green and blue coded blocks are arranged in a regular triangle. The coded blocks can flash. The six-dimensional pose information is calculated through image processing and perspective projection algorithms, and dynamic messages are transmitted through light and dark combination encoding.
It improves the recognition accuracy of visual landmarks under low-resolution cameras, expands the scope of application, realizes effective use in dark environments, and supports precise positioning and dynamic message delivery of robots, robotic arms, and unmanned vehicles.
Smart Images

Figure CN116128976B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual landmark technology, and in particular to a positioning method and a decoding method of a time-domain visual landmark based on RGB. Background Art
[0002] Visual landmarks are artificial landmarks that can be used for automatic detection and positioning. At present, visual landmarks usually adopt static landmarks such as QR codes, ARTags, and AprilTags. In visual landmarks such as QR codes, ARTags, and AprilTags, a single coding block is only black and white, and self-luminescence is not considered, so the scope of application is limited; in addition, visual landmarks usually have at least 49 coding blocks, which requires a high resolution of the camera; and once the labels corresponding to these visual landmarks are printed and used, the information represented is fixed and cannot be changed.
[0003] The QR code is square, only in black and white, and consists of about 268 coding blocks. The positioning pattern of the Chinese character "回" is printed in three of the four corners in the positive direction. Through the positioning pattern, users can recognize it without alignment. Except for the three positioning patterns of the Chinese character "回", the rest of the places can encode data. The QR code was first used in automobile manufacturers to track parts, and was later widely used in electronic ticketing, B2B and other fields. The main function is to encode information, and decoding and extracting information by scanning the code. ARTag looks similar to a QR code, but the encoding system is very different from that of a QR code. It is mostly used in camera calibration, robot positioning, augmented reality and other occasions. It is mainly used to reflect the posture relationship between the camera and the tag. Inspired by ARTag and others, AprilTag reduces the number of coding blocks and consists of 49-100 coding blocks, which improves the performance in low-resolution and long-distance scenes.
[0004] These artificial landmark systems are usually displayed on a screen or printed out for use. The disadvantages of these systems are:
[0005] 1. Limited usage scenarios: Printed labels are prone to errors in dark environments and cannot be used in dark environments; labels displayed on the screen cannot be separated from the screen and have high hardware requirements;
[0006] 2. When using low-resolution cameras, the accuracy is limited: From QR to ARTag, and then to AprilTag, the number of tag encoding blocks has been reduced successively to improve the applicability to low-resolution cameras, but there is still room for improvement;
[0007] 3. Only static use is allowed: Once generated and printed, the information on the label is fixed.
[0008] Therefore, using existing visual landmarks for camera positioning requires a relatively high resolution of the camera. When the resolution of the camera is low, the recognition accuracy is limited, resulting in poor positioning accuracy. In addition, existing visual landmarks can only be used statically and cannot transmit information.
[0009] For this reason, the present invention designs a time-domain visual landmark based on RGB. The visual landmark includes three coding blocks, namely red, green, and blue. The three coding blocks are arranged in an equilateral triangle, and the three coding blocks can all flash according to a certain rule. The time-domain visual landmark based on RGB effectively improves the applicable range of artificial visual landmarks, reduces the requirement for camera resolution, and can achieve precise positioning and dynamic message transmission using the time-domain visual landmark based on RGB.
[0010] Since there are definite mutual distances between any two of the three coding blocks in the time-domain visual landmark based on RGB designed by the present invention, and the visual landmark composed of the three coding blocks has definite six-dimensional pose information (including three-dimensional position information (positions on the x, y, and z axes) and three-dimensional angle information (including roll angle around the x-axis, pitch angle around the y-axis, and yaw angle around the z-axis)), therefore, the present invention proposes a method for positioning the pose of the landmark and a method for decoding the information transmitted by the landmark for the time-domain visual landmark based on RGB, so as to provide technical support for the positioning and message transmission of robots, robotic arms, unmanned vehicles, etc., and better correct the errors of robots, robotic arms, unmanned vehicles, etc. Summary of the Invention
[0011] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a method for positioning and decoding a time-domain visual landmark based on RGB, which can provide technical support for the positioning and message transmission of robots, robotic arms, unmanned vehicles, etc.
[0012] To achieve the above object of the present invention, according to the first aspect of the present invention, there is provided a method for positioning a time-domain visual landmark based on RGB. The visual landmark includes three coding blocks, wherein the three coding blocks are a red coding block, a green coding block, and a blue coding block respectively. The three coding blocks are arranged in an equilateral triangle, and each coding block is a static constant-brightness coding block with unchangeable brightness, or each coding block is a dynamic coding block with two display states of bright and not bright;
[0013] The positioning method includes the following steps:
[0014] Collect an image of the visual landmark through a camera to obtain a visual landmark image;
[0015] Screen the qualified visual landmark images from the collected visual landmark images according to the preset image screening rules to obtain the qualified visual landmark images;
[0016] Preprocess the qualified visual landmark images to obtain the R-channel color extreme value image, the G-channel color extreme value image, and the B-channel color extreme value image;
[0017] Extract data from the R-channel color extreme value image, the G-channel color extreme value image, and the B-channel color extreme value image respectively to obtain the number of pixel points corresponding to the colors of the three RGB channels and the image plane centroid coordinates of the color coding blocks corresponding to the three RGB channels;
[0018] According to the image plane centroid coordinates of the color coding blocks corresponding to the three RGB channels, use the perspective projection algorithm to calculate all possible distance information of the red coding block, the green coding block, and the blue coding block relative to the camera;
[0019] Determine the true distance information of the three coding blocks relative to the camera based on the number of pixel points corresponding to the colors of the three RGB channels obtained through multiple image acquisitions and image processing and all possible distance information of the three coding blocks relative to the camera calculated;
[0020] Calculate the three-dimensional angle information of the red coding block, the green coding block, and the blue coding block relative to the camera according to the true distance information of the three coding blocks relative to the camera.
[0021] Preferably, each coding block of the visual landmark is a dynamic coding block with two display states of bright and not bright. Before using the visual landmark for positioning, combine the bright and not bright states of all coding blocks to obtain a plurality of light and dark combination codings, and perform periodic coding on the plurality of light and dark combination codings with a set coding period to obtain a coding sequence. Among them, a coding period starts with all three coding blocks being bright, at least one coding block is not bright within a coding period, and the combination of the coding blocks being bright is used to code information. When the three coding blocks are all bright again next time, the coding period ends. Correspondingly, the positioning method further includes:
[0022] Set the image acquisition frame rate of the camera, where the method for setting the image acquisition frame rate is as follows:
[0023] If the duration of each light and dark combination of the three coding blocks is d and the length of the coding sequence of a coding period is N, then the length T of each coding period is
[0024] T=(N + 1)d
[0025] Among them, each period includes a full-brightness coding indicating the start of the period. To detect each light and dark combination, the image acquisition frame rate f suitable for positioning of the camera should satisfy the following relationship:
[0026]
[0027] Preferably, the preprocessing of the qualified visual landmark images to obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image includes:
[0028] Adopt the sliding average method to perform smoothing filtering on the images of the RGB three channels of the qualified visual landmark images to reduce the high-frequency noise of the images, and obtain the visual landmark filtered image;
[0029] Copy the visual landmark filtered image to obtain three visual landmark images;
[0030] Use the color extreme value function to perform R-channel color extreme value processing, G-channel color extreme value processing, and B-channel color extreme value processing on the three visual landmark images respectively;
[0031] Perform erosion and dilation processing on the three visual landmark images after color extreme value processing respectively to remove residual noise;
[0032] Perform connector analysis and masking processing on the three visual landmark images after removing residual noise respectively, and obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image including the masked area.
[0033] Preferably, the data extraction from the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image respectively to obtain the number of pixel points of the corresponding colors of the RGB three channels and the image plane centroid coordinates of the corresponding color coding blocks of the RGB three channels includes:
[0034] Use the gradient algorithm to perform image processing on the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image respectively to obtain the contours of the coding blocks included in the masked areas of the images;
[0035] Calculate the pixels in the masked areas of the images to obtain the number of pixel points of the corresponding colors of the RGB three channels;
[0036] Perform linear combination operations on the pixel coordinates of the corresponding colors of the images to obtain the image plane centroid coordinates of the corresponding color coding blocks of the RGB three channels.
[0037] Preferably, determining the true distance information of the three coding blocks relative to the camera based on the number of pixel points of the corresponding colors in the three RGB channels obtained through multiple image acquisitions and image processing and all possible distance information of the three coding blocks relative to the camera obtained by calculation includes:
[0038] Method 1: Determine the distance of each coding block from the camera based on the number of pixel points of the corresponding colors in the three RGB channels obtained through multiple image acquisitions and image processing, so as to screen out the unique true distance information of the three coding blocks relative to the camera from all possible distance information of the three coding blocks relative to the camera obtained by calculation;
[0039] Method 2: When the unique true distance information cannot be screened out by using Method 1, determine the distance of each coding block from the camera according to the relationship between the angles of the triangles in the phase plane image and the angles of equilateral triangles, so as to screen out the unique true distance information of the three coding blocks relative to the camera from all possible distance information of the three coding blocks relative to the camera obtained by calculation;
[0040] Method 3: When the unique true distance information cannot be screened out by using Method 2, calculate all possible corresponding distances according to the historical trajectory of the camera and the image information collected during the movement offset, and exclude the possible positions that do not match the historical trajectory and movement offset of the camera, so as to obtain the unique true distance information of the three coding blocks relative to the camera.
[0041] According to the second aspect of the present invention, the present invention also provides a decoding method for time-domain visual landmarks based on RGB. The visual landmarks include three coding blocks, where the three coding blocks are a red coding block, a green coding block, and a blue coding block respectively. The three coding blocks are arranged in an equilateral triangle. Each coding block is a dynamic coding block with two display states: bright and not bright; before using the visual landmarks for information transmission, combine the bright and not bright display states of all coding blocks to obtain multiple light and dark combination encodings, and perform periodic encoding on the multiple light and dark combination encodings with a set coding period to obtain a coding sequence. One coding period starts with all three coding blocks being bright, and at least one coding block is not bright within one coding period. Encode information with the combination of the coding blocks being bright, and end the coding period when the three coding blocks are all bright again for the next time;
[0042] The decoding method includes the following steps:
[0043] Continuously collect images of the visual landmarks by a camera at a preset image acquisition frame rate to obtain a visual landmark image sequence, where the visual landmark image sequence includes multiple visual landmark images arranged in the order of acquisition time;
[0044] Preprocess each frame of the visual landmark images in the visual landmark image sequence in sequence to obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of the visual landmark image;
[0045] Judge and screen each frame of the visual landmark image according to the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of the visual landmark image to obtain an image sequence of one coding period;
[0046] Obtain the code corresponding to each frame of the image in the image sequence of one coding period by querying a preset coding table.
[0047] Preferably, the decoding method further includes:
[0048] Set the image acquisition frame rate of the camera, wherein the method for setting the image acquisition frame rate is as follows:
[0049] If the duration of each light and dark combination of the three coding blocks is d and the length of the coding sequence of one coding period is N, then the length T of each coding period is
[0050] T = (N + 1)d
[0051] Wherein, each period includes a full-bright code indicating the start of the period. In order to detect each light and dark combination, the image acquisition frame rate f of the camera applicable to decoding should satisfy the following relational expression:
[0052]
[0053] Wherein, m is an integer greater than 0. At this time, within the duration d of each light and dark combination of the visual landmark, the camera acquires m frames of images.
[0054] Preferably, the preprocessing method in the preprocessing of each frame of the visual landmark image in the visual landmark image sequence in sequence is as follows:
[0055] Adopt the sliding average method to perform smoothing filtering on the images on the RGB three channels of the visual landmark image to reduce the high-frequency noise of the image and obtain the visual landmark filtered image;
[0056] Copy the visual landmark filtered image to obtain three visual landmark images;
[0057] Use the color extreme value function to perform R-channel color extreme value processing, G-channel color extreme value processing, and B-channel color extreme value processing on the three visual landmark images respectively;
[0058] Perform erosion and dilation processing on the three visual landmark images after color extreme value processing respectively to remove residual noise;
[0059] Perform connection piece analysis and masking processing on the three visual landmark images after removing residual noise respectively, to obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image containing the masked area.
[0060] Preferably, judging and screening each frame of visual landmark image according to the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image, the obtained image sequence of one coding period includes:
[0061] Use the gradient algorithm to perform image processing on the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image respectively to obtain the contours of the coding blocks contained in the masked area of each image;
[0062] Count the pixels of the masked area within the contour of the coding block to obtain the number of pixel points of the corresponding color in each image;
[0063] Statistically count the number of pixel points of the corresponding color in the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image, and judge whether there are coding blocks of the corresponding color in each image according to the preset pixel point number threshold;
[0064] Taking a frame of visual landmark image that simultaneously contains the three coding blocks of the red coding block, green coding block, and blue coding block as the starting image, and a frame of visual landmark image that simultaneously contains the three coding blocks of the red coding block, green coding block, and blue coding block next time as the ending image, all the frame images between these two frames are arranged in sequence to form an image sequence of one coding period.
[0065] Preferably, the encoding corresponding to each frame of image in the image sequence of one coding period obtained by querying the preset encoding table includes:
[0066] Overlap the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image within one coding period, and use the gradient algorithm to process the overlapped image to obtain the total contour of the coding blocks contained in the masked area;
[0067] Count the number of pixel points of the corresponding color of the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image within one coding period at the corresponding positions within the total contour, and judge whether the corresponding coding blocks in each overlapped image are in the lit state or unlit state according to the preset pixel point number threshold;
[0068] According to the light and dark combinations of each coding block in each frame of the visual landmark image, query the preset coding table to obtain the coding of each frame of the visual landmark image.
[0069] As can be seen from the above solution, the present invention provides a method for positioning and decoding a time-domain visual landmark based on RGB, which can quickly and accurately calculate the six-dimensional pose information of the camera relative to the visual landmark, so as to realize the accurate positioning of the camera, and can quickly decode the dynamic messages transmitted by the visual landmark, so as to provide technical support for the positioning and message transmission of robots, robotic arms, unmanned vehicles, etc., so as to better correct the errors of robots, robotic arms, unmanned vehicles, etc.
[0070] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0072] Figure 1 is a diagram of 8 change states of the light and dark combinations of all coding blocks of a time-domain visual landmark based on RGB in a preferred embodiment of the present invention;
[0073] Figure 2 is a schematic diagram of the operating state of a camera in a preferred embodiment of the present invention;
[0074] Figure 3 is a flowchart of a method for positioning a time-domain visual landmark based on RGB in a preferred embodiment of the present invention;
[0075] Figure 4 is a flowchart of a method for decoding a time-domain visual landmark based on RGB in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0077] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined.
[0078] As Figure 1 shown, it is a diagram of 8 change states of the light and dark combinations of all coding blocks of a preferred embodiment of the present invention's RGB-based time-domain visual landmark. The RGB-based time-domain visual landmark provided in this embodiment includes:
[0079] A coding block group, wherein the coding block group includes three coding blocks, the colors of the three coding blocks are red, green, and blue respectively, the three coding blocks are arranged in an equilateral triangle, and when viewed from the top (looking down), the red, green, and blue coding blocks are placed in a clockwise order (or counterclockwise), and each coding block has two states: on and off, and each coding block can switch between the on and off states according to a preset rule.
[0080] Specifically, although the three coding blocks only need to not be on a straight line to obtain the pose of the relative visual landmark during subsequent decoding, the equilateral triangle shape has no inclination on any side and has the best balance for images at any angle. Therefore, in this embodiment, the three coding blocks in the same coding block group are arranged in an equilateral triangle.
[0081] Specifically, from the implementation of the coding blocks, there are various colors of LED lamp beads, but red, green, and blue are the three most convenient colors to obtain and have the best economy. From the subsequent decoding, since color is usually encoded by three RGB channels, the three extreme values on the three channels are red, green, and blue respectively, with the best discrimination. Therefore, the three coding blocks in this visual landmark are respectively a red coding block, a green coding block, and a blue coding block.
[0082] It should be noted that in actual applications, a visual landmark can, as described in the above embodiment, only include one coding block group, and each coding block group includes three coding blocks of red, green, and blue. A visual landmark can also include a multi-coding block group obtained by combining multiple coding block groups, and each coding block group includes three coding blocks of red, green, and blue. From this, it can be understood that the scheme of combining multiple such visual landmarks for use should also be regarded as a specific implementation manner of the visual landmark of the present invention.
[0083] The embodiment of the present invention also provides a coding method for an RGB-based time-domain visual landmark, and the method includes the following steps:
[0084] S1. Represent 0 as bright and 1 as not bright. Combine the bright and not bright display states of all the coding blocks in all the coding block groups of the RGB-based time-domain visual landmark in claim 5 to obtain a coding block brightness and darkness combination table containing multiple brightness and darkness combination codings.
[0085] S2. Perform a number system conversion on all the brightness and darkness combination codings in the coding block brightness and darkness combination table to obtain a coding table.
[0086] S3. Perform a number system conversion on the information to be transmitted to obtain the information to be coded.
[0087] S4. Based on the coding table, code the information to be coded into a periodic coding sequence according to the time series, where each coding in the periodic coding sequence corresponds to a coding block brightness and darkness combination.
[0088] Specifically, in this embodiment, only one coding block group is adopted for the visual landmark. This coding block group includes three coding blocks arranged in an equilateral triangle. The three coding blocks are a red coding block, a green coding block, and a blue coding block respectively. The brightness and darkness combination of the three coding blocks in this embodiment can obtain a total of 8 brightness and darkness combination codings. The number system conversion of all the brightness and darkness combination codings in the coding block brightness and darkness combination table is performed to obtain the coding table as shown in the following table:
[0089] Light and dark combined coding Corresponding to septenary 000 Start / End 001 0 010 1 011 2 100 3 101 4 110 5 111 6
[0090] For example, if it is necessary to represent the number 0 under ASCII code, the corresponding binary code is 00110000, which is converted to decimal 48 and then to septenary 66. Querying the above coding table, it can be known that the corresponding coding sequence is the brightness and darkness combination 000111111000.
[0091] Specifically, in step S4, a cycle starts with all three coding blocks being bright; within the cycle, at least one coding block is not bright, and the combination of the lit coding blocks is used to code the information; when the three coding blocks are all bright again next time, the cycle ends.
[0092] In this embodiment, the cycle of dynamic change is determined by the all-bright coding. For subsequent detection and decoding, only when this visual landmark is detected can further cycle sampling and decoding be performed. In a dark environment, when all the coding blocks are not bright, it cannot be detected. When one or two coding blocks are bright, the artificial feature is not obvious and it is easy to be confused with other devices or light sources. Therefore, in this embodiment, all three coding blocks being bright is adopted to determine the dynamic change cycle of the coding blocks, and one information transmission is completed within one dynamic change cycle. If multiple information transmissions are required, between two adjacent coding cycles, all the coding blocks of the visual landmark being bright is adopted as the demarcation point between two adjacent information transmissions.
[0093] In this embodiment, within a cycle, the combination of light and dark of three coding blocks is used to encode information. Although each coding block can also blink independently and use a coding similar to Morse code to transmit information independently on three channels, since the full-brightness coding has been used to determine the dynamic change cycle, the blinking of the three channels is not completely independent, which will increase the complexity of information encoding, detection, and decoding. Therefore, the combination of light and dark of the coding blocks is used to encode information without independent control.
[0094] The time domain in the time-domain visual landmark of the present invention refers to that the coding blocks will change in the combination of light and dark over time according to the needs of transmitting information. The user can freely choose the number of combinations transmitted in each cycle, that is, choose the cycle length, or use multiple cycles to transmit information.
[0095] As Figure 2 shown, it is a schematic diagram of the working state of a camera in a preferred embodiment of the present invention. Since there is a definite mutual distance between any two of the three coding blocks in the RGB-based time-domain visual landmark designed by the present invention, and the visual landmark composed of the three coding blocks has definite six-dimensional pose information (including three-dimensional position information (positions on the x, y, and z axes) and three-dimensional angle information (including roll angle around the x-axis, pitch angle around the y-axis, and yaw angle around the z-axis)), therefore, for this RGB-based time-domain visual landmark, the present invention proposes a method for positioning the landmark pose and a method for decoding the information transmitted by the landmark.
[0096] As Figure 3 shown, an embodiment of the present invention provides a method for positioning an RGB-based time-domain visual landmark. The visual landmark includes three coding blocks, where the three coding blocks are a red coding block, a green coding block, and a blue coding block respectively. The three coding blocks are arranged in an equilateral triangle. Each coding block is a static constant-brightness coding block with an unchangeable brightness, or each coding block is a dynamic coding block with two display states of bright and not bright;
[0097] The positioning method includes the following steps:
[0098] S101, collecting an image of the visual landmark through a camera to obtain a visual landmark image;
[0099] Specifically, when the visual landmark is a static landmark, that is, each coding block is a static constant-brightness coding block with an unchangeable brightness, the image of the visual landmark is directly collected through the camera at a set image acquisition frame rate;
[0100] When the visual landmark is a dynamic landmark, that is, each coding block of the visual landmark is a dynamic coding block with two display states of bright and not bright, before positioning using the visual landmark, the bright and not bright display states of all coding blocks are combined to obtain multiple light and dark combination codings, and the multiple light and dark combination codings are cyclically coded with a set coding period to obtain a coding sequence. Among them, a coding period starts with all three coding blocks being bright, at least one coding block is not bright within one coding period, and the combination of coding blocks being bright is used to code information. When the next time all three coding blocks are bright, the coding period ends. Correspondingly, the positioning method further includes:
[0101] Set the image acquisition frame rate of the camera. Among them, the method for setting the image acquisition frame rate is as follows:
[0102] If the duration of each light and dark combination of the three coding blocks is d, and the length of the coding sequence of one coding period is N, then the length T of each coding period is
[0103] T=(N + 1)d
[0104] Among them, each period includes a full-bright coding indicating the start of the period. In order to detect each light and dark combination, the applicable image acquisition frame rate f of the camera for positioning should satisfy the following relational expression:
[0105]
[0106] By setting the image acquisition frame rate of the camera through the above method, it can be ensured that the camera can collect each light and dark combination of each coding block of the visual landmark.
[0107] S102, screen qualified visual landmark images from the collected visual landmark images according to a preset image screening rule to obtain qualified visual landmark images;
[0108] Specifically, the qualified visual landmark images include images of three coding blocks, and the screening method is as follows:
[0109] If the visual landmark image collected by the camera has no image of other light sources with similar colors interfering, and includes images of three coding blocks of the visual landmark, it is a qualified visual landmark image; if there is a light source with a similar color interfering in the image, the interference can be excluded manually or using machine vision methods, and the image of the visual landmark can be selected to form a qualified visual landmark image.
[0110] S103, preprocess the qualified visual landmark images to obtain an R-channel color extreme value image, a G-channel color extreme value image, and a B-channel color extreme value image;
[0111] Specifically, the preprocessing method is as follows:
[0112] The sliding average method is adopted to perform smoothing filtering on the images of the three RGB channels of the qualified visual landmark images to reduce the high-frequency noise of the images, and a visual landmark filtered image is obtained;
[0113] The visual landmark filtered image is copied to obtain three visual landmark images;
[0114] The color extreme value function is used to perform R-channel color extreme value processing, G-channel color extreme value processing, and B-channel color extreme value processing on the three visual landmark images respectively; specifically, when performing R-channel color extreme value processing on one of the visual landmark images, the pixels with the largest proportion of the R channel relative to the other channels in this visual landmark image are retained, and the remaining pixels are modified to black. The processing methods for performing G-channel color extreme value processing and B-channel color extreme value processing on the other two visual landmark images are similar to the foregoing method and will not be elaborated here;
[0115] The corrosion and dilation processing are respectively performed on the three visual landmark images after the color extreme value processing to remove the residual noise;
[0116] The connection piece analysis and masking processing are respectively performed on the three visual landmark images after removing the residual noise to obtain an R-channel color extreme value image, a G-channel color extreme value image, and a B-channel color extreme value image including the masking area.
[0117] S104, data extraction is respectively performed on the R-channel color extreme value image, the G-channel color extreme value image, and the B-channel color extreme value image to obtain the number of pixel points corresponding to the colors of the three RGB channels and the image plane centroid coordinates of the color coding blocks corresponding to the three RGB channels;
[0118] Specifically, the process is as follows:
[0119] The gradient algorithm is used to perform image processing on the R-channel color extreme value image, the G-channel color extreme value image, and the B-channel color extreme value image respectively to obtain the contours of the coding blocks included in the masking area of each image;
[0120] The pixels in the masking area of each image are calculated to obtain the number of pixel points corresponding to the colors of the three RGB channels;
[0121] The pixel coordinates corresponding to the colors of each image are subjected to a linear combination operation to obtain the image plane centroid coordinates of the color coding blocks corresponding to the three RGB channels.
[0122] S105, according to the image plane centroid coordinates of the color coding blocks corresponding to the three RGB channels, the perspective projection algorithm is used to solve all possible distance information of the three coding blocks, namely the red coding block, the green coding block, and the blue coding block, relative to the camera;
[0123] Specifically, the process is as follows:
[0124] Based on the six - dimensional pose information of the visual landmarks, the three - dimensional centroid coordinates of the three coding blocks are obtained using geometric relationships;
[0125] According to the image - plane centroid coordinates of the three color - coding blocks and the three - dimensional centroid coordinates of the coding blocks, using the method of polynomial root - finding, all four possible distances between the camera and the three coding blocks are solved.
[0126] S106, Determine the true distance information of the three coding blocks relative to the camera based on the number of pixel points of the corresponding colors in the RGB three channels obtained through multiple image acquisitions and image processing and all possible distance information of the three coding blocks relative to the camera obtained by calculation;
[0127] The process of determining the true distance information of the three coding blocks relative to the camera is as follows:
[0128] Method 1: Determine the distance of each coding block from the camera based on the number of pixel points of the corresponding colors in the RGB three channels obtained through multiple image acquisitions and image processing, and thus screen out the unique true distance information of the three coding blocks relative to the camera from all possible distance information of the three coding blocks relative to the camera obtained by calculation;
[0129] Method 2: When the unique true distance information cannot be screened out by Method 1, determine the distance of each coding block from the camera according to the relationship between the angles in the triangle in the phase - plane image and the angles of an equilateral triangle, and thus screen out the unique true distance information of the three coding blocks relative to the camera from all possible distance information of the three coding blocks relative to the camera obtained by calculation;
[0130] Method 3: When the unique true distance information cannot be screened out by Method 2, calculate all possible distances corresponding to the historical trajectory of the camera and the image information collected during the movement offset, and exclude the possible positions that do not match the historical trajectory and movement offset of the camera to obtain the unique true distance information of the three coding blocks relative to the camera.
[0131] S107, Calculate the three - dimensional angle information of the red coding block, green coding block, and blue coding block relative to the camera according to the true distance information of the three coding blocks relative to the camera. The specific calculation process is as follows:
[0132] According to the true distance information of the camera from the three coding blocks and the three - dimensional centroid coordinates of the three coding blocks, establish and further solve a system of equations using geometric relationships to obtain the three - dimensional angle information of the three coding blocks relative to the camera, thereby realizing the positioning of the camera.
[0133] In this embodiment, according to the image of the two-dimensional image plane collected by the camera, the six-dimensional pose information of the camera relative to the time-domain landmark is inversely calculated, so as to realize the positioning of the camera.
[0134] As Figure 4 shown, an embodiment of the present invention provides a decoding method for a time-domain visual landmark based on RGB.
[0135] The visual landmark includes three coding blocks. Among them, the three coding blocks are a red coding block, a green coding block, and a blue coding block respectively. The three coding blocks are arranged in an equilateral triangle. Each coding block is a dynamic coding block with two display states: bright and not bright. Before using the visual landmark for information transmission, the bright and not bright display states of all coding blocks are combined to obtain a plurality of light and dark combination codings, and the plurality of light and dark combination codings are periodically coded with a set coding period to obtain a coding sequence. Among them, a coding period starts with all three coding blocks being bright, and at least one coding block is not bright within one coding period. The combination of the coding blocks being bright is used to code information, and when the three coding blocks are all bright again next time, this coding period ends.
[0136] The decoding method includes the following steps:
[0137] S201, continuously collect images of the visual landmark by the camera at a preset image acquisition frame rate to obtain a visual landmark image sequence, where the visual landmark image sequence includes multiple visual landmark images arranged in the order of acquisition time.
[0138] Specifically, the method for setting the image acquisition frame rate is as follows:
[0139] If the duration of each light and dark combination of the three coding blocks is d, and the length of the coding sequence of one coding period is N, then the length T of each coding period is
[0140] T=(N + 1)d
[0141] Among them, each period includes a full-bright coding that marks the start of the period. In order to detect each light and dark combination, the applicable image acquisition frame rate f of the camera for decoding should satisfy the following relational expression:
[0142]
[0143] where m is an integer greater than 0. At this time, within the duration d of each light and dark combination of the visual landmark, the camera collects m frames of images. Specifically, in this embodiment, m = 1. It should be noted that when m>1, the final decoding result needs to be downsampled m times.
[0144] S202. Preprocess each frame of the visual landmark images in the visual landmark image sequence in order to obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark images;
[0145] Specifically, the preprocessing method is as follows:
[0146] Use the sliding average method to perform smoothing filtering on the images of the three RGB channels of the visual landmark images to reduce the high-frequency noise of the images, and obtain the visual landmark filtered images;
[0147] Copy the visual landmark filtered images to obtain three copies of visual landmark images;
[0148] Use the color extreme value function to perform R-channel color extreme value processing, G-channel color extreme value processing, and B-channel color extreme value processing on the three copies of visual landmark images respectively; specifically, when performing R-channel color extreme value processing on one of the visual landmark images, retain the pixels in the R-channel of this visual landmark image with the largest ratio relative to the other channels, and modify the remaining pixels to black. The processing methods for performing G-channel color extreme value processing and B-channel color extreme value processing on the other two visual landmark images are similar to the above method and will not be elaborated here;
[0149] Perform erosion and dilation processing on the three copies of visual landmark images after color extreme value processing respectively to remove residual noise;
[0150] Perform connector analysis and masking processing on the three copies of visual landmark images after removing residual noise respectively to obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image containing the masked area.
[0151] S203. Judge and screen each frame of visual landmark images according to the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark images to obtain an image sequence for one coding period;
[0152] Specifically, the judgment and screening process is as follows:
[0153] Use the gradient algorithm to perform image processing on the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark images respectively to obtain the contours of the coding blocks contained in the masked areas of the images;
[0154] Count the pixels in the masked areas within the contours of the coding blocks to obtain the number of pixel points of the corresponding colors in each image;
[0155] Count the number of color pixels corresponding to the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of the visual landmark image, and determine whether there is a coding block of the corresponding color in each image according to a preset pixel number threshold; if the counted number of corresponding color pixels is greater than the preset pixel number threshold, it means that there is a coding block of the corresponding color in the image, otherwise it is judged that there is no coding block of the corresponding color in the image. The pixel number threshold can be set according to experience and experiments;
[0156] Taking a frame of visual landmark image that simultaneously contains three coding blocks: a red coding block, a green coding block, and a blue coding block as the starting image, and the next frame of visual landmark image that simultaneously contains these three coding blocks as the ending image, all the frame images between these two frames are arranged in order to form an image sequence of a coding cycle.
[0157] S204. Obtain the code corresponding to each frame image in the image sequence of a coding cycle by querying a preset coding table. The specific implementation process is as follows:
[0158] Overlap the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of the visual landmark image within a coding cycle, and use the gradient algorithm to process the overlapped image to obtain the total contour of the coding blocks contained in the mask area;
[0159] Count the number of pixels of the corresponding color pixels of the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of the visual landmark image within a coding cycle within the total contour at the corresponding position, and determine whether the corresponding coding block in each overlapped image is in the lit state or the unlit state according to a preset pixel number threshold; specifically, when the counted number of corresponding color pixels within the total contour at the corresponding position exceeds the preset pixel number threshold, it is determined that the corresponding coding block is in the lit state, otherwise it is determined that it is in the unlit state. The pixel number threshold can be set according to experience and experiments;
[0160] According to the light and dark combination of each coding block in each frame of the visual landmark image, query the preset coding table to obtain the code of each frame of the visual landmark image, which is the complete decoding process. By combining the codes of each frame of the visual landmark image obtained by decoding in order, the information content transmitted by the visual landmark within the coding cycle can be obtained through parsing.
[0161] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0162] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0163] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0164] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A positioning method for time-domain visual landmarks based on RGB, characterized in that The visual landmark includes three coding blocks, where the three coding blocks are a red coding block, a green coding block, and a blue coding block respectively. The three coding blocks are arranged in an equilateral triangle. Each coding block is a static constant-brightness coding block with unchangeable brightness, or each coding block is a dynamic coding block with two display states of bright and not bright; The positioning method includes the following steps: Collect an image of the visual landmark through a camera to obtain a visual landmark image; Select qualified visual landmark images from the collected visual landmark images according to a preset image screening rule to obtain qualified visual landmark images; Preprocess the qualified visual landmark image to obtain an R-channel color extreme value image, a G-channel color extreme value image, and a B-channel color extreme value image; Extract data from the R-channel color extreme value image, the G-channel color extreme value image, and the B-channel color extreme value image respectively to obtain the number of pixel points of the corresponding colors in the three RGB channels and the image plane centroid coordinates of the corresponding color coding blocks in the three RGB channels; Solve all possible distance information of the three coding blocks, namely the red coding block, the green coding block, and the blue coding block, relative to the camera according to the image plane centroid coordinates of the corresponding color coding blocks in the three RGB channels by using a perspective projection algorithm; Determine the true distance information of the three coding blocks relative to the camera based on the number of pixel points of the corresponding colors in the three RGB channels obtained through multiple image acquisitions and image processings and all possible distance information of the three coding blocks relative to the camera obtained by solving; Solve the three-dimensional angle information of the three coding blocks, namely the red coding block, the green coding block, and the blue coding block, relative to the camera according to the true distance information of the three coding blocks relative to the camera; where, When each coding block of the visual landmark is a dynamic coding block with two display states of bright and not bright, before using the visual landmark for positioning, combine the bright and not bright display states of all coding blocks to obtain multiple light and dark combination codings, and perform periodic coding on the multiple light and dark combination codings at a set coding period to obtain a coding sequence. Wherein, a coding period starts with all three coding blocks being bright, at least one coding block is not bright within one coding period, and the combination of the coding blocks being bright is used to code information. When the three coding blocks are all bright again next time, end the coding period. Correspondingly, the positioning method further includes: Set the image acquisition frame rate of the camera, where the method for setting the image acquisition frame rate is as follows: If the duration of each light and dark combination of the three coding blocks is d and the length of the coding sequence of one coding period is N, then the length T of each coding period is T=(N + 1)d where each period includes a full-bright coding indicating the start of the period. In order to detect each light and dark combination, the applicable image acquisition frame rate f of the camera for positioning should satisfy the following relational expression:
2. The positioning method of the RGB-based time-domain visual landmark according to claim 1, wherein The preprocessing of the qualified visual landmark image to obtain an R-channel color extreme value image, a G-channel color extreme value image, and a B-channel color extreme value image includes: The sliding average method is used to perform smoothing filtering on the images of the RGB three channels of the qualified visual landmark images to reduce the high-frequency noise of the images, and a visual landmark filtered image is obtained; The visual landmark filtered image is copied to obtain three visual landmark images; The color extreme value functions are used to perform R-channel color extreme value processing, G-channel color extreme value processing, and B-channel color extreme value processing on the three visual landmark images respectively; The corrosion and dilation processing are respectively performed on the three visual landmark images after the color extreme value processing to remove the residual noise; The connection piece analysis and masking processing are respectively performed on the three visual landmark images after removing the residual noise, and an R-channel color extreme value image, a G-channel color extreme value image, and a B-channel color extreme value image including the masked area are obtained.
3. The method for positioning a time-domain visual landmark based on RGB according to claim 2, characterized in that, The data extraction is respectively performed on the R-channel color extreme value image, the G-channel color extreme value image, and the B-channel color extreme value image to obtain the number of pixel points of the corresponding colors of the RGB three channels and the image plane centroid coordinates of the corresponding color coding blocks of the RGB three channels, including: The gradient algorithm is used to perform image processing on the R-channel color extreme value image, the G-channel color extreme value image, and the B-channel color extreme value image respectively to obtain the contours of the coding blocks included in the masked areas of the respective images; The pixels in the masked areas of the respective images are calculated to obtain the number of pixel points of the corresponding colors of the RGB three channels; The linear combination operation is performed on the pixel coordinates of the corresponding colors of the respective images to obtain the image plane centroid coordinates of the corresponding color coding blocks of the RGB three channels.
4. The RGB-based time-domain visual landmark positioning method according to any one of claims 1-3, characterized in that The true distance information of the three coding blocks relative to the camera is determined by the number of pixel points of the corresponding colors of the RGB three channels obtained through multiple image acquisitions and image processing and all possible distance information of the three coding blocks relative to the camera obtained by calculation, including: Method 1: The number of pixel points of the corresponding colors of the RGB three channels obtained through multiple image acquisitions and image processing is used to determine the distance of each coding block from the camera, so as to screen out the unique true distance information of the three coding blocks relative to the camera from all possible distance information of the three coding blocks relative to the camera obtained by calculation; Method 2: When the unique true distance information cannot be screened out by using Method 1, according to the relationship between the triangle angle and the equilateral triangle angle in the phase plane image, the distance of each coding block from the camera is determined, so as to screen out the unique true distance information of the three coding blocks relative to the camera from all possible distance information of the three coding blocks relative to the camera obtained by calculation; Method 3: When the unique true distance information cannot be screened out by using Method 2, according to the historical trajectory of the camera and the image information collected during the remote motion offset, all possible distances are calculated, and the possible positions that do not match the historical trajectory and remote motion offset of the camera are excluded, and the unique true distance information of the three coding blocks relative to the camera is obtained.
5. A decoding method for time-domain visual landmarks based on RGB, characterized in that The visual landmark includes three coding blocks, where the three coding blocks are a red coding block, a green coding block, and a blue coding block respectively, and the three coding blocks are arranged in an equilateral triangle. Each coding block is a dynamic coding block with two display states: bright and not bright. Before using the visual landmark for information transmission, the bright and not bright display states of all coding blocks are combined to obtain a plurality of light and dark combination codings, and the plurality of light and dark combination codings are periodically coded with a set coding period to obtain a coding sequence. Wherein, a coding period starts with all three coding blocks being bright, at least one coding block is not bright within one coding period, and the combination of the coding blocks being bright is used to code information. When all three coding blocks are bright again next time, the coding period ends; The decoding method includes the following steps: Continuously collect images of the visual landmark by a camera at a preset image acquisition frame rate to obtain a visual landmark image sequence, where the visual landmark image sequence includes multiple frames of visual landmark images arranged in the order of acquisition time; Preprocess each frame of the visual landmark image sequence in sequence to obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of the visual landmark image; Judge and screen each frame of the visual landmark image according to the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of the visual landmark image to obtain an image sequence of one coding period; Obtain the coding corresponding to each frame of the image in the image sequence of one coding period by querying a preset coding table; The decoding method further includes: Set the image acquisition frame rate of the camera, where the method for setting the image acquisition frame rate is as follows: If the duration of each light and dark combination of the three coding blocks is d and the length of the coding sequence of one coding period is N, then the length T of each coding period is T=(N + 1)d Wherein, each period includes the all-bright coding indicating the start of the period. In order to detect each light and dark combination, the applicable image acquisition frame rate f of the camera for decoding should satisfy the following relational expression: Where m is an integer greater than 0. At this time, within the duration d of each light and dark combination of the visual landmark, the camera acquires m frames of images.
6. The decoding method of RGB-based time-domain visual landmarks according to claim 5, characterized in that The preprocessing method in the step of preprocessing each frame of the visual landmark image sequence in sequence is as follows: Use the sliding average method to perform smoothing filtering on the images on the RGB three channels of the visual landmark image to reduce the high-frequency noise of the image and obtain a visual landmark filtered image; Copy the visual landmark filtered image to obtain three visual landmark images; Perform R-channel color extreme value processing, G-channel color extreme value processing, and B-channel color extreme value processing on the three visual landmark images respectively by using a color extreme value function; Perform erosion and dilation processing on the three visual landmark images after color extreme value processing respectively to remove residual noise; Perform connector analysis and masking on the three visual landmark images after removing residual noise respectively to obtain the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image containing the masked area.
7. The decoding method of RGB-based time-domain visual landmarks according to claim 6, characterized in that Judging and screening each frame of visual landmark image according to the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image, the obtained image sequence of one coding period includes: Use the gradient algorithm to perform image processing on the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image respectively to obtain the contours of the coding blocks contained in the masked area of each image; Count the pixels in the masked area within the contour of the coding block to obtain the number of pixel points of the corresponding color in each image; Statistically count the number of pixel points of the corresponding color in the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image, and judge whether there is a coding block of the corresponding color in each image according to the preset pixel point number threshold; Taking a frame of visual landmark image that simultaneously contains the three coding blocks of the red coding block, green coding block, and blue coding block as the starting image, and the next frame of visual landmark image that simultaneously contains the three coding blocks of the red coding block, green coding block, and blue coding block as the ending image, all the frame images between these two frames are arranged in order to form an image sequence of one coding period.
8. The decoding method of the RGB-based time-domain visual landmark according to claim 7, wherein The encoding corresponding to each frame of image in the image sequence of one coding period obtained by querying the preset encoding table includes: Overlap the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image within one coding period, and use the gradient algorithm to process the overlapped image to obtain the total contour of the coding blocks contained in the masked area; Count the number of pixel points of the corresponding color pixel points of the R-channel color extreme value image, G-channel color extreme value image, and B-channel color extreme value image of each frame of visual landmark image within one coding period within the pixels of the total contour at the corresponding positions, and judge whether the corresponding coding blocks in each overlapped image are in the lit state or unlit state according to the preset pixel point number threshold; According to the light and dark combination of each coding block in each frame of visual landmark image, query the preset encoding table to obtain the encoding of each frame of visual landmark image.
Citation Information
Patent Citations
Artificial vision landmark and coding method thereof
CN112270715A
Decoding and positioning method for artificial vision landmark
CN112270716A