Internet of Things-based Book Inventory Robot System and Method

By designing a book inventory robot system based on the Internet of Things, using a variety of sensors and image acquisition technologies, the problem of low efficiency of book shelf problems and the high cost, high time-consuming and inaccurate positioning of existing RFID systems is solved, and the fast, accurate positioning and efficient management of book locations are achieved.

CN114092802BActive Publication Date: 2025-06-27ZHONGKAI UNIV OF AGRI & ENG +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111281622.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-06-27
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

The prior art has low efficiency in handling the problem of book shelves in libraries, and the RFID system is expensive, time-consuming and inaccurate positioning.

Method used

A book inventory robot system based on the Internet of Things is designed, using robot base mechanism, main controller, infrared sensor, magnetic navigation sensor, ultrasonic radar, gyroscope, bookshelf positioning module and image acquisition and processing unit to realize automatic positioning and data acquisition of book locations through image acquisition and recognition technology.

Benefits of technology

It realizes the rapid and accurate positioning of book locations, reduces the workload of manual inventory, improves the efficiency of book management, and reduces the manufacturing cost of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092802B_ABST
    Figure CN114092802B_ABST
Patent Text Reader

Abstract

An Internet of Things-based book inventory robot system and method, which is provided with a robot base mechanism. A main controller is arranged in the robot base mechanism. The input ends of the main controller are respectively connected with an infrared sensor, a magnetic navigation sensor, an ultrasonic radar, a gyroscope, a bookshelf positioning module and an image acquisition and processing unit; the output end is connected with a motor drive module, and the communication end is connected with a communication module. The method includes: one, starting; two, cruising; three, detecting the first electronic tag of the bookshelf; four, adjusting the lifting structure and collecting the book spine image data; five, detecting the second electronic tag of the bookshelf, turning around, and raising the camera mechanism by one row height to collect the book spine image data; six, detecting the first electronic tag of the bookshelf, and judging whether the data collection of the bookshelf is completed according to the number of times of detecting the first electronic tag and the second electronic tag of the bookshelf. If so, enter the next step, otherwise enter step five; seven, cruising, continue cruising to complete the collection of all the book spine image data of the bookshelves.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the Internet of Things, and specifically to a book inventory robot system and method based on the Internet of Things. Background Art

[0002] At present, the book management work in libraries is mainly based on the library information management system. However, for the mis-shelving of books, even if the library has a very perfect information management system and sufficient staff, it is still helpless in the face of the mis-shelving problem. In real life, there is a situation where the books shown in the library management system are not found at their placed positions, which increases the difficulty of book inventory.

[0003] In order to accurately locate the positions and relevant information of books and facilitate the inventory and error correction in libraries, a small number of libraries introduce RFID book management systems to manage books. Although RFID can reduce the workload of staff inventory, there are still some deficiencies. First, because each book needs to be equipped with an RFID tag, the overall cost of RFID tags is too high; second, since a large amount of time is required for attaching tags and inputting information, a large amount of manual workload is consumed; finally, the tags are prone to interference with each other, and there are still problems of inaccurate positioning and low recognition rate. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention proposes a book inventory robot system and method based on the Internet of Things with high intelligence, simple control, and low manufacturing cost. The specific technical solutions are as follows:

[0005] A book inventory robot system based on the Internet of Things is provided with a robot base mechanism. A main controller is arranged in the robot base mechanism. The input end of the main controller is respectively connected with an infrared sensor, a magnetic navigation sensor, an ultrasonic radar, a gyroscope, a bookshelf positioning module, and an image acquisition and processing unit; the output end is connected with a motor drive module, and the communication end is connected with a communication module.

[0006] The method of the book inventory robot system based on the Internet of Things specifically comprises the following steps:

[0007] Step 1: Start the robot system;

[0008] Step 2: Start the magnetic navigation sensor, the motor drive module is started to drive the wheels to rotate, and cruise according to the set route;

[0009] Step 3: Start the bookshelf positioning module. The bookshelf positioning module detects the first electronic tag at the starting point of the bookshelf passage, records the number of read electronic tags, and stops the vehicle;

[0010] Step 4: Start the image acquisition and processing unit. Through the lifting mechanism, align the camera mechanism in the image acquisition and processing unit with the bottom spine of the bookshelf. Start the motor drive module to drive the wheels to rotate and collect the spine image data.

[0011] Step 5: When the bookshelf positioning module detects the second electronic tag at the end of the bookshelf channel and records the number of read electronic tags,

[0012] Compare the number of times of detecting the first electronic tag and the second electronic tag of the bookshelf with the information of the number of bookshelf layers to determine whether the data acquisition of the current bookshelf is completed. If yes, go to Step 8; otherwise, go to Step 6.

[0013] Step 6: Start the motor drive module to drive the wheels to rotate. The robot base mechanism rotates 180 degrees to complete a U-turn. The camera mechanism in the image acquisition and processing unit rises by one row height, aligns with the spine, starts the motor drive module to drive the wheels to rotate, and collect the spine image data.

[0014] Step 7: The bookshelf positioning module detects the first electronic tag of the bookshelf again and stops. Determine whether the data acquisition of the current bookshelf is completed according to the number of times of detecting the first electronic tag and the second electronic tag of the bookshelf. If yes, go to Step 8; otherwise, go to Step 6.

[0015] Step 8: Continue to cruise according to the set route and enter the next bookshelf.

[0016] Step 9: Complete the cruise of the set route to realize the acquisition of the spine image data of all bookshelves.

[0017] As an optimization: A set of motor drive modules are respectively arranged on the two opposite sides of the robot base mechanism. The output end of the motor drive module is connected with wheels. A lifting mechanism is arranged on the top of the robot base mechanism. An image acquisition and processing unit is arranged on the top of the lifting mechanism. Ultrasonic radars are respectively arranged on the front and rear parts of the robot base mechanism. At least one infrared sensor is also arranged on the top of the robot base mechanism. A bookshelf positioning module is arranged at the bottom of the robot base mechanism. A main controller and a gyroscope are arranged inside the robot base mechanism.

[0018] As an optimization: In Step 6, when the robot base mechanism rotates 180 degrees to complete a U-turn, specifically, start the ultrasonic radar and the infrared sensor to avoid obstacles, and adjust the steering acceleration according to the inclination data collected by the gyroscope to complete the steering and U-turn.

[0019] As an optimization: The specific process of image acquisition and processing in Step 4 is as follows:

[0020] 4.1: Image preprocessing;

[0021] 4.2: Image skew correction;

[0022] 4.3: Complete the positioning and recognition of the book spine from the image;

[0023] 4.4: Detect the text area;

[0024] 4.5: Text recognition;

[0025] 4.6: Post-processing of text, optimizing the recognition results.

[0026] As an optimization: The specific steps of the above-mentioned step 4.1 image preprocessing are as follows:

[0027] 4.1.1. Grayscale processing: The conversion from a color image to a grayscale image can be completed by a linear transformation, satisfying the following formula P(x, y) = kR(x, y) + lG(x, y) + mB(x, y), where k + l + m = 1, R(x, y), G(x, y), B(x, y) are the R, G, B values of the pixel point (x, y) respectively, k, l, m are pre-determined parameters, and the grayscale value of P(x, y) can be obtained;

[0028] 4.1.2. Binarization processing: Since the images captured by the lifting camera are color images, and the color images contain a huge amount of information, which will affect the recognition effect of the text. Therefore, in order to recognize the text more clearly, binarization processing is required. A set threshold is selected, and the pixels with grayscale values greater than the set threshold are defined as white pixels, while the pixels less than the set threshold are defined as black pixels, and finally a binary image can be converted;

[0029] 4.1.3. Noise removal.

[0030] As an optimization: The above-mentioned 4.1.3. Noise removal is specifically as follows. The first step is basic estimation:

[0031] For each target patch, find at most one similar patch nearby. To avoid the influence of noise, the patches are transformed by 2D and then the Euclidean distance is used to measure the similarity degree. After sorting by distance from small to large, take at most the first one and stack them into a three-dimensional array;

[0032] For the third dimension of the 3D array, that is, after the patches are stacked, the array composed of the pixel points at the same position of each patch, after DCT transformation, the components less than the hyperparameter formula are set to 0 by the hard threshold method, and at the same time, the number of non-zero components is counted as a reference for the subsequent weights, and then the third dimension is inversely transformed;

[0033] After inverse-transforming these patches and putting them back in place, use the non-zero component number statistics to superimpose the weights. Finally, divide the stacked image by the weight of each point to obtain the image of the basic estimate. At this time, most of the noise in the image has been removed;

[0034] Step 2, Final Estimation:

[0035] Since the basic estimation greatly eliminates noise, for each target patch of the noisy original image, the similarity degree can be directly measured by the Euclidean distance of the corresponding basic estimation patch. After sorting in ascending order of distance, take at most the first hyperparameter number, and stack the basic estimation patches and the noisy original image patches into two three-dimensional arrays respectively;

[0036] Perform DCT transformation on the third dimension of the 3D array containing the basic estimation, that is, after the patches are stacked, the array composed of the pixel points at the same position of each patch;

[0037] Multiply the coefficients by the noisy 3D patches and put them back in place, and finally perform weighted average adjustment to obtain the final estimation image, which restores more details of the original image compared to the basic estimation image.

[0038] As an optimization: The 4.2, Image Skew Correction is specifically as follows: When the robot scans the text on the book spine, if the text image is skewed, this skew of the image will affect the final character recognition accuracy. To avoid the problem of low recognition accuracy, the method of using Hough transform for image skew correction is adopted. At the same time, to overcome the disadvantage of large computational amount of Hough transform, the variable-resolution image pyramid strategy is adopted. At the same time, the skew angle of the scanned image is quickly and accurately measured, and it has high anti-noise performance and application adaptability. The basic strategy of Hough transform is to calculate the possible trajectory of the reference point in the parameter space from the coordinates of the target object pixels in the image space, and count the calculated reference points in an accumulator.

[0039] As an optimization: The positioning and recognition of the book spine from the image is specifically as follows: First, the book spine in the bookshelf image needs to be cut out. Since the contour of the book spine image can be approximated as a rectangle, usually the straight line extraction algorithm is used to identify the two long edge straight lines of the book spine, and the position of the book spine is determined by the two parallel long edge straight lines. The LSD straight line detection algorithm can be used. The LSD algorithm mainly includes three parts: region growing, rectangular region approximation, and region validity detection. The algorithm mainly finds the pixel point regions in the image that share the same gradient direction. In the LSD line segment detection algorithm, this kind of image region is called a line support region, also called a level-line. The direction of the detected line segment is approximately the average direction in the line support region.

[0040] As an optimization: For step 4.4, detecting the text area, specifically, the DB (Differentiable Binarization) algorithm is used for text detection. The scene text detection based on segmentation is to convert the probability map generated by the segmentation method into a bounding box and a text area, which also includes the post-processing process of binarization. First, the image passes through the resnet50-vd layer of the feature pyramid structure, and the output of the feature pyramid is transformed into the same size by upsampling, and features and feature layers are generated by cascading; then, the probability map and the text probability map are predicted through the feature layer, which is used to calculate the probability that the pixel belongs to the text to form the text probability map. Then, the dynamic threshold map is formed according to the dynamic threshold of each pixel. Next, the DB binary map is generated through the text probability map and the dynamic threshold map. Finally, the label is generated by expanding the DB binary map to form a text box. In the training stage, supervision is applied to the threshold map, the probability map, and the approximate binary map, and the latter two share the same supervision; in the inference stage, the bounding box can be easily obtained from the latter two.

[0041] 4.5: Text recognition, specifically: After the text area in the picture is located through text detection, the text in the area needs to be recognized. The CNN+RNN+Attention algorithm is adopted. This method is a text recognition algorithm based on visual attention. First, the model runs a sliding CNN on the input picture to extract features. Then, the obtained feature sequence is input into the LSTM stacked on top of the CNN for encoding the feature sequence. Then, the attention model is used for decoding and the label sequence is output. The attention model adopted by the CNN+RNN+Attention algorithm allows the decoder to calculate the variable context vector by weighted averaging of the hidden state of the encoder at each step of the decoding process. Therefore, the most relevant information can be read at all times without completely relying on the hidden state of the previous moment.

[0042] 4.6: Post-processing of text, optimizing the recognition result, specifically: Post-processing is used to optimize the recognition result, which can be corrected through a language model. At the same time, the recognized images of OCR often contain a large amount of text, and there are complex situations in the typesetting and font sizes of these texts. In the post-processing, the recognition result is formatted to unify the typesetting rules and improve the recognition efficiency.

[0043] The beneficial effects of the present invention are as follows: 1. By using the cameras on both sides of the robot to simultaneously recognize the book spines of the bookshelves on both sides to obtain books, time is saved and the recognition efficiency is improved;

[0044] 2. The electronic tags at the starting point and the ending point of the passage between the library bookshelves are used to control the 180° turning of the robot and to detect the electronic tags on the ground bookshelves.

[0045] 3. The image acquisition lifting device controls the lifting of the camera platform to the horizontal position of the bookshelf to obtain book information at different levels, which is convenient to control and has high efficiency.

[0046] 4. By determining the number of times the robot detects the electronic tag on the ground and comparing it with the number of layers of the bookshelf, it is judged that the robot has completed the acquisition of all book information on the current bookshelf, so that the robot can enter the next aisle. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic structural diagram of the present invention.

[0048] Figure 2 It is a schematic diagram of the bookshelf and the robot walking track in the present invention.

[0049] Figure 3 It is a schematic diagram of the distribution of the bookshelf and the robot in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0050] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0051] Such as Figure 1Shown: An Internet of Things-based book inventory robot system, which includes a robot base, robot drive wheels, and drive wheel DC motors. The robot drive wheels are movably connected to the robot base. The drive wheel DC motors are arranged on the robot base and connected to the robot drive wheels. An image acquisition device, an image acquisition lifting device, a lifting device motor, a bookshelf positioning reader, and a server are provided. The image acquisition lifting device is located in the middle of the robot base and connected to the lifting device motor, and its top is connected to the image acquisition device. The bookshelf positioning reader is located at the bottom of the robot base. The server is signal-connected to the image acquisition device and the bookshelf positioning reader and stores the information read by the image acquisition device and the bookshelf positioning reader. The bookshelf positioning reader detects the bookshelf number electronic tags located on the ground and transmits the detected information to the central processing unit. The image acquisition device has a pair of cameras located on both sides of the top of the image acquisition lifting device and is used to simultaneously acquire the information of the book spines on both sides. On the basis of the above technology, preferably, it further includes a drive module, a communication module, a debugging interface, and a power module. The drive module realizes the signal conversion with the DC motor and the encoder and is used to control the movement of the robot in the library and the rotation of the wheels. The power module provides power for the inventory robot. Further preferably, it further includes a data acquisition module, a magnetic navigation sensor, an ultrasonic sensor, and an infrared sensor. The magnetic navigation sensor has a pair of magnetic navigation sensors located at the front and rear positions at the bottom of the machine and is used to jointly acquire the ground magnetic stripe information and return it to the control device through the data link. The ultrasonic sensor has a pair located at the front and rear of the robot and is used to detect whether there are obstacles in the front and rear directions of the robot. The infrared sensor is located at the four corners of the robot and is used to monitor whether there are obstacles in the vertical direction of the trolley, such as books protruding from the bookshelf. The data acquisition module realizes the information interaction between the magnetic navigation sensor, the ultrasonic sensor, and the infrared sensor and the main controller.

[0052] The layout of the library is as Figure 2 shown, and it consists of bookshelves, a fork magnetic navigation guide rail 51, a main fork navigation line 52, and a branch fork navigation line 53. Among them, the bookshelves are arranged in sequence on both sides of the branch fork navigation line 53 and are used for the robot to obtain book information. The main fork navigation line 52 is arranged side by side at both ends of the bookshelves. The main fork navigation line 52 and the branch fork navigation line 53 are connected by the fork magnetic navigation guide rail 51 and are used for the robot to switch between different bookshelves.

[0053] The bookshelf is a general layered bookshelf and can be divided into left and right sides. There are several electronic tags pasted on the branch fork navigation line 53 between the bookshelves, which are located at both ends of the navigation line. The electronic tag information mainly includes the bookshelf number and the area number where the bookshelf is located. In this way, the robot can know the relevant information of the bookshelf by reading the information of the electronic tags.

[0054] As Figure 3As shown: When the robot detects the bookshelf tag number 31, the lifting device motor at the bottom of the lifting structure in the robot controls the image acquisition lifting device to lift the image acquisition device to the horizontal position of the book. The robot continues to travel along the line, and the image acquisition device continuously records the information of the book spine. When the bookshelf positioning reader in the robot detects the bookshelf electronic tag again, the robot makes a 180° turn and records the number of times the electronic tag is detected. The lifting device motor controls the image acquisition lifting device to lift the image acquisition device to the horizontal position of the next layer of books again. The robot continues to travel along the line, and the image acquisition device continuously records the information of the book spine. This process continues until all the books on the bookshelves on both sides of the channel are recorded. In this way, by using the image acquisition device to collect the information of the book spine, the image acquisition lifting device, and combining with the OCR recognition technology, the acquisition of book information can be realized.

[0055] The robot is enabled to switch between the main branch navigation line and the branch navigation line. When traveling on the main branch navigation line and detecting the electronic tag on the ground, the robot enters the branch navigation line. When switching from the main branch navigation line to the branch navigation line, the number of layers of the bookshelf is pre-recorded on the main control chip. By comparing it with the number of times the detected electronic tag is recorded, if the two are equal, it indicates that all the book information in the current channel has been recorded, and the robot can enter the next channel.

[0056] To enable the robot to travel along the main branch navigation line, a magnetic navigation sensor and a magnetic strip are required. The magnetic strip is set at both ends of the library bookshelves to form a main navigation route. By setting the magnetic field intensity signal range of the magnetic strip, the robot can be controlled to travel along the main branch navigation line of the library bookshelves. The magnetic navigation sensor is set at the bottom of the robot to detect the magnetic field intensity signal on the ground in real time and transmit it to the main controller. The main controller has pre-set the magnetic field intensity signal range of the magnetic strip. As shown: To enable the robot to travel along the branch navigation line, in addition to the magnetic navigation sensor and the magnetic strip, an infrared sensor and an ultrasonic radar are also required. Four infrared sensors are respectively set at the four corners of the top of the robot chassis and connected to the main controller. They are used to effectively avoid the books protruding from the bookshelves. The ultrasonic radar is fixedly set on the front and rear sides of the chassis and is used to detect the distance signal between the chassis and the bookshelves on both sides and transmit it to the main control. To realize the turning around and steering of the robot, in addition to the infrared sensor and the ultrasonic radar, a gyroscope is also required. The gyroscope is connected to the main controller and is used to detect the angular acceleration of the chassis and transmit it to the main controller. The main controller has pre-set the angular acceleration of the steering. When the robot needs to turn and turn around, the main controller controls the motor to turn to the corresponding angle. If the robot deviates from the navigation line during the route or turning around, the ultrasonic radar and the gyroscope cooperate together to enable the robot to maintain the correct posture and continue to travel. The power system of the robot mainly consists of a main controller, a motor, and small wheels. The robot uses a lifting camera combined with Paddle OCR to collect and process the spine information of the books on each layer of the bookshelf to obtain the book information of a certain bookshelf. Here, Paddle OCR is mainly used for text recognition, and the recognized spine information is returned to the background server. Specifically, Paddle OCR performs text recognition and processing on the pictures collected by the lifting camera, and returns the extracted spine information to the background server. The main steps are as follows:

[0057] Preprocessing:

[0058] Grayscale processing: The conversion from a color image to a grayscale image can be completed by a linear transformation, which satisfies the following formula P(x, y) = kR(x, y) + lG(x, y) + mB(x, y). Where k + l + m = 1, R(x, y), G(x, y), B(x, y) are the R, G, B values of the pixel point (x, y) respectively, and k, l, m are pre-determined parameters. By calculating the grayscale value of P(x, y), the conversion can be achieved.

[0059] Binarization processing: Since the images captured by the lift camera are color images, and color images contain a huge amount of information, which will affect the recognition effect of text. Therefore, in order to recognize text more clearly, binarization processing is required. In this process, a specific threshold can be selected. Pixels with a gray value greater than this threshold are defined as white pixels, and pixels less than this threshold are defined as black pixels. Finally, a binary image can be converted.

[0060] Noise removal: Since digital images in reality are often affected by noise interference from imaging devices and external environments during the digitization and transmission processes, the quality of noise reduction processing will directly affect the accuracy of image recognition. The BM3D algorithm can be used to solve this problem. The main process is as follows:

[0061] The first step, basic estimation:

[0062] For each target patch, find up to MAXN1 (hyperparameter) similar patches nearby. To avoid the influence of noise, the patches are transformed by 2D transformation (DCT transformation is used in the code) and then the Euclidean distance is used to measure the similarity. After sorting by distance from small to large, take the first MAXN1 at most. Stack them into a three-dimensional array.

[0063] For the third dimension of the 3D array, that is, after the patches are stacked, the array composed of pixel points at the same position of each patch. After performing DCT transformation, the components less than the hyperparameter [formula] are set to 0 in a hard threshold manner. At the same time, the number of non-zero components is counted as a reference for subsequent weights. Then the third dimension is inversely transformed.

[0064] After inversely transforming these patches and putting them back in place, use the non-zero component number statistics to superimpose weights. Finally, divide the stacked image by the weight of each point to obtain the image of the basic estimate. At this time, most of the noise in the image has been removed.

[0065] The second step, final estimation:

[0066] Since the basic estimate greatly eliminates noise, for each target patch of the noisy original image, the Euclidean distance of the corresponding basic estimate patch can be directly used to measure the similarity. After sorting by distance from small to large, take the first MAXN1 at most. Stack the basic estimate patches and the noisy original image patches into two three-dimensional arrays respectively.

[0067] For the third dimension of the 3D array containing the basic estimate, that is, after the patches are stacked, the array composed of pixel points at the same position of each patch, perform DCT transformation.

[0068] Multiply the coefficients by the noisy 3D patches and put them back in place. Finally, perform weighted average adjustment to obtain the final estimate image. Compared with the basic estimate image, more details of the original image are restored.

[0069] Skew Correction: When the robot scans the text on the book spine, the uploaded image will more or less be skewed to some extent. This skew of the image will affect the final character recognition accuracy. To avoid the problem of low recognition accuracy, the image can be corrected by software methods. The method of Hough transform can be used for image skew correction. At the same time, to overcome the disadvantage of large computational amount of Hough transform, this method adopts the variable-resolution image pyramid strategy. At the same time, this method can quickly and accurately measure the skew angle of the scanned image, and has high anti-noise performance and application adaptability. The basic strategy of Hough transform is to calculate the possible trajectories of reference points in the parameter space from the coordinates of the target object pixels in the image space, and count the calculated reference points in an accumulator.

[0070] Book Spine Location and Recognition:

[0071] After preprocessing, book spine location and recognition are essential processes. First, the book spine in the bookshelf image needs to be cut out. Since the contour of the book spine image can be approximated as a rectangle, the straight line extraction algorithm is usually used to identify the two long edge straight lines of the book spine, and the position of the book spine is determined by the two parallel long edge straight lines. The LSD straight line detection algorithm can be used. The LSD algorithm mainly includes three parts: region growing, rectangular region approximation, and region validity detection. The algorithm mainly finds the pixel point regions in the image that share the same gradient direction. In the LSD line segment detection algorithm, this kind of image region is called the line support region, also known as the level-line. The direction of the detected line segment is roughly the average direction in the line support region. This algorithm has the advantages of self-adaptation and fast detection speed. It is very suitable for book spine recognition that requires real-time performance and self-adaptability.

[0072] Text Detection:

[0073] This patent uses the DB (Differentiable Binarization) algorithm for text detection. DB is a segmentation-based text detection algorithm, whose full name is differentiable binarization processing. Segmentation-based scene text detection means converting the probability map (heat map) generated by the segmentation method into bounding boxes and text regions, which also includes the post-processing process of the binarization operation. The algorithm process is shown in the following figure. First, the image passes through the resnet50-vd layer of the feature pyramid structure, and the output of the feature pyramid is transformed into the same size through upsampling, and features and feature layers are generated by cascading; then, the probability map (probability map) and text probability map are predicted through the feature layer, which is used to calculate the probability that the pixel belongs to the text to form the text probability map. Then, the dynamic threshold map is formed according to the dynamic threshold of each pixel. Then, the DB binary map is generated through the text probability map and the dynamic threshold map. Finally, the label is generated by expanding the DB binary map to form a text box. In the training stage, supervision is applied to the threshold map, probability map, and approximate binary map, where the latter two share the same supervision; in the inference stage, the bounding box can be easily obtained from the latter two.

[0074] Text recognition:

[0075] After localizing the text area in the image through text detection, it is also necessary to recognize the text in the area. This patent mainly uses the CNN+RNN+Attention algorithm, which is a text recognition algorithm based on visual attention. The main process is shown in the following figure. First, the model runs a sliding CNN on the input image to extract features, and then the obtained feature sequence is input into the LSTM stacked on top of the CNN for encoding the feature sequence. Then, the attention model is used for decoding and the label sequence is output. The attention model adopted by the CNN+RNN+Attention algorithm allows the decoder to calculate the variable context vector by weighted averaging the hidden state of the encoder at each step of the decoding process. Therefore, it can always read the most relevant information without completely relying on the hidden state of the previous moment.

[0076] Post-processing:

[0077] Post-processing is used to optimize the recognition results and can be corrected through a language model. At the same time, the recognized images of OCR often contain a large amount of text, and these texts have complex situations in typesetting and font sizes. In post-processing, the recognition results are formatted, the typesetting rules are unified, and the recognition efficiency is improved.

Claims

1. A book inventory robot system based on the Internet of Things, characterized in that: A robot base mechanism is provided, and a main controller is provided in the robot base mechanism. The input ends of the main controller are respectively connected to an infrared sensor, a magnetic navigation sensor, an ultrasonic radar, a gyroscope, a bookshelf positioning module, and an image acquisition and processing unit; the output end is connected to a motor drive module, and the communication end is connected to a communication module; The method of the book inventory robot system based on the Internet of Things is specifically as follows: Step 1: Start the robot system; Step 2: Start the magnetic navigation sensor, the motor drive module starts, drives the wheels to rotate, and cruises according to the set route; Step 3: Start the bookshelf positioning module. The bookshelf positioning module detects the first electronic tag at the starting point of the bookshelf passage, records the number of read electronic tags, and stops the vehicle; Step 4: Start the image acquisition and processing unit. Through the lifting mechanism, the camera mechanism in the image acquisition and processing unit is aligned with the bottom spine of the bookshelf. The motor drive module starts, drives the wheels to rotate, and acquires the spine image data; Step 5: When the bookshelf positioning module detects the second electronic tag at the end of the bookshelf passage and records the number of read electronic tags, According to the number of times of detecting the first electronic tag and the second electronic tag of the bookshelf, compare with the bookshelf layer information to judge whether the data acquisition of the current bookshelf is completed. If yes, enter Step 8; otherwise, enter Step 6; Step 6: The motor drive module starts, drives the wheels to rotate, the robot base mechanism rotates 180 degrees to complete a U-turn, the camera mechanism in the image acquisition and processing unit rises to the height of one row, aligns with the spine, the motor drive module starts, drives the wheels to rotate, and acquires the spine image data; Step 7: The bookshelf positioning module detects the first electronic tag of the bookshelf again, stops the vehicle, and judges whether the data acquisition of the current bookshelf is completed according to the number of times of detecting the first electronic tag and the second electronic tag of the bookshelf. If yes, enter Step 8; otherwise, enter Step 6; Step 8: Continue to cruise according to the set route and enter the next bookshelf; Step 9: Complete the cruise of the set route and realize the acquisition of the spine image data of all bookshelves.

2. The book inventory robot system based on the Internet of Things according to claim 1, wherein: A set of motor drive modules are respectively provided on the opposite sides of the robot base mechanism. The output end of the motor drive module is connected to a wheel. A lifting mechanism is provided on the top of the robot base mechanism. An image acquisition and processing unit is provided on the top of the lifting mechanism. Ultrasonic radars are respectively provided at the front and rear of the robot base mechanism. At least one infrared sensor is also provided on the top of the robot base mechanism. A bookshelf positioning module is provided at the bottom of the robot base mechanism. A main controller and a gyroscope are provided inside the robot base mechanism.

3. The method of the book inventory robot system based on the Internet of Things according to claim 1, characterized in that: In Step 6, the robot base mechanism rotates 180 degrees to complete a U-turn. Specifically, start the ultrasonic radar and the infrared sensor to avoid obstacles, and adjust the steering acceleration according to the inclination data collected by the gyroscope to complete the steering and U-turn.

4. The method of the book inventory robot system based on the Internet of Things according to claim 1, wherein: the image acquisition and processing process in Step 4 is specifically as follows: 4.1: Image preprocessing; 4.2: Image skew correction; 4.3: Complete the positioning and recognition of the book spine from the image; 4.4: Detect the text area; 4.5: Text recognition; 4.6: Post-processing of the text to optimize the recognition result.

5. The method of the book inventory robot system based on the Internet of Things according to claim 4, characterized in that: The specific image preprocessing in step 4.1 is as follows: 4.1.

1. Grayscale processing: The conversion from a color image to a grayscale image can be completed by a linear transformation that satisfies the following formula P(x, y) = kR(x, y) + lG(x, y) + mB(x, y), where k + l + m = 1, R(x, y), G(x, y), and B(x, y) are the R, G, and B values of the pixel point (x, y) respectively, and k, l, and m are pre-determined parameters. By calculating P(x, y), the grayscale value can be obtained; 4.1.

2. Binarization processing: Since the images captured by the lift camera are color images, which contain a large amount of information and will affect the text recognition effect. Therefore, for clearer text recognition, binarization processing is required. A set threshold is selected, and the pixels with grayscale values greater than the set threshold are defined as white pixels, while the pixels less than the set threshold are defined as black pixels. Finally, a binary image can be obtained; 4.1.

3. Noise removal.

6. The method of the book inventory robot system based on the Internet of Things according to claim 5, characterized in that: The specific noise removal in 4.1.3 is as follows. The first step is the basic estimation: For each target patch, find at most one similar patch in the vicinity. To avoid the influence of noise, the patches are transformed by 2D and then the Euclidean distance is used to measure the similarity. After sorting by distance from small to large, take at most the first one and stack them into a three-dimensional array; For the third dimension of the 3D array, that is, after the patches are stacked, the array composed of the pixel points at the same position of each patch is subjected to DCT transformation. Then, the components less than the hyperparameter formula are set to 0 by the hard threshold method, and the number of non-zero components is counted as a reference for the subsequent weight. Then, the inverse transformation is performed on the third dimension; After the inverse transformation of these patches and putting them back in place, the weights are superimposed using the non-zero component count. Finally, the stacked image is divided by the weight of each point to obtain the basic estimated image, and at this time, most of the noise in the image is removed; The second step is the final estimation: Since the basic estimation greatly eliminates noise, for each target patch in the original noisy image, the similarity can be directly measured by the Euclidean distance of the corresponding basic estimated patch. After sorting by distance from small to large, take at most the first one, and stack the basic estimated patches and the original noisy patches into two three-dimensional arrays respectively; For the third dimension of the 3D array containing the basic estimation, that is, after the patches are stacked, the array composed of the pixel points at the same position of each patch is subjected to DCT transformation; Multiply the coefficients by the noisy 3D patches and put them back in place. Finally, weighted average adjustment is performed to obtain the final estimated image, which restores more details of the original image compared to the basic estimated image.

7. The method of the book inventory robot system based on the Internet of Things according to claim 4, characterized in that: 4.

2. Image skew correction is specifically as follows: When the robot scans the text on the book spine, if the text image is skewed, this skew will affect the final character recognition accuracy. To avoid the problem of low recognition accuracy, the method of using Hough transform for image skew correction is adopted. At the same time, to overcome the disadvantage of large computational amount of Hough transform, the variable-resolution image pyramid strategy is adopted. Meanwhile, the skew angle of the scanned image is quickly and accurately measured, and it has high anti-noise performance and application adaptability. The basic strategy of Hough transform is to calculate the possible trajectory of the reference point in the parameter space from the coordinates of the target object pixels in the image space, and count the calculated reference points in an accumulator.

8. The method of the book inventory robot system based on the Internet of Things according to claim 4, characterized in that: The positioning and recognition of the book spine from the image are specifically as follows: First, the book spine in the bookshelf image needs to be cut out. Since the contour of the book spine image can be approximated as a rectangle, usually the straight line extraction algorithm is used to identify the two long edge straight lines of the book spine, and the position of the book spine is determined by the two parallel long edge straight lines. The LSD straight line detection algorithm can be used. The LSD algorithm mainly includes three parts: region growing, rectangular region approximation, and region validity detection. The algorithm mainly finds the pixel point regions in the image that share the same gradient direction. In the LSD line segment detection algorithm, this kind of image region is called the line support region, also known as the level-line. The direction of the detected line segment is approximately the average direction in the line support region.

9. The method of the book inventory robot system based on the Internet of Things according to claim 4, characterized in that: 4.

4. Detecting the text region is specifically as follows: The DB (Differentiable Binarization) algorithm is used for text detection. Scene text detection based on segmentation is to convert the probability map generated by the segmentation method into bounding boxes and text regions, which also includes the post-processing process of binarization operation. First, the picture passes through the resnet50-vd layer of the feature pyramid structure, and the output of the feature pyramid is transformed into the same size by upsampling, and features and feature layers are cascaded. Then, the probability map (probability map) and text probability map are predicted through the feature layer to calculate the probability that the pixel belongs to the text to form the text probability map. Then, the dynamic threshold map is formed according to the dynamic threshold of each pixel. Then, the DB binary map is generated through the text probability map and the dynamic threshold map. Finally, the text box is formed by expanding the label according to the DB binary map. In the training stage, supervision is applied to the threshold map, probability map, and approximate binary map, and the latter two share the same supervision; in the inference stage, the bounding box can be easily obtained from the latter two. 4.5: Text recognition, specifically: After text detection to locate the text area in the picture, it is also necessary to recognize the text in the area. The CNN+RNN+Attention algorithm is adopted. This method is a text recognition algorithm based on visual attention. First, the model runs a sliding CNN on the input picture to extract features, then inputs the obtained feature sequence into an LSTM stacked on top of the CNN for encoding the feature sequence, and then uses an attention model for decoding and outputs a label sequence. The attention model adopted by the CNN+RNN+Attention algorithm allows the decoder to calculate a variable context vector by weighted averaging of the hidden states of the encoder at each step of the decoding process. Therefore, it can read the most relevant information at all times without completely relying on the hidden state of the previous moment; 4.6: Post-processing of text, optimizing the recognition result, specifically: Post-processing is used to optimize the recognition result and can be corrected through a language model. At the same time, the recognized images of OCR often contain a large amount of text, and there are complex situations in the typesetting and font sizes of these texts. In post-processing, the recognition result is formatted, the typesetting rules are unified, and the recognition efficiency is improved.

Citation Information

Patent Citations

  • Library operation robot and operation method thereof

    CN111687853A