A method for intelligent inventory and shelf arrangement assistance in libraries based on machine vision
By deploying cameras in the library and using machine vision technology for book image processing and recognition, the efficiency and accuracy of book search and management in the prior art are solved, and efficient and accurate book management and unsensed use are achieved.
Patent Information
- Application Number
- CN202410355690.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-03-27
AI Technical Summary
The prior art has problems with efficiency and accuracy in the search and management of books in libraries, especially when the placement of books on bookshelf causes the camera to capture mainly spine images, and it becomes difficult to identify books through detection and recognition of spines.
Using intelligent inventory and shelf assistance methods based on machine vision, the library bookshelf is transformed by deploying cameras, the book image information is obtained, the book image is restored and enhanced, and the Hough line transformation and Mask-RCNN model are used for image segmentation and recognition, and the book title and book number are obtained to achieve accurate positioning and efficient management of books.
It significantly improves the efficiency and accuracy of book search and management in the library, solves the difficult problems in the process of borrowing and returning books, and realizes true unsensible use, is not affected by external environmental factors, and is significantly cost-effective.
Smart Images

Figure CN118277641B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of library book search, and in particular to a machine vision-based library intelligent inventory and shelf arrangement auxiliary method. Background Art
[0002] Since the placement of books on the bookshelf causes the camera to mainly capture the image of the book spine, it becomes feasible to identify the book by detecting and identifying the book spine. A series of vision-based book spine detection and recognition methods have been developed in this field. At present, the use of line detection algorithms for book spine instance segmentation is the main method in this research field. A book spine visual recognition algorithm that combines Hough transform and state machine theory is proposed. The core of this algorithm is to combine the state machine theory on the basis of Hough transform to realize state machine-like transformation to identify the book spine, which effectively solves the impact of the book spine gap problem on recognition.
[0003] A visual book spine recognition algorithm based on an improved Hough transform was also proposed. This method uses parallel lines to exclude falsely detected straight lines, thereby significantly improving the accuracy of spine edge detection. A visual book spine recognition system based on wavelet analysis and probabilistic Hough transform has emerged. The Hough transform algorithm in this system is different because it can adjust different Hough parameters according to the thickness of the spine, thereby improving the robustness of the algorithm. These studies show that the efficiency and accuracy of book spine recognition in libraries can be significantly improved through refined image processing and computer vision technology.
[0004] Existing vision-based book spine recognition methods mostly use text detection and recognition methods to obtain the text on the book spine, and then use text retrieval methods to achieve book spine recognition. The invention adopts a different idea from text detection and recognition, and directly uses the features of the book spine image to achieve book recognition. Summary of the invention
[0005] The present invention provides a method for intelligent inventory counting and shelving assistance in libraries based on machine vision, which has solved the current problem of difficulty in borrowing and returning books. By using computer vision to locate books, the retrieval efficiency and accuracy are greatly improved.
[0006] To achieve the above object, the technical solution adopted by the present invention is:
[0007] A method for intelligent inventory counting and shelf arrangement assistance in a library based on machine vision comprises the following steps:
[0008] S1. Transform the library shelves by deploying cameras;
[0009] S2. Restore the book image information obtained by S1 based on the image edge points;
[0010] S3. Enhance the restored book image based on SRGAN;
[0011] S4. Preprocessing of book image rotation correction based on Hough line transform;
[0012] S5. Use the Mask-RCNN model to perform image segmentation on the preprocessed data to obtain the spine image of a single book;
[0013] S6. Obtain the book title and book number in the image through OCR, and integrate the library lending system with the book title and book number as key indexes.
[0014] Furthermore, the step S1 includes the following steps:
[0015] S11. Refer to the bookshelf of Nanjing University of Aeronautics and Astronautics Library. The bookshelf width is 1000mm, the single-sided bookshelf depth is 300mm, the height is 2000mm, and one layer can store about 50 books;
[0016] S12. Install two Xiaomi cameras with 3 million pixels on a six-story bookshelf, with one camera installed on the top of the bookshelf and the other installed in the middle between the third and fourth layers of the bookshelf. This layout ensures that the two cameras can cover and capture the image information of all books on the adjacent bookshelves.
[0017] Furthermore, the step S2 comprises the following steps:
[0018] S12. Considering that the distance between adjacent bookshelves is relatively close, the captured images are not only affected by the light, but also by the tilt of the book spines due to the overhead shooting, so the images need to be preprocessed;
[0019] S21. Standardize the detection image, use the trained model to detect the standardized distorted image, and obtain the rectangular detection frame of the target, multiple pairs of Bezier control points on the edge of the target, and the Mask of the target. The target detection frame Rect is defined by the coordinates of the upper left corner point and the width and height of the circumscribed rectangle: Rect = x, y, w, h, x is the horizontal coordinate of the upper left corner point of the target detection frame Rect, y is the vertical coordinate of the upper left corner point of the target detection frame Rect, w is the width of the circumscribed rectangle, and h is the height of the circumscribed rectangle;
[0020] The result of control point regression in the network model is the normalized relative distance from the control point to the upper left corner of the document detection rectangle Rect, x j is the horizontal coordinate of the jth control point, y j is the ordinate of the jth control point, Δx and Δy are the results of control point regression in the MaskRCNN network model, and d x dy is the normalized relative distance, w x 、w y is the standardized weight value; calculated as:
[0021] Δx=x j -x
[0022] Δy=y j -y
[0023]
[0024]
[0025] The coordinate x of the jth control point is obtained j ,y j The relative distance d obtained by regression x ,d y And the document rectangle detection box obtained by regression is calculated:
[0026] S22. The coordinates of the points on the target edge curve are calculated through the control points of the Bezier curve. These initial points are used to calculate the rectangular template after image correction. The corresponding coordinate points on the rectangular template are the target points. These initial points and target points correspond to each other up and down and left and right. When constructing the rectangular template, this study selected the same number of points as the corresponding edge points in the original image as the left and right edge points of the template. The coordinates of these points in the corrected image are defined as target points. Assuming that the number of edge point pairs is k, the total number of edge points is 2k. w is the width of the circumscribed rectangle, and h is the height of the circumscribed rectangle. At the left boundary of the template, k points are evenly selected from top to bottom. The edge points of the left boundary of the rectangular template are represented as: 0, Similarly, k edge points are also selected at the corresponding position of the right boundary. The points on the right boundary of the rectangular template are as shown in formula 0.
[0027] S23. Use the TPS transformation algorithm for the initial point and target point obtained in S12, calculate the transformation matrix of the transformation algorithm, and define two corresponding point sets, namely point set S and point set T: S = m 1 ,n 1 ,...,m n ,n n , T = p 1 ,q 1 ,...,p n ,q n , the S point set is called the template point set, the T point set is the target point set, and m 1 ,n 1 ,...,m n ,n nis the coordinate of the template point set, p 1 ,q 1 ,...,p n ,q n are the coordinates of the target point set, which are all control points of the TPS algorithm;
[0028] S24. Use the transformation matrix obtained in S23 to transform the original image in S11 to obtain a corrected target image;
[0029] S25. Using the ordered point set of the Mask circumscribed polygon obtained in S21, and using the transformation matrix calculated in S13, transform the ordered point set to obtain a transformed ordered point set;
[0030] S26. Calculate the bounding rectangle of the transformed ordered point set obtained in S25;
[0031] S27. Obtain the coordinates of the circumscribed rectangle through S26, perform cropping on the corrected image obtained in S14, and accurately crop the required area from the corrected image to obtain the final corrected image.
[0032] Furthermore, step S3 includes the following steps:
[0033] S31.SRGAN redefines the loss function and names it perceptual loss. SR It consists of two parts: in, For content loss, For adversarial losses;
[0034] S32. Content loss A pre-trained 19-layer VGG network was used, in which the ReLU activation layer was used as the basis for the VGG loss calculation, which involved the calculation of the Euclidean distance between feature representations. After training, feature maps were extracted from specific layers of the VGG model, which were then compared and analyzed with the actual image; the content loss of the VGG model The calculation formula is expressed as:
[0035]
[0036] in: It represents the feature map output after the convolution layer a of the VGG network and before the maximum pooling layer b; c represents the column index, and the range of c is from 1 to W a,b , W a,b Represents the feature map in the VGG network The width of the row, d represents the row index, and the range of d is from 1 to H a,b , H a,bRepresentation feature map Height; G θG is a mapping function from a low-resolution image to a high-resolution image, I LR It is a low-resolution image. The part after the minus sign is This is the reconstructed high-resolution image, the part before the minus sign It is a real high-resolution image;
[0037] S33. Adversarial Loss in: is the reconstructed image, is the probability of a natural HR image, the part in log is the output of the discriminator for generating super-resolution images, n is the number of feature maps, and N is the number of samples;
[0038] S34. The initial step includes taking a low-resolution image as input, which first passes through a convolutional layer, which processes the image using 64 9*9 filters. The output of this convolutional layer is fed into a parameterized ReLU function for nonlinear transformation. The processed data flows to multiple residual blocks, which contain a series of standard operations and form the core of the network. Each residual block contains a convolutional layer with 64 channels and 3*3 filters, followed by a parameterized ReLU layer and a batch normalization layer, followed by another convolution layer, followed by batch normalization. The last step of each residual block is to perform an element-wise summation operation on the input of the block and its output. The output of the block is passed to the next residual block, and the above sequence of operations within the residual block is repeated until the last residual block is processed. Finally, the network ends with a convolutional layer to generate a high-resolution super-resolution image as output;
[0039] S35. The discriminator uses a convolutional neural network model, which first applies a convolutional layer to the input image to extract features from the image. The extracted features are then passed to the Leaky ReLU function for nonlinear transformation. The core of the discriminator consists of multiple blocks, each of which contains a convolutional layer, a batch normalization layer, and a Leaky ReLU layer. These blocks act continuously on the image data to further process and refine the features. After this series of processing, the image data is directed to the final dense layer network, which includes a dense layer, followed by a Leaky ReLU layer, and then another dense layer. Through this process, the discriminator can effectively evaluate and analyze the input image and finally output the evaluation result of a high-resolution image.
[0040] Furthermore, the step S4 comprises the following steps:
[0041] S41. Use the Canny operator to detect the edge of the book spine and obtain a binary edge image;
[0042] S42. Perform Hough line transform on the edge image to calculate the angle of spine inclination. The principle of Hough line transform is as follows: In the (u, v) coordinate space of the image, the line passing through the point (u i , v i ) is represented by: c =-u i k c +v i ,y c represents the dependent variable, u i represents the independent variable, k c It represents the slope of the line;
[0043] S43. The original image is rotated to ensure that the book spine is perpendicular to the horizontal plane in the image. This process takes into account the actual placement of the book spine in the bookcase. Therefore, the range of the rotation angle θ is set to θ∈[-π / 2, π / 2). In order to systematize the rotation operation, we set every 10 degrees as a rotation unit. Therefore, for the fth rotation operation, the rotated image is represented as I f , combining all the rotation operations, we get a set of all the rotated images, denoted as I f = {I f |f=1,2,...,18}.
[0044] Furthermore, the step S5 comprises the following steps:
[0045] S51. Use the Mask-RCNN model to detect and segment each image in the image set. Assume that represents the number of bounding boxes and mask blocks detected and segmented in the rth image. The spine bounding boxes and mask blocks in the image are represented by the set B = {B r |r=1,2,...,8} and set M={M r |r=1, 2, ..., 18};
[0046] S52. Specific rules are used to filter out invalid bounding boxes and mask blocks to ensure that these elements do not negatively affect the final result. For valid mask blocks, two key rules are defined:
[0047] ① The width of the bounding box must be less than the preset maximum threshold maxD. If the bounding box is from the upper left vertex (x 1 ,y 1 ) and the lower right vertex (x 2 ,y 2 ) definition, then it is expressed as x 2 -x 1 ≤maxD,
[0048] ② The ratio of the area covered by the segmentation mask to the area of its corresponding bounding box needs to be greater than a specific ratio τ. Here, assuming that the width of the predicted image is Ww and the height is Hh, the mask area S m and the bounding box area S b Respectively expressed as:
[0049]
[0050] Among them: the function f(is, js) represents the predicted segmentation value of the pixel point (is, js) in the image. If it is a book spine class, it is 1, and if it is not a book spine class, it is 0. Rule ② is expressed as
[0051] The spine mask set M obtained for each rotated image is , using the above ①② rules for filtering, since the predicted spine mask is not segmented on the original image, each mask block needs to be rotated, where the rotation angle is the opposite of the previous rotation angle, and finally a new mask block set M is obtained for each image js Then the total set of mask blocks is M js ;
[0052] S53. When processing images that have been rotated multiple times, there is a possibility that the same book spine instance may be repeatedly segmented in multiple rotated images. After S43, the same instance will overlap, so the segmentation mask set M needs to be js Deduplication, the deduplication process involves comparing the set M js The overlapping area Sc between any two mask blocks in is , if the overlapping area of two mask blocks exceeds a certain threshold, they are judged to belong to the same book spine instance, and deduplication processing is performed accordingly.
[0053] Furthermore, step S6 includes the following steps:
[0054] S61. Locate the text in the spine image, and then the OCR system parses and converts the text to convert the text content in the image into an editable and searchable text format;
[0055] S62. Search for the book in the library lending system according to the book number. If the book is found, verify it with the book title. If the book title is the same, it proves that the book is found, and returns the location of the book to the borrower.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention mainly relies on miniature camera hardware devices. These devices are used to collect images of book spines from bookshelves, and then the algorithm in the present invention is applied to detect and identify the spines of these images, so as to accurately obtain the relevant information of each book and its specific location on the bookshelf. With the rapid development of artificial intelligence and image recognition technology, the cost of cameras is gradually decreasing, especially miniature cameras. Therefore, the method proposed in the present invention has significant cost-effectiveness. At the same time, due to its flexibility and convenience, the method is not easily affected by external environmental factors. It is worth noting that the method does not require any form of modification to the books, realizing truly non-perceptual use, that is, users can take and place books at will without affecting the operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a flow chart of the present invention.
[0059] Figure 2 Model diagram of the SRGAN network used to increase image pixels.
[0060] Figure 3 Flowchart of the Mask-RCNN model for spine image segmentation network. DETAILED DESCRIPTION
[0061] The present invention will be further described below in conjunction with the embodiments.
[0062] The present invention discloses a method for intelligent inventory counting and shelving assistance in a library based on machine vision, comprising the following steps: firstly, by deploying an image acquisition device (camera) on a library bookshelf, image information of books on adjacent bookshelves is obtained; secondly, preprocessing operations are performed on the obtained images, including image enhancement and image restoration, etc., in order to improve the clarity of the image and correct the distorted image morphology; then, the Mask-RCNN model is used to segment the image data, so as to accurately define the storage area of each book spine, and further identify specific areas containing key information such as the book title, author name and book number; finally, by taking the book number as a key index, the system is integrated with the library lending system, so as to improve the retrieval efficiency and precise positioning of books, and realize efficient management of books.
[0063] Example 1
[0064] A method for intelligent inventory counting and shelf arrangement assistance in a library based on machine vision comprises the following steps:
[0065] S1. Reconstruct the library bookshelves by deploying cameras, specifically:
[0066] S11. Refer to the bookshelf of Nanjing University of Aeronautics and Astronautics Library. The bookshelf width is 1000mm, the single-sided bookshelf depth is 300mm, the height is 2000mm, and one layer can store about 50 books;
[0067] S12. Install two Xiaomi cameras with 3 million pixels on a six-story bookshelf, with one camera installed on the top of the bookshelf and the other installed in the middle between the third and fourth layers of the bookshelf. This layout ensures that the two cameras can cover and capture the image information of all books on the adjacent bookshelves.
[0068] S2. Image restoration processing based on image edge points, specifically:
[0069] S12. Considering that the distance between adjacent bookshelves is relatively close, the captured images are not only affected by the light, but also by the tilt of the book spines due to the overhead shooting, so the images need to be preprocessed;
[0070] S21. Standardize the detection image, use the trained model to detect the standardized distorted image, obtain the rectangular detection frame of the target, multiple pairs of Bezier control points of the target edge and the Mask of the target, and the target detection frame Rect is defined by the coordinates of the upper left corner point and the width and height of the circumscribed rectangle: Rect = x, y, w, h, x is the horizontal coordinate of the upper left corner point of the target detection frame Rect, y is the vertical coordinate of the upper left corner point of the target detection frame Rect, w is the width of the circumscribed rectangle, and h is the height of the circumscribed rectangle;
[0071] The result of control point regression in the network model is the normalized relative distance from the control point to the upper left corner of the document detection rectangle Rect, x j is the horizontal coordinate of the jth control point, y j is the ordinate of the jth control point, Δx and Δy are the results of control point regression in the MaskRCNN network model, and d x d y is the normalized relative distance, w x 、w y is the standardized weight value; calculated as:
[0072] Δx=x j -x
[0073] Δy=y j -y
[0074]
[0075]
[0076] The coordinate x of the jth control point is obtained j ,y jThe relative distance d obtained by regression x , d y And the document rectangle detection box obtained by regression is calculated:
[0077] S22. The coordinates of the points on the target edge curve are calculated through the control points of the Bezier curve. These initial points are used to calculate the rectangular template after image correction. The corresponding coordinate points on the rectangular template are the target points. These initial points and target points correspond to each other up and down and left and right. When constructing the rectangular template, this study selected the same number of points as the corresponding edge points in the original image as the left and right edge points of the template. The coordinates of these points in the corrected image are defined as target points. Assuming that the number of edge point pairs is k, the total number of edge points is 2k. w is the width of the circumscribed rectangle, and h is the height of the circumscribed rectangle. At the left boundary of the template, k points are evenly selected from top to bottom. The edge points of the left boundary of the rectangular template are represented as: 0, Similarly, k edge points are also selected at the corresponding position of the right boundary. The points on the right boundary of the rectangular template are as shown in formula 0.
[0078] S23. Use the TPS transformation algorithm for the initial point and target point obtained in S12, calculate the transformation matrix of the transformation algorithm, and define two corresponding point sets, namely point set S and point set T: S = m 1, n 1 , ..., m n , n n , T = p 1 ,q 1 , .., p n ,q n , the S point set is called the template point set, the T point set is the target point set, and m 1 , n 1 , ..., m n , n n is the coordinate of the template point set, p 1 ,q 1 , .., p n ,q n are the coordinates of the target point set, which are all control points of the TPS algorithm;
[0079] S24. Use the transformation matrix obtained in S23 to transform the original image in S11 to obtain a corrected target image;
[0080] S25. Using the ordered point set of the Mask circumscribed polygon obtained in S21, and using the transformation matrix calculated in S13, transform the ordered point set to obtain a transformed ordered point set;
[0081] S26. Calculate the bounding rectangle of the transformed ordered point set obtained in S25;
[0082] S27. Obtain the coordinates of the circumscribed rectangle through S26, perform cropping on the corrected image obtained in S14, and accurately crop the area from the corrected image to obtain the final corrected image.
[0083] S3. Image enhancement processing based on SRGAN, specifically:
[0084] S31.SRGAN redefines the loss function and names it perceptual loss. SR It consists of two parts: in, For content loss, For adversarial losses;
[0085] S32. Content loss A pre-trained 19-layer VGG network was used, in which the ReLU activation layer was used as the basis for the VGG loss calculation, which involved the calculation of the Euclidean distance between feature representations. After training, feature maps were extracted from specific layers of the VGG model, which were then compared and analyzed with the actual image; the content loss of the VGG model The calculation formula is expressed as:
[0086]
[0087] in: It represents the feature map output after the convolution layer a of the VGG network and before the maximum pooling layer b; c represents the column index, and the range of c is from 1 to W a,b , W a,b Represents the feature map in the VGG network The width of the row, d represents the row index, and the range of d is from 1 to H a,b , H a,b Representation feature map Height; G θG is a mapping function from a low-resolution image to a high-resolution image, I LR It is a low-resolution image. The part after the minus sign is This is the reconstructed high-resolution image, the part before the minus sign It is a real high-resolution image;
[0088] S33. Adversarial Loss in: is the reconstructed image, is the probability of a natural HR image, the part in log is the output of the discriminator for generating super-resolution images, n is the number of feature maps, and N is the number of samples;
[0089] S34. The initial step includes taking a low-resolution image as input, which first passes through a convolutional layer, which processes the image using 64 9*9 filters. The output of this convolutional layer is fed into a parameterized ReLU function for nonlinear transformation. The processed data flows to multiple residual blocks, which contain a series of standard operations and form the core of the network. Each residual block contains a convolutional layer with 64 channels and 3*3 filters, followed by a parameterized ReLU layer and a batch normalization layer, followed by another convolution layer, followed by batch normalization. The last step of each residual block is to perform an element-wise summation operation on the input of the block and its output. The output of the block is passed to the next residual block, and the above sequence of operations within the residual block is repeated until the last residual block is processed. Finally, the network ends with a convolutional layer to generate a high-resolution super-resolution image as output;
[0090] S35. The discriminator uses a convolutional neural network model, which first applies a convolutional layer to the input image to extract features from the image. The extracted features are then passed to the Leaky ReLU function for nonlinear transformation. The core of the discriminator consists of multiple blocks, each of which contains a convolutional layer, a batch normalization layer, and a Leaky ReLU layer. These blocks act continuously on the image data to further process and refine the features. After this series of processing, the image data is directed to the final dense layer network, which includes a dense layer, followed by a Leaky ReLU layer, and then another dense layer. Through this process, the discriminator can effectively evaluate and analyze the input image and finally output the evaluation result of a high-resolution image.
[0091] S4. Preprocessing of the book image for rotation correction based on Hough line transform, specifically:
[0092] S41. Use the Canny operator to detect the edge of the book spine and obtain a binary edge image;
[0093] S42. Perform Hough line transform on the edge image to calculate the angle of spine inclination. The principle of Hough line transform is as follows: In the (u, v) coordinate space of the image, the line passing through the point (u i , v i ) is represented by: c =-u i k c +v i ,y c represents the dependent variable, u i represents the independent variable, k c It represents the slope of the line;
[0094] S43. The original image is rotated to ensure that the book spine is perpendicular to the horizontal plane in the image. This process takes into account the actual placement of the book spine in the bookcase. Therefore, the range of the rotation angle θ is set to θ∈[-π / 2, π / 2). In order to systematize the rotation operation, we set every 10 degrees as a rotation unit. Therefore, for the fth rotation operation, the rotated image is represented as I f , combining all the rotation operations, we get a set of all the rotated images, denoted as I f = {I f |f=1,2,...,18}.
[0095] S5. Use the Mask-RCNN model to perform image segmentation on the preprocessed data to obtain the spine image of a single book, specifically:
[0096] S51. Use the Mask-RCNN model to detect and segment each image in the image set. Assume that represents the number of bounding boxes and mask blocks detected and segmented in the rth image. The spine bounding boxes and mask blocks in the image are represented by the set B = {B r |r=1,2,...,8} and set M={M r |r=1, 2, ..., 18};
[0097] S52. Specific rules are used to filter out invalid bounding boxes and mask blocks to ensure that these elements do not negatively affect the final result. For valid mask blocks, two key rules are defined:
[0098] ① The width of the bounding box must be less than the preset maximum threshold maxD. If the bounding box is from the upper left vertex (x 1 ,y 1 ) and the lower right vertex (x 2 ,y 2 ) definition, then it is expressed as x 2 -x 1 ≤maxD,
[0099] ② The ratio of the area covered by the segmentation mask to the area of its corresponding bounding box needs to be greater than a specific ratio τ. Here, assuming that the width of the predicted image is Ww and the height is Hh, the mask area S m and the bounding box area S b Respectively expressed as:
[0100]
[0101] Among them: the function f(is, js) represents the predicted segmentation value of the pixel point (is, js) in the image. If it is a book spine class, it is 1, and if it is not a book spine class, it is 0. Rule ② is expressed as
[0102] The spine mask set M obtained for each rotated image is , use the above ①② rules to filter. Since the predicted spine mask is not segmented on the original image, each mask block needs to be rotated, where the rotation angle is the opposite of the previous rotation angle. For each image, a new mask block set M is finally obtained. js Then the total set of mask blocks is M js ;
[0103] S53. When processing images that have been rotated multiple times, there is a possibility that the same book spine instance may be repeatedly segmented in multiple rotated images. After S43, the same instance will overlap, so the segmentation mask set M needs to be js Deduplication, the deduplication process involves comparing the set M js The overlapping area Sc between any two mask blocks in is , if the overlapping area of two mask blocks exceeds a certain threshold, they are judged to belong to the same book spine instance, and deduplication processing is performed accordingly.
[0104] S6. Obtain the book title and book number in the image through OCR, and integrate the library lending system with the book title and book number as key indexes; specifically:
[0105] S61. Locate the text in the spine image, and then the OCR system parses and converts the text to convert the text content in the image into an editable and searchable text format;
[0106] S62. Search for the book in the library lending system according to the book number. If the book is found, verify it with the book title. If the book title is the same, it proves that the book is found, and returns the location of the book to the borrower.
[0107] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for intelligent inventory and shelf arrangement assistance in libraries based on machine vision, which is characterized by: The following steps are involved: S1. The library bookshelves are transformed by deploying cameras, and two cameras are used to capture the image information of all books on adjacent bookshelves; S2. Restore the book image information obtained by S1 based on the image edge points; S3. Enhance the restored book image based on SRGAN; S4. Preprocessing the book image for rotation correction based on Hough line transform to obtain a set of all rotated image sets of the book image; S5. Use the Mask-RCNN model to perform image segmentation on the preprocessed data to obtain the spine image of a single book; The step S5 comprises the following steps: S51. Use the Mask-RCNN model to detect and segment each image in the image set, and set the spine bounding box and mask block of the rth image to be represented as a set B = {B r |r=1,2,...,8} and set M={M r |r=1,2,...,18}; S52. Specific rules are used to filter out invalid bounding boxes and mask blocks; The spine mask set M obtained for each rotated image is , use the above rules to filter. Since the predicted spine mask is not segmented on the original image, each mask block needs to be rotated, where the rotation angle is the opposite of the previous rotation angle. For each image, a new mask block set M is finally obtained. js Then the total set of mask blocks is M js ; S53. When processing an image that has been rotated multiple times, the deduplication process involves comparing the set M js The overlapping area Sc between any two mask blocks in , if the overlapping area of two mask blocks exceeds a certain threshold, they are judged to belong to the same book spine instance, and deduplication processing is performed accordingly; S6. Obtain the book title and book number in the image through OCR, and integrate the library lending system with the book title and book number as key indexes.
2. The method for intelligent inventory counting and shelf arrangement assistance in libraries based on machine vision according to claim 1 is characterized in that: The step S2 comprises the following steps: S21. Considering that the distance between adjacent bookshelves is close, the captured image is not only affected by the light, but also by the tilt of the book spine due to the overhead shooting, so the original image needs to be preprocessed; S22. Standardize the detection image, use the trained model to detect the standardized distorted image, obtain the rectangular detection frame of the target, multiple pairs of Bezier control points on the edge of the target and the Mask of the target, the target detection frame Rect is defined by the coordinates of the upper left corner point and the width and height of the circumscribed rectangle: Rect = x, y, w, h, x is the horizontal coordinate of the upper left corner point of the target detection frame Rect, y is the vertical coordinate of the upper left corner point of the target detection frame Rect, w is the width of the circumscribed rectangle, and h is the height of the circumscribed rectangle; The result of control point regression in the network model is the normalized relative distance from the control point to the upper left corner of the document detection rectangle Rect, x j is the horizontal coordinate of the jth control point, y j is the ordinate of the jth control point, △x and △y are the results of control point regression in the MaskRCNN network model, and d x ,d y is the normalized relative distance, w x 、w y is the standardized weight value; calculated as: Δx=x j -x Δy=y j -y The coordinate x of the jth control point is obtained j ,y j The relative distance d obtained by regression x ,d y And the document rectangle detection box obtained by regression is calculated: S23. Calculate the coordinates of the points on the target edge curve through the control points of the Bezier curve, calculate the initial points of these nodes, and calculate the rectangular template after image correction. The corresponding coordinate points on the rectangular template are the target points. These initial points and target points correspond to each other up and down and left and right. When constructing the rectangular template, a set of points with the same number of corresponding edge points as those in the original image is selected as the left and right edge points of the template. The coordinates of these points in the corrected image are defined as target points. Let the number of edge point pairs be k, then the total number of edge points is 2k, w is the width of the circumscribed rectangle, h is the height of the circumscribed rectangle, and k points are evenly selected from top to bottom on the left boundary of the template. The edge points on the left boundary of the rectangular template are represented as: 0, Similarly, k edge points are also selected at the corresponding position of the right boundary. The points on the right boundary of the rectangular template are as shown in formula 0. S24. Use the TPS transformation algorithm for the initial point and target point obtained in S23, calculate the transformation matrix of the transformation algorithm, and define two corresponding point sets, namely point set S and point set T: S = m1, n1, ..., m n ,n n , T=p1,q1,...,p n ,q n , the S point set is called the template point set, the T point set is the target point set, m1,n1,...,m n ,n n is the coordinate of the template point set, p1,q1,...,p n ,q n are the coordinates of the target point set, which are all control points of the TPS algorithm; S25. Use the transformation matrix obtained in S24 to transform the original image in S21 to obtain a corrected target image; S26. Using the ordered point set of the Mask circumscribed polygon obtained in S22, and using the transformation matrix calculated in S24, the ordered point set is transformed to obtain a transformed ordered point set; S27. Calculate the bounding rectangle of the transformed ordered point set obtained in S26; S28. Obtain the coordinates of the circumscribed rectangle through S27, perform cropping on the corrected image obtained in S25, and accurately crop the required area from the corrected image to obtain the final corrected image.
3. The method for intelligent inventory counting and shelf arrangement assistance in libraries based on machine vision according to claim 2 is characterized in that: The step S3 comprises the following steps: S31.SRGAN redefines the loss function and names it perceptual loss. SR It consists of two parts: in, For content loss, For adversarial losses; S32. Content loss A pre-trained 19-layer VGG network was used, in which the ReLU activation layer was used as the basis for the VGG loss calculation, which involved the calculation of the Euclidean distance between feature representations. After training, feature maps were extracted from specific layers of the VGG model, which were then compared and analyzed with the actual image; the content loss of the VGG model The calculation formula is expressed as: in: It represents the feature map output after the convolution layer a of the VGG network and before the maximum pooling layer b; c represents the column index, and the range of c is from 1 to W a,b , W a,b Represents the feature map in the VGG network The width of the row, d represents the row index, and the range of d is from 1 to H a,b , H a,b Representation feature map Height; G θG is a mapping function from a low-resolution image to a high-resolution image, I LR It is a low-resolution image. The part after the minus sign is This is the reconstructed high-resolution image, the part before the minus sign It is a real high-resolution image; S33. Adversarial Loss in: is the reconstructed image, is the probability of a natural HR image, the part in log is the output of the discriminator for generating super-resolution images, n is the number of feature maps, and N is the number of samples; S34. The initial step includes taking a low-resolution image as input. The input image first passes through a convolutional layer. The convolutional layer processes the image using 64 9*9 filters. The output of the convolutional layer is fed into a parameterized ReLU function for nonlinear transformation. The processed data flows to multiple residual blocks. These blocks contain a series of standard operations and form the core of the network. Each residual block contains a convolutional layer with 64 channels and 3*3 filters, followed by a parameterized ReLU layer and a batch normalization layer, followed by another convolution layer, followed by batch normalization. The last step of each residual block is to perform an element-wise summation operation on the input of the block and its output. The output of the block is passed to the next residual block. The above sequence of operations within the residual block is repeated until the last residual block is processed. Finally, the network ends with a convolutional layer to generate a high-resolution super-resolution image as output; S35. The discriminator uses a convolutional neural network model, which first applies a convolutional layer to the input image to extract features from the image. The extracted features are then passed to the Leaky ReLU function for nonlinear transformation. The core of the discriminator consists of multiple blocks, each of which contains a convolutional layer, a batch normalization layer, and a Leaky ReLU layer. These blocks act continuously on the image data to further process and refine the features. After this series of processing, the image data is directed to the final dense layer network, which includes a dense layer, followed by a Leaky ReLU layer, and then another dense layer. Through this process, the discriminator can effectively evaluate and analyze the input image and finally output the evaluation result of a high-resolution image.
4. The method for intelligent inventory counting and shelf arrangement assistance in libraries based on machine vision according to claim 3 is characterized in that: The step S4 comprises the following steps: S41. Use the Canny operator to detect the edge of the book spine and obtain a binary edge image; S42. Perform Hough line transform on the edge image to calculate the angle of spine tilt. The principle of Hough line transform is as follows: In the (u, v) coordinate space of the image, the line passing through the point (u i ,v i ) is represented by: c =-u i k c +v i ,y c represents the dependent variable, u i represents the independent variable, k c It represents the slope of the line; S43. The original image is rotated to ensure that the spine is perpendicular to the horizontal plane in the image. The range of the rotation angle θ is set to θ∈[-π / 2, π / 2). In order to systematize the rotation operation, every 10 degrees is set as a rotation unit. For the fth rotation operation, the rotated image is represented as I f , combining all the rotation operations, we get a set of all rotated images, represented by I f = {I f |f=1,2,...,18}.
5. The method for intelligent inventory counting and shelf arrangement assistance in libraries based on machine vision according to claim 4 is characterized in that: The step S6 comprises the following steps: S61. Locate the text in the spine image, and then the OCR system parses and converts the text to convert the text content in the image into an editable and searchable text format; S62. Search for the book in the library lending system according to the book number. If the book is found, verify it with the book title. If the book title is the same, it proves that the book is found, and returns the location of the book to the borrower.
Citation Information
Patent Citations
Spine extraction method and device of vision-based book inventory system
CN111368856A
Book checking method based on computer vision
CN114882483A