Intelligent library seat state monitoring method based on YOLO

Through the YOLO-based intelligent library seat status monitoring method, combined with database recording and deep learning object detection technology, the problems of high manpower occupation and low management efficiency in library seat management are solved, and efficient and low-cost seat status monitoring and management are achieved.

CN120356146APending Publication Date: 2025-07-22KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337675.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing library seat status monitoring methods have problems such as high manpower occupation, high management costs and low management efficiency. The existing technologies such as infrared sensors are expensive, RFID technology is difficult to popularize, facial recognition system technology is difficult and may encounter social resistance.

Method used

Using the YOLO-based intelligent library seat status monitoring method, by recording seat and user information in the database, using optical character recognition technology OCR to identify seat numbers, and combining deep learning object detection technology, the monitoring logic of seat occupation, seat arrival and return seat is realized, reducing hardware requirements, and improving recognition accuracy and management efficiency.

Benefits of technology

It can efficiently and accurately detect and handle library seating situations without administrator intervention, reduce management costs, improve seat resource utilization and management efficiency, and is suitable for general GPU or CPU equipment and standard definition cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356146A_ABST
    Figure CN120356146A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent library seat state monitoring method based on YOLO, and belongs to the technical field of computer vision. Firstly, library seat information, use conditions, user information, seat occupation records and other data are recorded in a database. Secondly, performing preprocessing and seat number recognition on the image by using an optical character recognition (OCR) technology; and then, through a deep learning-based YOLO target detection technology, the library seat condition is detected in real time. And finally, designing seat occupation monitoring, seat arrival monitoring and seat return monitoring logics to detect and process seat using behaviors of the user, thereby realizing intelligent management of the library seat state. According to the invention, the YOLO algorithm is applied to library seat state monitoring, automatic detection and processing of seat occupying behaviors are realized through image acquisition equipment and background data processing, the library seat occupying behaviors can be efficiently and accurately monitored and processed without manual intervention, the management cost is reduced, and the library service quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent library seat status monitoring method based on YOLO, belonging to the technical field of computer vision. Background Art

[0002] With the continuous improvement of educational concepts, current university education has gradually focused on cultivating students' autonomous learning ability. As a hub of higher education resource literature, the library integrates collection, retrieval, borrowing, and reading services, and adopts an open management mode, naturally becoming an ideal place for students to explore and learn on their own. Its quiet and comfortable environment and strong learning atmosphere have attracted a large number of students to choose to study here.

[0003] However, with more and more students choosing to study in the library, seat resources have become increasingly scarce, and seat occupancy behaviors have become more and more common. In order to ensure a place to study, many students have to arrive at the library long in advance to grab seats, and some even use personal items to occupy seats for a long time and do not give up the resources even when temporarily leaving. The phenomenon of "one seat is hard to get" has not only become the focus of students' daily complaints but also questioned the service quality of the library. This situation not only affects students' normal right to use the library but also intensifies the contradictions among students to a certain extent, as well as the tense relationship between students and library management and even school management.

[0004] Solving the problem of library seat shortage and improving resource utilization and service efficiency are the main purposes of the present invention. Traditional solutions such as seat reservation systems, extended opening hours, and strengthened administrator supervision have problems of large workload, high manpower requirements, and low efficiency. The development of information technology and intelligent technology has brought new methods, such as mobile sensing, seat pressure sensing, infrared sensors or RFID technology, and camera face recognition, etc., to improve management efficiency. However, these methods have their own disadvantages: infrared sensors are expensive and vulnerable to environmental interference; RFID technology requires specific hardware support and is difficult to popularize; the face recognition system requires a large amount of face information, has high technical difficulty, and may encounter social resistance. Therefore, the present invention aims to achieve real-time monitoring and management of seat status in a more efficient and intelligent way, reduce manpower occupancy in the management process, lower management costs, and improve management efficiency. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent library seat status monitoring method based on YOLO, aiming to solve the technical problems of high manpower occupancy, high management cost, and low management efficiency existing in the existing library seat status monitoring methods.

[0006] To achieve the above object, the technical solution of the present invention is: Based on the field of computer vision, an efficient and low-complexity library seat status monitoring method is realized. Specifically, an intelligent library seat status monitoring method based on YOLO with high accuracy and low cost is provided. The specific steps are as follows:

[0007] Step1: Record library seat information and user information in the database;

[0008] Step2: Use optical character recognition technology OCR to recognize seat numbers;

[0009] Step3: Through object detection technology based on deep learning, recognize the occupancy status of library seats;

[0010] Step4: Based on the verification result of the seat number and seat information, use the logic of seat occupancy monitoring, arrival monitoring, and return monitoring to detect the user's behavior of using the seat.

[0011] The specific steps of Step1 are as follows:

[0012] Step1.1: Record library seat information in the database, including seat number, current user ID, seat status. The seat status includes "idle", "in use", "occupied", "selecting seat", covering all states of library seats;

[0013] Step1.2: Record library user information in the database, including user ID, user name, user login password, user permissions, user email, user occupancy status, last occupancy information, blacklist time. The user occupancy status includes five states: "0 occupancy times", "first occupancy", "second occupancy", "third occupancy", "blacklist". The state of "0 occupancy times" indicates the state where no occupancy behavior has occurred. The "blacklist" means that the user cannot perform seat selection operations within γ time in the future.

[0014] The specific steps of Step2 are as follows:

[0015] Step2.1: Obtain the video stream from the image acquisition device and extract a single-frame image as the processing object to provide an image basis for subsequent processing;

[0016] Step2.2: Convert the obtained color frame image into a grayscale image to further reduce the amount of data, reduce the complexity of image processing, and retain the key information of the image at the same time;

[0017] Step2.3: Perform an inversion operation on the grayscale image to convert the white seat number into black and the black background into white, forming an effect of black characters on a white background to enhance the contrast between the text and the background;

[0018] Step2.4: Perform adaptive histogram equalization to enhance the contrast of the inverted image, and use Gaussian filtering to denoise the image, retaining the edge information of the seat number to improve the recognition accuracy of the seat number;

[0019] Step2.5: Perform adaptive threshold binarization on the denoised image. By adjusting the threshold, divide the image into foreground and background;

[0020] Specifically, the adaptive threshold binarization process adjusts the threshold according to the dynamic characteristics of the local area of the image, divides the image into multiple small areas, calculates a local threshold within the neighborhood of each small area, and divides the pixels within the neighborhood into foreground and background according to the threshold. The implementation formula is:

[0021] T(x,y) = gaussian(x,y) - C

[0022] where T(x,y) is the local threshold, gaussian(x,y) is the Gaussian weighted average value, and C is an adjustment parameter used to fine-tune the calculation of the local threshold, so as to better handle the illumination changes and noise in the image. The effect is to set the pixel values of the text area to the background color (black), and the pixel values of the remaining areas to the foreground color (white), improving the contrast of the text, and thus improving the recognition accuracy of subsequent OCR for the text.

[0023] Step2.6: Use an edge detection algorithm to perform edge detection on the binarized image and extract the contour information in the image;

[0024] Step2.7: Perform morphological transformation on the image, including dilation and erosion techniques, to enhance the connectivity and integrity of the edge contours and improve the recognition degree of the contours;

[0025] Step2.8: After the image processing operations are completed, use contour detection technology to extract contours;

[0026] Step2.8.1: Each seat has a seat number box. The seat number box sets an outer contour, and the outer contour is responsible for locating the area in the image that conforms to the characteristics of the seat number box. Define the aspect ratio of the outer contour according to the aspect ratio of the seat number box to reduce the interference of other box items in the image on the recognition of the seat number box;

[0027] Step2.8.2: On the basis of locating the outer contour, define the aspect ratio of the inner contour according to the aspect ratio of the seat number digits, and the inner contour locates the area in the seat number box that conforms to the seat number characteristics;

[0028] Step2.9: Use a dual aspect ratio positioning algorithm to achieve the positioning of the seat number area;

[0029] Step 2.10: Perform character recognition on the located seat number area to extract the seat number.

[0030] The specific steps of the said Step 2.9 are as follows:

[0031] Step 2.9.1: For the first aspect ratio positioning, screen the extracted contours, obtain the largest contour whose aspect ratio conforms to the library seat label, calculate the minimum rectangular boundary enclosing the specified contour, and intercept the image enclosed by it as the region of interest;

[0032] Step 2.9.2: For the second aspect ratio positioning, perform contour extraction on the region of interest, retain the regions where all contours with aspect ratios conforming to digital features are located, and mark them as potential digital regions;

[0033] Step 2.9.3: Exclude the potential digital regions with areas smaller than a predetermined threshold, and screen out the potential digital regions that meet the preset threshold conditions.

[0034] Through two rounds of contour screening involving double aspect ratios, determine the positions of the library seat number label and the seat number digits, and then intercept the seat number area image for subsequent precise character recognition and obtaining the seat number.

[0035] The specific steps of the said Step 3 are as follows:

[0036] Step 3.1: During the seat usage process, use an image acquisition device to collect real-time video information, extract frame images, scale them to the size of λ×λ, input the scaled images into the CNN network of YOLO, and use the Input (input preprocessing) part of the network structure to perform preprocessing operations on the pictures, including Mosaic data augmentation, adaptive anchor box calculation, and adaptive picture scaling;

[0037] Step 3.2: Use the Backbone part of the YOLO network structure to extract features from the preprocessed image, generating three feature maps with sizes of ɑ×ɑ×γ, β×β×δ, and ε×ε×δ respectively;

[0038] The said Backbone mainly uses the ELAN (Efficient Layer Aggregation) network structure and the MP (McCulloch-Pitts) model, aiming to improve the accuracy and robustness of the object detection algorithm by effectively aggregating the feature information of different layers;

[0039] The described ELAN is an efficient layer aggregation network structure that fuses feature information from different layers in a specific way, controls the shortest and longest gradient paths, and enhances the model's learning ability for target features. ELAN usually contains multiple branches, each branch processes feature maps of different sizes, and feature fusion is performed between branches through different methods such as concat (concatenation) and add (addition) to achieve the aggregation of multi-scale features. In addition, the network structure of ELAN can reduce the problems of gradient disappearance and gradient explosion, and improve the training stability and convergence speed of the model.

[0040] The described MP model refers to MaxPooling (maximum pooling), which is a commonly used pooling operation and a common downsampling technique in deep learning. MaxPooling usually selects the maximum value within a local area as the output, and this process performs a sliding window operation with a certain stride. YOLO uses the MP model to gradually reduce the size of the feature map so that subsequent layers can process higher-level semantic information and helps maintain the real-time processing ability of the network.

[0041] Step3.3: Use the Neck part of the YOLO network structure to perform feature fusion on the feature map, and obtain three enhanced feature layers through the FPN in the Backbone part.

[0042] Specifically, the Neck part of the YOLO network structure adopts the FPN+PAN structure. Neck fuses the three effective feature layers obtained from the backbone network to combine feature information of different scales. During the feature fusion process, not only upsampling of features is performed to achieve feature fusion, but also downsampling of features is performed again to achieve feature fusion. Through FPN, three enhanced effective feature layers can be obtained. Each feature layer contains width, height, and number of channels. At this time, each feature map can be regarded as a set of multiple feature points, and there are three prior boxes on each feature point, and each prior box has the number of channels (3 for RGB pictures obtained by image acquisition devices) of features.

[0043] The FPN refers to the Feature Pyramid Network, which aims to solve the problem of multi-scale feature fusion. Its working process is as follows: First, extract feature layers through a convolutional neural network to form the basis of the feature pyramid; second, upsample the high-level features to make their resolutions consistent with the low-level features; finally, perform lateral connections to fuse the upsampled high-level features with the corresponding low-level features to enhance the semantic information of the low-level features.

[0044] The PAN mentioned above refers to the Path Aggregation Network, which is designed to enhance the feature pyramid ability and aims to improve the accuracy of object detection. It is extended based on the concept of FPN, not only realizing the top-down path but also adding a bottom-up path to better fuse feature information at different levels;

[0045] Step3.4: Use the Head part of the YOLO network structure to determine whether there is an object corresponding to the prior box on the feature point, obtain the size, position of the object bounding box, and the class score of the object prediction, and then output the obtained data by the output layer;

[0046] Step3.5: Adopt the Non-Maximum Suppression algorithm NMS to solve the problem of the object being detected multiple times. Perform iterative selection and suppression on each bounding box, and remove overlapping redundant prediction results according to the IoU (Intersection over Union) threshold. Finally, obtain the object detection result of the image.

[0047] The specific steps of Step4 are as follows:

[0048] Step4.1: After the user selects a seat number using the mobile device, the background database processes the seat information, registers the user's identity, and modifies the seat status to "Seat selection in progress";

[0049] Step4.2: For seats with the status of "Seat selection in progress", perform periodic arrival detection through the image acquisition device. For users who do not arrive at the seat within the specified σ time, the background will cancel the seat selection, give a reminder, and record the behavior of illegal use;

[0050] Step4.3: For seats where the user arrives normally, modify the seat status in the database to "In use", and perform periodic occupancy detection and return seat detection through the image acquisition device. For users who occupy the seat and do not return to the seat within the specified time α, the background will cancel the seat selection, give a reminder, and record the behavior of illegal use;

[0051] Specifically, according to the actual situation, for the situation where items are left after the user has vacated the seat and are detected, the background will remind the library administrator to handle it;

[0052] Step4.4: For users whose number of consecutive illegal use behaviors reaches the specified threshold β, the background modifies the user to the "Blacklist" status, sends a blacklist reminder message to the user, and this user cannot select a seat normally within the specified restricted time.

[0053] The present invention is highly efficient and can completely detect and process the situation of occupied seats in the library without any manual intervention by the administrator. The accuracy of target detection is high, and the robustness is strong, which can cover and quickly and accurately identify various occupied seat situations. At the same time, the present invention has low costs and low requirements for hardware. General GPUs or CPUs can run relevant algorithms, and a standard-definition camera can meet the performance requirements of the image acquisition device.

[0054] The beneficial effects of the present invention are as follows:

[0055] (1) The present invention proposes a target detection method to realize the monitoring of seat usage behavior, improve the recognition accuracy, and reduce the management cost;

[0056] (2) The present invention redefines the role of YOLO in the seat monitoring environment. It is no longer a pre-work for executing technologies such as face recognition, but as a part of the seat usage behavior monitoring algorithm, and completely participates in the discovery and processing of occupied seat behaviors. It realizes the monitoring of seat status in a way with lower complexity and high precision;

[0057] (3) The present invention designs the algorithm logic for occupied seat detection, arrival monitoring, and return seat detection. Through the periodic detection of the image acquisition device and the background data processing, it detects and processes occupied seat behaviors, and improves the management efficiency of library seats and the utilization rate of such scarce public resources as seats. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is the execution architecture diagram of the present invention;

[0059] Figure 2 is the flowchart of image preprocessing in the seat number recognition step of the present invention;

[0060] Figure 3 is the effect diagram of image preprocessing in the seat number recognition step of the present invention;

[0061] Figure 4 is the effect diagram of library monitoring of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0063] Example 1: As Figure 1 shown, an intelligent library seat status monitoring method based on YOLO specifically includes the following steps:

[0064] Step1: Record the library seat information and user information in the database.

[0065] Step1.1: Record the library seat information in the database, including seat number, current user ID, and seat status. The seat status includes "idle", "in use", "occupied", and "selecting a seat", covering all the states of library seats.

[0066] Step1.2: Record the library user information in the database, including user ID, user name, user login password, user permissions, user email, user occupancy status, last occupancy information, and blacklist time. The user occupancy status includes five states: "0 occupancy", "first occupancy", "second occupancy", "third occupancy", and "blacklist". The state of "0 occupancy" indicates that there has never been an occupancy behavior, and the "blacklist" means that the user cannot perform seat selection operations within the next γ time.

[0067] Step2: Use the optical character recognition technology OCR to recognize the seat number.

[0068] Step2.1: Obtain the video stream from the image acquisition device and extract a single-frame image as the processing object to provide an image basis for subsequent processing.

[0069] Step2.2: Convert the obtained color frame image into a grayscale image to further reduce the data volume, lower the image processing complexity, and at the same time retain the key information of the image.

[0070] Step2.3: Perform an inversion operation on the grayscale image to convert the white seat number into black and the black background into white, forming an effect of black characters on a white background and enhancing the contrast between the text and the background.

[0071] Step2.4: Take adaptive histogram equalization to enhance the contrast of the inverted image and use Gaussian filtering to denoise the image, retaining the edge information of the seat number and improving the recognition accuracy of the seat number.

[0072] Step2.5: Perform adaptive threshold binarization processing on the denoised image. By adjusting the threshold, the image is divided into foreground and background.

[0073] Specifically, the adaptive threshold binarization processing is to adjust the threshold according to the dynamic characteristics of the local area of the image, divide the image into multiple small areas, calculate a local threshold within the neighborhood of each small area, and divide the pixels within the neighborhood into foreground and background according to the threshold. The implementation formula is:

[0074] T(x,y)=gaussian(x,y)-C

[0075] Among them, T(x, y) is the local threshold, gaussian(x, y) is the Gaussian weighted average value, and C is an adjustment parameter used to fine-tune the calculation of the local threshold, so as to better handle the illumination changes and noise in the image. In this way, the pixel values in the text area are set to the background color (black), and the pixel values in the remaining areas are set to the foreground color (white), improving the contrast of the text, and thus improving the accuracy of subsequent OCR for text recognition.

[0076] Step2.6: Use an edge detection algorithm to perform edge detection on the binary image and extract the contour information in the image;

[0077] Specifically, the present invention uses the Canny algorithm to perform edge detection on the binary image. The Canny algorithm is used to identify edges in digital images, and its goal is to find an optimal edge detection solution or locate the positions with the strongest gray intensity changes in an image. The main steps of the Canny algorithm are Gaussian filtering for denoising, gradient calculation, non-maximum suppression, double-threshold detection, and edge tracking. Among them, Gaussian filtering for denoising is expressed as:

[0078] I′ = I × G

[0079] In the formula, I is the input image, G is the Gaussian kernel, and I′ is the image after Gaussian filtering. Gradient calculation is expressed as:

[0080]

[0081] In the formula, G x and G y are the partial derivatives of the Gaussian kernel G in the x direction and y direction respectively. M represents the gradient magnitude, and θ represents the gradient direction. The process of non-maximum suppression is: for each pixel, check the adjacent pixels in its gradient direction. If the gradient magnitude of the current pixel is not the local maximum in that direction, then suppress it (set it to 0). The process of double-threshold detection is: select two thresholds T high and T low , where T high > T low . According to these two thresholds, classify the image after non-maximum suppression. Pixels with a gradient magnitude greater than T high are defined as strong edges, pixels with a gradient magnitude between T high and T low are defined as weak edges, and pixels with a gradient magnitude less than T low are defined as non-edges. The process of edge retention is: only retain those weak edges that are directly or indirectly connected to strong edges, and the remaining weak edges are regarded as noise and removed;

[0082] Step2.7: Perform morphological transformation on the image, including dilation and erosion techniques, to enhance the connectivity and integrity of the edge contours and improve the recognition of the contours.

[0083] Step2.8: After the image processing operations are completed, use contour detection technology to extract contours.

[0084] Specifically, the present invention uses the cv2.findContours function to extract contours.

[0085] Step2.8.1: Each seat has a seat number box, and an outer contour is set for the seat number box. The outer contour is responsible for locating the area in the image that conforms to the characteristics of the seat number box. Define the aspect ratio of the outer contour according to the aspect ratio of the seat number box to reduce the interference of other box items in the image on the recognition of the seat number box.

[0086] Step2.8.2: On the basis of locating the outer contour, define the aspect ratio of the inner contour according to the aspect ratio of the seat number digits. The inner contour locates the area in the seat number box that conforms to the seat number characteristics.

[0087] Step2.9: Use a dual aspect ratio positioning algorithm to achieve the positioning of the seat number area.

[0088] Step2.9.1: The first aspect ratio positioning screens the extracted contours, obtains the largest contour whose aspect ratio conforms to the library seat label, calculates the minimum rectangular boundary enclosing the specified contour, and intercepts the image it encloses as the region of interest.

[0089] Step2.9.2: The second aspect ratio positioning extracts contours from the region of interest, retains the regions where all contours with aspect ratios conforming to the digit characteristics are located, and marks them as potential digit regions.

[0090] Step2.9.3: Exclude the potential digit regions with an area smaller than a predetermined threshold, and screen out the potential digit regions that meet the preset threshold conditions.

[0091] Through the two contour screenings involved in the dual aspect ratio, determine the positions of the library seat number labels and the seat number digits, and then intercept the seat number area image for subsequent accurate character recognition and obtaining the seat number.

[0092] Step2.10: Perform character recognition on the located seat number area to extract the seat number.

[0093] Step3: Use object detection technology based on deep learning to recognize the occupancy status of library seats.

[0094] Step3.1: During the use of the seat, collect real-time video information using an image acquisition device, extract frame images, scale them to a size of 640×640, input the scaled images into the CNN network of YOLO, and use the Input part of the network structure to perform preprocessing operations on the pictures, including Mosaic data augmentation, adaptive anchor box calculation, and adaptive picture scaling;

[0095] Step3.2: Use the Backbone part of the YOLO network structure to extract features from the preprocessed image, generating three feature maps with sizes of 80×80×512, 40×40×1024, and 20×20×1024 respectively;

[0096] Step3.3: Use the Neck part of the YOLO network structure to perform feature fusion on the feature maps, and through the FPN of the Backbone part, obtain three enhanced feature layers;

[0097] Step3.4: Use the Head part of the YOLO network structure to determine whether there is an object corresponding to the prior box on the feature point, obtain the size and position of the target bounding box and the class score of the target prediction, and then output the obtained data by the output layer;

[0098] Step3.5: Use the non-maximum suppression algorithm NMS to solve the problem of the target being detected multiple times, perform iterative selection and suppression on each bounding box, and remove overlapping redundant prediction results according to the IoU threshold, and finally obtain the target detection result of the image.

[0099] Step4: Based on the verification result of the seat number and seat information, use the logic of occupancy monitoring, arrival monitoring, and return monitoring to detect the user's behavior of using the seat.

[0100] Step4.1: After the user selects a seat using the mobile device, the background database processes the seat information, registers the user's identity, and modifies the seat status to "selecting seat";

[0101] Step4.2: For seats with the status of "selecting seat", perform periodic arrival detection through the image acquisition device. For users who do not arrive at the seat within the specified σ time, the background cancels the seat reservation, gives a reminder, and records the behavior of illegal use;

[0102] Step4.3: For seats where the user arrives normally, modify the seat status in the database to "in use", perform periodic occupancy detection and return detection through the image acquisition device. For users who occupy the seat and do not return to the seat within the specified time α, the background cancels the seat reservation, gives a reminder, and records the behavior of illegal use;

[0103] Specifically, according to the actual situation, in the case where an item is detected as remaining after the user has vacated the seat and left, the background will remind the library administrator to handle it;

[0104] Step4.4: For users whose number of consecutive violations reaches the specified threshold β, the background modifies the user to the "blacklist" status, sends a blacklist reminder message to the user, and this user is not allowed to select a seat normally within the specified restricted time.

[0105] Furthermore, in this embodiment, first, frame images are obtained from the video stream input by the video collector, and the frame images are used as the seat analysis object. Then, image preprocessing operations are performed on the images to reduce the influence of other irrelevant items in the images on the seat number recognition. After the preprocessing is completed, the area where the seat number is located is located by extracting two specific rectangular boxes with a width-to-height ratio of α:β. Then, the optical character recognition technology OCR is used to recognize the seat number from the image information obtained by the image acquisition device. After the seat number recognition is completed, the different states of the seats in the database are read, and the corresponding state monitoring logic algorithms are called, namely the arrival monitoring algorithm, the occupancy monitoring algorithm, and the return monitoring algorithm. At the same time, the YOLO algorithm is started to perform object detection. This detection first intercepts frame images from the video stream, performs image preprocessing and image segmentation, and inputs them into the backbone network of the object detection algorithm; in the backbone network, the ELAN network structure and the MP model are used to extract the feature information in the images and generate three feature maps of different sizes; then, the Neck of the algorithm extracts information from the feature maps of different scales through FPN, and further enhances the cross-layer feature fusion using PAN; next, the Head part of the algorithm judges the feature points and performs the prediction task, and outputs the object classification, confidence, and the position of the bounding box on its output layer; finally, NMS is used to eliminate redundant prediction results to complete the object detection. The specific execution results of the seat status monitoring logic algorithm are as Figure 4 shown.

[0106] Figure 2It is a flowchart of image preprocessing before OCR seat number recognition. First, a video stream of the library seat area is obtained through an image acquisition device, and a single-frame image is extracted from it as the processing object. This step is the basis for all subsequent image processing operations. Next, the color frame image is converted into a grayscale image, which can reduce the data volume, lower the complexity of image processing, and at the same time retain the key information of the image. Subsequently, an inversion operation is performed on the grayscale image to convert the original seat number (white) into black and the background (black) into white, forming an effect of black characters on a white background. This operation significantly enhances the contrast between the text and the background and facilitates the subsequent recognition steps. To further improve the image quality, contrast enhancement is performed on the inverted image. The adaptive histogram equalization technique is used to enhance the contrast of the image, and Gaussian filtering is used for denoising to retain the edge information of the seat number. These processing steps help improve the accuracy of subsequent OCR text recognition. After the image preprocessing is completed, the image is divided into foreground and background through adaptive threshold binarization. In this step, the threshold is adjusted according to the dynamic characteristics of the local area of the image, and the image is divided into multiple small areas. A local threshold is calculated in each neighborhood, and the pixels in the neighborhood are divided into foreground and background according to the threshold. Next, the Canny algorithm is used to perform edge detection on the binarized image to extract the contour information existing in the image. After the edge detection is completed, morphological transformation is performed on the image, including dilation and erosion operations, to enhance the connectivity and integrity of the edge contour and improve the recognition degree of the contour. Finally, the cv2.findContours function is used for contour extraction. The aspect ratio of the outer contour is defined according to the aspect ratio of the seat number box, and the aspect ratio of the inner contour is defined according to the aspect ratio of the seat number digits. Through two combinations of contour screening, the positioning of the seat number area is realized, and the image of the located seat number area is passed to OCR for further recognition operations. The specific implementation effect is as Figure 3 shown.

[0107] The specific implementation manners of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above implementation manners. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.

Claims

1. An intelligent library seat status monitoring method based on YOLO, characterized in that: Step1: Record library seat information and user information in the database; Step2: Use the optical character recognition technology OCR to recognize seat numbers; Step3: Identify the occupancy status of library seats through object detection technology based on deep learning; Step4: Based on the verification result of the seat number and seat information, use the logic of seat occupancy monitoring, arrival monitoring, and return monitoring to detect the behavior of users using seats.

2. The intelligent library seat status monitoring method based on YOLO according to claim 1, characterized in that The specific steps of Step1 are as follows: Step1.1: Record library seat information in the database, including seat number, current user ID, seat status, and the seat status includes "idle", "in use", "occupied", "selecting a seat", covering all states of library seats; Step1.2: Record library user information in the database, including user ID, user name, user login password, user permissions, user email, user occupancy status, last occupancy information, and blacklist time. The user occupancy status includes five states: "0 occupancy", "first occupancy", "second occupancy", "third occupancy", and "blacklist". The state of "0 occupancy" means that there has never been an occupancy behavior, and the "blacklist" means that the user cannot perform seat selection operations within γ time in the future.

3. A method for monitoring the seat status of an intelligent library based on YOLO according to claim 1, characterized in that, The specific steps of Step2 are as follows: Step2.1: Obtain the video stream from the image acquisition device and extract a single-frame image as the processing object; Step2.2: Convert the obtained color frame image into a grayscale image; Step2.3: Perform an inversion operation on the grayscale image to convert the white seat number into black and the black background into white, forming an effect of black characters on a white background; Step2.4: Take adaptive histogram equalization to enhance the contrast of the inverted image, and use Gaussian filtering to denoise the image, retaining the edge information of the seat number; Step2.5: Perform adaptive threshold binaryzation processing on the denoised image, and divide the image into foreground and background by adjusting the threshold; Step2.6: Use an edge detection algorithm to perform edge detection on the binary image and extract the contour information in the image; Step2.7: Perform morphological transformation on the image, including dilation and erosion techniques; Step2.8: After the image processing operation is completed, use contour detection technology to extract contours; Step2.8.1: Each seat has a seat number box, and the seat number box is set with an outer contour. The outer contour is responsible for locating the area in the image that conforms to the characteristics of the seat number box, and defines the aspect ratio of the outer contour according to the aspect ratio of the seat number box; Step2.8.2: On the basis of locating the outer contour, define the aspect ratio of the inner contour according to the aspect ratio of the seat number digits, and the inner contour locates the area in the seat number box that conforms to the seat number characteristics; Step2.9: Use a dual aspect ratio positioning algorithm to achieve the positioning of the seat number area; Step 2.10: Perform character recognition on the located seat number area to extract the seat number.

4. The intelligent library seat status monitoring method based on YOLO according to claim 3, wherein, The specific steps of Step 2.9 are as follows: Step 2.9.1: The first aspect ratio positioning filters the extracted contours to obtain the largest contour whose aspect ratio conforms to the library seat label, calculates the minimum rectangular boundary enclosing the specified contour, and intercepts the image enclosed by it as the region of interest. Step 2.9.2: The second aspect ratio positioning extracts contours from the region of interest, retains the regions where all contours conform to the digital characteristics in terms of aspect ratio, and marks them as potential digital regions. Step 2.9.3: Exclude the potential digital regions with an area smaller than a predetermined threshold, and screen out the potential digital regions that meet the preset threshold conditions.

5. The intelligent library seat status monitoring method based on YOLO according to claim 1, wherein, The specific steps of Step 3 are as follows: Step 3.1: During the seat usage process, collect real-time video information using an image acquisition device, extract frame images, scale them to the size of λ×λ, input the scaled images into the CNN network of YOLO, and use the Input part of the network structure to perform preprocessing operations on the pictures, including Mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling. Step 3.2: Use the Backbone part of the YOLO network structure to extract features from the preprocessed image, generating three feature maps with sizes of ɑ×ɑ×γ, β×β×δ, and ε×ε×δ respectively. Step 3.3: Use the Neck part of the YOLO network structure to perform feature fusion on the feature maps, and obtain three enhanced feature layers through the FPN of the Backbone part. Step 3.4: Use the Head part of the YOLO network structure to determine whether there is an object corresponding to the prior box at the feature point, obtain the size, position of the target bounding box, and the class score of the target prediction, and then output the obtained data by the output layer. Step 3.5: Use the non-maximum suppression algorithm NMS to perform iterative selection and suppression on each bounding box, eliminate overlapping redundant prediction results according to the IoU threshold, and finally obtain the object detection result of the image.

6. The intelligent library seat status monitoring method based on YOLO according to claim 1, wherein, The specific steps of Step 4 are as follows: Step 4.1: After the user selects a seat using the mobile device, the background database processes the seat information, registers the user's identity, and modifies the seat status to "seating selected". Step 4.2: For seats with the status of "seating selected", perform periodic arrival detection through the image acquisition device. For users who do not arrive at the seat within the specified σ time, the background vacates the seat, gives a reminder, and records the behavior of illegal use. Step 4.3: For seats where the user arrives normally, modify the seat status in the database to "in use", perform periodic occupancy detection and return detection through the image acquisition device. For users who occupy the seat and do not return within the specified time α, the background vacates the seat, gives a reminder, and records the behavior of illegal use. Step 4.4: For users whose number of consecutive violations reaches the specified threshold β, the background modifies the user's status to "blacklist", and the user cannot select seats normally within the specified restricted time.