Portable intelligent literature cataloguing system based on edge calculation AI

By using a portable intelligent document cataloging system based on edge computing AI, the problems of high cost, large space occupation, and poor flexibility of heavy-duty robotic cataloging systems have been solved. This system enables low-cost, flexible document cataloging in multiple scenarios, reduces equipment procurement and operating costs, and improves data processing speed and security.

CN121304079APending Publication Date: 2026-01-09莫少强
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511495637.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing heavy-duty robot cataloging systems are costly, space-consuming, inflexible, poorly adaptable to various scenarios, and complex to maintain, making them difficult to adapt to diverse work environments.

Method used

A portable intelligent cataloging system based on edge computing AI is adopted, including a handheld terminal and a PC. The handheld terminal has a built-in Rockchip RK3588S chip, a Sony IMX586 camera and a laser rangefinder sensor for local image processing and data preprocessing, and the PC generates standardized cataloging documents.

Benefits of technology

It achieves low cost, low space occupation, high flexibility, strong scenario adaptability, and simple maintenance. It can handle multiple document types, reduce equipment procurement and operating costs, and improve data processing speed and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304079A_ABST
    Figure CN121304079A_ABST
Patent Text Reader

Abstract

The invention discloses a portable literature intelligent cataloguing system based on edge calculation AI, and belongs to the technical field of intelligent libraries. The system comprises a hand-held terminal and a PC terminal. The handheld terminal is provided with a main camera, a light supplementing lamp and a laser distance measuring sensor and is used for collecting literature images and carrying out local AI processing; and the PC end is connected with the handheld terminal through a USB interface, receives the structured data sent by the handheld terminal, generates a metadata text file according to an MARC format, and associates and stores an image file. According to the method, optimized lightweight hardware is combined with local AI processing, a heavy artificial intelligence cataloguing system is replaced by a low-cost and small-size handheld device, flexible, efficient and offline literature intelligent cataloguing is achieved, the device cost, the space requirement and the maintenance complexity are remarkably reduced, and the method is suitable for popularization and application. The method is suitable for large, medium and small libraries, mobile literature interview and other non-standard environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent library document information processing technology, and specifically relates to a portable intelligent document cataloging device and its method and system based on edge computing AI. Background Technology

[0002] In the construction of smart libraries, the digitization and intelligent cataloging of document resources is a core foundation. Existing technologies, such as high-end intelligent cataloging systems (patent number CN202411153319.8), employ large industrial robots in automated production lines to automatically handle, turn pages, scan, and extract information from books. However, such solutions suffer from extremely high equipment procurement, deployment, and maintenance costs; require hundreds of square meters of fixed space; have poor adaptability to document size, thickness, and binding methods; are highly integrated, difficult to disassemble and move, and cannot adapt to non-standard work scenarios such as mobile book carts and temporary document acquisition points; and have complex mechanical structures requiring professional technicians for maintenance, resulting in high manufacturing and operating costs. Therefore, there is an urgent need for a low-cost, flexible, easy-to-maintain, and lightweight intelligent document cataloging solution that can adapt to various work scenarios. Summary of the Invention

[0003] This invention aims to solve the problems of high cost, large space occupation, poor flexibility, weak scene adaptability and complex maintenance of existing heavy robot cataloging systems, and provides a portable intelligent document cataloging system based on edge computing AI, which can achieve efficient and flexible multi-scene document cataloging with low hardware cost and small space requirements.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] A portable intelligent document cataloging system based on edge computing AI includes a handheld terminal (100) and a PC (200). The handheld terminal (100) and the PC (200) are physically connected and communicate with each other via a USB interface. The specific structure is as follows:

[0006] Handheld terminal (100)

[0007] This is used to collect document images and perform local AI recognition and data preprocessing, including:

[0008] 1. Core computing module (101): A system-on-a-chip with built-in NPU, the system-on-a-chip is Rockchip RK3588S, which can run optimized OCR neural network model and image classification neural network model to realize local real-time processing of document images;

[0009] 2. Image acquisition module (102): includes at least one main camera, a ring light, and a laser rangefinder; the main camera uses a Sony IMX586 48MP camera module for high-definition acquisition of document page images; the ring light is used to provide uniform illumination to avoid image reflection or shadows affecting recognition accuracy; the laser rangefinder is used to assist the main camera in quickly focusing and to determine the appropriate scanning distance between the handheld terminal (100) and the document (preferably a scanning distance of 15-25cm);

[0010] 3. Data communication module (103): It adopts a USB Type-C interface to transmit the structured data packets generated by the handheld terminal (100) to the PC (200), and supports high-speed data transmission (transmission rate not less than 100Mbps);

[0011] 4. Power module (104): It uses a 6000mAh rechargeable lithium battery to power the core computing module (101), image acquisition module (102) and data communication module (103). A single full charge can support continuous scanning of no less than 500 pages of documents.

[0012] 5. Human-computer interaction structure: The handheld terminal (100) adopts an ergonomic design. The right side is provided with a finger support groove and a soft anti-slip covering layer (the soft anti-slip covering layer is made of silicone material with a Shore hardness of 50-60HA). The top front end is provided with a physical scanning button with a pressing stroke of 0.8-1.2mm, which is convenient for users to operate with one hand and reduce fatigue during long-term use.

[0013] PC version (200)

[0014] Used to receive and parse data transmitted by a handheld terminal (100) and generate standardized catalog files, which include:

[0015] 1. Data receiving module (201): Built-in USB driver, supports recognition of the USB Type-C interface of the handheld terminal (100), can automatically receive structured data packets sent by the handheld terminal (100), and supports batch data reception (can receive data packets of no less than 100 cataloging tasks at a time);

[0016] 2. File generation module (202): used to parse structured data packets, generate .txt text files from the extracted metadata in MARC format, and save document images as .jpg format files; the MARC format text files contain the 010 International Standard Book Number field, the 200 Title and Author field, the 210 Publication field, and other core necessary fields, and the file header contains a unique key value used to associate all image files in the same scanning task;

[0017] 3. Local storage module (203): The generated files are stored using a preset directory structure. The preset directory includes a "Metadata" subdirectory (storing MARC format .txt text files) and an "Images" subdirectory (storing .jpg image files). The directory path can be set to a fixed path on the local hard drive of the PC (e.g., D:\ScanOutput\). It supports automatically naming files according to unique key values ​​(e.g., text files are named "ABC123.txt", and corresponding image files are named "ABC123_copyright.jpg" and "ABC123_cover.jpg").

[0018] Beneficial effects

[0019] Compared with existing technologies, this invention has significant advantages, specifically in terms of low cost, small space occupation, high flexibility, strong scenario adaptability, simple maintenance, and efficient data security: In terms of cost, the hardware cost is only one percent of that of traditional heavy-duty robotic cataloging systems, and there are no annual maintenance costs, which can significantly reduce the economic threshold for grassroots libraries; in terms of space requirements, no dedicated space is needed, only a desk is required to carry out cataloging work, reducing space requirements by more than 90%; in terms of processing scope, it can handle documents of any size and type, such as books, periodicals, ancient books, and CDs, achieving full coverage of intelligent cataloging; in terms of usage scenarios, it has strong scenario adaptability, the equipped handheld terminal is lightweight, preferably ≤500g, and has good portability, which can be used in various scenarios such as in the library, grassroots service points, and mobile document acquisition; in terms of maintenance and operation, the electronic equipment is highly integrated, without complex mechanical structures, and users can operate it after simple training, without the need for a professional maintenance team; in terms of data processing, AI recognition and data processing are both completed locally on the handheld terminal, without the need to connect to a network server, resulting in fast response speed and high quality. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the portable intelligent document cataloging system provided by the present invention;

[0021] Figure 2 This is a flowchart illustrating the overall workflow of the present invention;

[0022] Figure 3 This is a flowchart of the logic for handheld terminal and metadata extraction in this invention;

[0023] Figure 4 This is a schematic diagram of the external structure of the handheld terminal (100) of the present invention (front view, side view, top view);

[0024] (Note: The components in the attached diagram are labeled as follows: 100-Handheld terminal, 101-Core computing module, 102-Image acquisition module, 103-Data communication module, 104-Power module, 200-PC terminal, 201-Data receiving module, 202-File generation module, 203-Local storage module; The handheld terminal's external features include: finger rest groove, soft anti-slip covering layer, physical scanning button, ring light, laser rangefinder sensor, and USB Type-C interface.) Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0026] This embodiment provides a portable intelligent document cataloging system based on edge computing AI, with the following structure: Figure 1 As shown:

[0027] - Handheld terminal (100): Uses Rockchip RK3588S as the core computing module (101), equipped with a Sony IMX58648MP main camera, a ring LED fill light (3W power, 5500K color temperature), and a laser rangefinder (measurement range 5-50cm, accuracy ±1mm); the data communication module (103) is a USB 3.2 Type-C interface; the power module (104) is a 6000mAh lithium battery supporting 18W fast charging; the dimensions are 180mm×80mm×18mm, the weight is 450g, the right-side finger rest groove depth is 3mm, and the soft anti-slip coating thickness is 2mm. The handheld terminal runs a lightweight Linux operating system. The AI ​​models deployed on it include:

[0028] Page classification model: The lightweight MobileNetV3 network is used for training to identify page types such as cover, copyright page, and body text page. After training, the model undergoes pruning and quantization to convert it into an INT8 precision model, and inference is accelerated using the NPU built into the Rockchip RK3588S chip.

[0029] OCR model: It adopts a text recognition pipeline based on CNN+RNN+Attention mechanism and has been optimized for training on book page text.

[0030] All models have undergone pruning and quantization optimization to adapt to the INT8 operation of Rockchip RK3588S NPU, ensuring inference speed and accuracy.

[0031] - PC (200): Uses Intel Core i5 or higher processor, memory ≥8GB, hard disk storage space ≥500GB; the PC runs a background service program and communicates with the handheld terminal via USB CDC protocol. The data receiving module (201) supports USB 3.2 protocol; the file generation module (202) has a built-in MARC format parsing plugin, which can automatically generate .txt files that conform to the "Chinese Machine-Readable Directory Format"; the local storage module (203) has a default directory of D:\ScanOutput\, where the Metadata subdirectory stores metadata files, and the Images subdirectory creates subfolders based on unique key values ​​to store the corresponding image files.

[0032] This case study utilizes the aforementioned hardware and software, along with an optimized lightweight MobileNetV3 classification model, to achieve real-time and accurate page classification on the handheld terminal locally. This is one of the key technological foundations for achieving beneficial effects such as 'data security and efficiency', 'fast processing speed', and 'low power consumption'.

[0033] The usage steps of the system in this embodiment are as follows: Figure 2 As shown:

[0034] Step S1: The user holds the handheld terminal (100) with one hand, places the finger in the finger rest groove, and points the main camera at the document page (such as the copyright page or cover). The laser range sensor detects the distance and prompts the appropriate scanning position (the indicator light illuminates when the distance is 15-25cm).

[0035] Step S2: Press the top physical scan button to turn on the ring fill light, the main camera captures document images, and the core computing module (101) calls the local NPU to run the OCR model, identify the page type and extract metadata (such as ISBN, book title, author, publisher, etc.).

[0036] Step S3: The handheld terminal (100) generates a unique key value (such as "XYZ789"), associates the metadata with the image file, and encapsulates it into a structured data packet;

[0037] Step S4: Connect the handheld terminal (100) to the PC (200) via a USB Type-C data cable, and the data receiving module (201) will automatically receive the data packets;

[0038] Step S5: The file generation module (202) parses the data packet, generates "XYZ789.txt" (containing fields such as 010, 200, 210, etc.) in the Metadata subdirectory, and generates "XYZ789_copyright.jpg" and "XYZ789_cover.jpg" in the Images\XYZ789\ directory;

[0039] Step S6: The cataloger retrieves the above-mentioned files from the library integrated system (such as ILAS), verifies the accuracy of the metadata, and completes the data entry operation.

[0040] It should be noted that the core computing module (101) in this invention can also be other system-on-a-chip with NPU (such as Huawei HiSilicon Hi3559A), the main camera can be selected with a higher pixel module as needed, and the PC (200) can also be a laptop computer, all of which are within the scope of protection of this invention.

[0041] Figure 3 This is a flowchart and algorithm for intelligent image processing on document pages, demonstrating the complete process from image input to data association, encapsulation, and local temporary storage:

[0042] 1. Image Input (S2 Trigger): Receives document page images acquired in step S1, which can be a single page or multiple pages.

[0043] 2. Image preprocessing: With the computing power of the RK3588S NPU, image noise reduction, tilt correction, and sharpness optimization are completed, laying a solid foundation for subsequent processing.

[0044] 3. Page type determination: The page attributes are identified through the image classification model (this AI logic 1). If it is a copyright page / title page, the OCR text recognition model is triggered; if it is a cover / other page, only the image feature information is retained and OCR is not started.

[0045] 4. Key Metadata Extraction and Analysis: For pages that trigger OCR, the OCR will parse out fields such as 010ISBN, 200Title and Author, and 210Publication and Distribution after recognition. The field format will also be pre-validated to ensure compliance with MARC specifications. For pages such as covers, the retained image feature information will also be used for subsequent data association.

[0046] 5. Data association and encapsulation: Generate a unique key (e.g., ABC123) to associate two parts of content: one is the metadata extracted from the copyright page branch, and the other is the image feature information of the cover subpackage.

[0047] 6. Local temporary storage (awaiting S3 transmission): The associated data is temporarily stored in the local storage of the handheld terminal, waiting for the S3 step to encapsulate it into a structured data packet.

[0048] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A portable intelligent document cataloging system based on edge computing AI, characterized in that, It includes a handheld terminal (100) and a PC (200), wherein the handheld terminal (100) and the PC (200) are connected via a USB interface and communicate with each other. The handheld terminal (100) includes a core computing module (101), an image acquisition module (102), a data communication module (103), and a power module (104); the core computing module (101) is a system-on-a-chip with an NPU; the image acquisition module (102) includes a main camera, a ring light, and a laser rangefinder; the data communication module (103) has a USB Type-C interface; the power module (104) is a rechargeable lithium battery; the handheld terminal (100) has a finger rest groove and a soft anti-slip covering layer on the right side, and a physical scanning button on the top front end; The PC (200) includes a data receiving module (201), a file generation module (202), and a local storage module (203); the data receiving module (201) has a built-in USB driver; the file generation module (202) can parse data and generate MARC format .txt text files and .jpg image files; the local storage module (203) has preset subdirectories "Metadata" and "Images".

2. The portable intelligent document cataloging system based on edge computing AI according to claim 1, characterized in that, The system-on-a-chip of the core computing module (101) is Rockchip RK3588S, which can run OCR neural network models and image classification neural network models.

3. The portable intelligent document cataloging system based on edge computing AI according to claim 1, characterized in that, The main camera of the image acquisition module (102) is a Sony IMX58648MP camera module, and the laser range sensor has a measurement range of 5-50cm and an accuracy of ±1mm.

4. A portable intelligent document cataloging system based on edge computing AI according to claim 1, characterized in that, The power module (104) is a 6000mAh rechargeable lithium battery, which supports continuous scanning of no less than 500 pages of documents on a single full charge.

5. A portable intelligent document cataloging system based on edge computing AI according to claim 1, characterized in that, The handheld terminal (100) has dimensions of 180mm×80mm×18mm and a weight of ≤500g. The soft, non-slip covering layer is made of silicone with a Shore hardness of 50-60HA.

6. A portable intelligent document cataloging system based on edge computing AI according to claim 1, characterized in that, The MARC format .txt text file generated by the file generation module (202) contains the 010 International Standard Book Number field, the 200 Title and Author field, the 210 Publication field, and other core necessary fields. The file header has a unique key value.

7. A portable intelligent document cataloging system based on edge computing AI according to claim 1, characterized in that, The preset directory path of the local storage module (203) is a fixed path on the local hard disk of the PC. The "Images" subdirectory creates subfolders to store the corresponding image files according to unique key values.

Citation Information

Patent Citations

  • Book collecting and editing method and system and computer readable storage medium

    CN119089924A