Paper standard digital embedding system

Through the digital processing and data analysis of paper standards, the problem of lack of data mining in existing systems has been solved, the scientificity and practicality of standards have been improved, and efficient data management and query services have been provided.

CN120234460APending Publication Date: 2025-07-01CHINA AEROSPACE STANDARDIZATION INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510189152.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing paper-based digital embedded systems lack data analysis and mining functions, and cannot assist in the improvement and improvement of standards.

Method used

A paper standard digital embedded system is designed, including document digital module, data and storage module, data processing and analysis module, search and query module, user management and permission module and system integration and interface module. Through technical means such as optical character recognition, natural language processing, database management, and user permission control, digital processing and data analysis of paper standards are realized.

Benefits of technology

It improves the scientificity and practicality of standards, supports the formulation, revision and improvement of standards, reduces resource consumption, improves management efficiency and data security, and provides a convenient standard query and update mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234460A_ABST
    Figure CN120234460A_ABST
Patent Text Reader

Abstract

The invention discloses a paper standard digital embedding system, which belongs to the technical field of data identification, and comprises a document digital module, an image processing module, an optical character identification module and a format conversion module, the data and storage module comprises a database management system, a data storage framework and a storage server; the data processing and analysis module comprises a text cleaning and preprocessing module, a semantic analysis sub-module and a data annotation sub-module; a retrieval and query module; a user management and authority module; and a system integration and interface module. According to the paper standard digital embedding system, data in the digital standard can be mined and analyzed, data support is provided for formulation, revision and perfection of the standard, and scientificity and practicability of the standard are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data recognition, and particularly relates to a paper standard digital embedding system. Background Art

[0002] A paper standard digital embedding system is a tool that converts paper standard documents into digital form and embeds them into relevant business systems. It mainly includes scanning paper standard documents, using optical character recognition (OCR) technology to extract text, and converting it into editable digital text. After format processing and quality inspection, these digital contents will be embedded into business systems such as enterprise resource planning (ERP) systems and quality management systems, so as to facilitate enterprises to quickly query and reference standard terms in production, management and other links, and ensure that operations meet standard requirements. At present, most paper standard digital embedding systems directly scan documents and embed them into different systems, but do not have the function of analyzing and mining data, and cannot assist in the improvement and refinement of standards. Summary of the Invention

[0003] The purpose of the present invention is to provide a paper standard digital embedding system to solve the problem that most current paper standard digital embedding systems directly scan documents and embed them into different systems, but do not have the function of analyzing and mining data, and cannot assist in the improvement and refinement of standards as mentioned in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions: A paper standard digital embedding system includes a document digitization module, a data and storage module, a data processing and parsing module, a retrieval and query module, a user management and permission module, and a system integration and interface module. The document digitization module includes a scanning sub-module, an image preprocessing sub-module, an optical character recognition sub-module, and a format conversion sub-module; the data and storage module includes a database management system, a data storage architecture, and a storage server; the data processing and parsing module includes a text cleaning and preprocessing module, a semantic parsing sub-module, and a data annotation sub-module; the retrieval and query module includes a retrieval engine, a retrieval interface, and a result display and arrangement sub-module; the user management and permission module includes a user registration and authentication sub-module, a permission management sub-module, and a user behavior monitoring and auditing sub-module, and the system integration and interface module includes a user program interface and middleware and message passing.

[0005] In a further embodiment, the scanning sub-module is the front-end entry of the system, used to convert paper standard documents into electronic images. It needs to be equipped with a high-quality scanner that can handle papers of different sizes and materials. The resolution, color mode, and scanning speed of the scanner are key parameters;

[0006] The image preprocessing sub-module preprocesses the image, including operations such as image rotation correction, denoising, and contrast enhancement;

[0007] The optical character recognition sub-module is the core part, which converts the text information in the preprocessed image into a text format recognizable by a computer. It uses complex character recognition algorithms that analyze features such as the shape, structure, and strokes of characters. To improve the recognition accuracy, OCR software usually pre-trains character models and optimizes them for different fonts, font sizes, languages, etc.;

[0008] The format conversion sub-module: Converts the text information recognized by OCR into a suitable digital document format. This sub-module typesets the text content according to the specifications of the corresponding format based on the system requirements and user selections, including font settings, paragraph formats, chart insertions, etc.

[0009] In a further embodiment, a suitable database is selected through a database management system to store the digitized standard document data. Inside the database, a reasonable data storage architecture needs to be designed, which can be classified and stored according to standard types (such as industry standards, national standards, local standards), fields (such as mechanical manufacturing, food hygiene, environmental protection), etc. At the same time, to facilitate data retrieval and management, indexes are established, such as indexes for important information such as standard numbers and keywords, to speed up querying. In addition, a data backup and recovery mechanism needs to be considered, and a combination of regular full backups and real-time incremental backups is adopted to ensure the security and integrity of the data;

[0010] The storage server is the physical carrier for data storage, and its performance directly affects the storage and access efficiency of data. Factors such as the storage capacity, read and write speed, and reliability of the server need to be considered.

[0011] In a further embodiment, text cleaning and preprocessing are used to clean the stored text data, remove interfering information such as extra spaces, line breaks, and special symbols in the text, and at the same time, standardize the text;

[0012] The semantic parsing sub-module: Uses natural language processing (NLP) technology to perform semantic parsing on the text content of the standard document. It can identify information such as sentence structure, part of speech, entities (such as organization names, technical terms), etc., and understands the meaning of the text through technologies such as constructing syntax trees and named entity recognition;

[0013] The data annotation sub-module: Adds annotation information to the parsed text content for better understanding and retrieval. The annotation content can include the importance level of standard clauses, topic classification, citation relationships, etc.

[0014] In a further embodiment, the retrieval engine: is the core functional component of the system, responsible for processing users' retrieval requests. It can adopt a full-text retrieval engine (such as Lucene) or the retrieval function built into the database;

[0015] The retrieval interface: provides a user-friendly retrieval interface for users to conveniently input retrieval conditions. The interface design should conform to users' usage habits, including a simple and clear retrieval box and the selection of various retrieval methods (such as simple search, advanced search);

[0016] The result display and sorting sub-module: After the retrieval is completed, it displays the results to the users. The display content includes the basic information (name, number, issuing agency, etc.) of the standard documents and the fragments related to the retrieval keywords. At the same time, in order to facilitate users to quickly find the most relevant documents, the retrieval results will be sorted. The sorting method can be based on various methods such as relevance scores (obtained by calculating factors such as the occurrence frequency and position of the retrieval terms in the documents), release time (the latest released documents are ranked at the front), etc.

[0017] In a further embodiment, the user registration and authentication sub-module: is responsible for user registration and identity authentication. Users need to provide basic information (such as name, unit, contact information) for registration. After registration, the system can perform identity authentication through methods such as username / password, digital certificate, biometric recognition, etc.;

[0018] The permission management sub-module: assigns different permissions according to users' roles and responsibilities. Common user roles include system administrators, standard compilers, ordinary users, etc.;

[0019] The user behavior monitoring and auditing sub-module: monitors and audits users' behaviors in the system, and records information such as users' login times, retrieval contents, download behaviors, etc.

[0020] In a further embodiment, the application programming interface: provides an open API to enable the system to be integrated with other external systems (such as an enterprise's quality management system, production management system, etc.).

[0021] Middleware and message passing: utilizes middleware to achieve communication and data transfer between the system and external systems.

[0022] The technical effects and advantages of the present invention:

[0023] This paper-based standard digital embedding system can mine and analyze the data in the digital standards through the data processing and parsing module, such as counting the usage frequency of the standards, analyzing the correlation relationships between different standards, etc., providing data support for the formulation, revision, and improvement of the standards, and improving the scientificity and practicality of the standards;

[0024] The digitized standards can be accessed anytime and anywhere via the network, eliminating the need to rummage through large amounts of paper documents, saving time and effort. Users only need to enter keywords or specific conditions in the system to quickly locate the required standards, improving work efficiency;

[0025] When modifying, supplementing, or updating the standard content, simply perform operations in the system, and it can be updated and synchronized to all user terminals in real-time, ensuring that users obtain the latest version, avoiding work mistakes or risks caused by using old standards. Centralize the storage of a large number of paper standards in the database, saving physical space and facilitating management and maintenance. At the same time, the system can automatically back up data regularly to prevent data loss caused by natural disasters, human errors, etc., ensuring the security and integrity of the standard data;

[0026] Avoid the printing, distribution, and storage of paper standards, reduce the consumption of resources such as paper and ink, and reduce the impact on the environment, meeting the requirements of sustainable development. The digital system can achieve automated process management and permission control, reducing manual operation and management costs, and improving management efficiency and accuracy. This paper standard digitization embedding system can mine and analyze the data in the digital standards, providing data support for the formulation, revision, and improvement of the standards, and enhancing the scientificity and practicality of the standards. Brief Description of the Drawings

[0027] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a schematic diagram of the present invention. Detailed Description of the Invention

[0029] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some well-known technical features in the art are not described.

[0030] Unless otherwise defined, the up, down, left, right, front, back, inner, and outer directions involved in this article are based on the up, down, left, right, front, back, inner, and outer directions in the drawings shown in the present invention, and are hereby explained together.

[0031] Please refer to Figure 1, the present invention provides a paper standard digital embedding system, including a document digitization module, a data and storage module, a data processing and parsing module, a retrieval and query module, a user management and permission module, and a system integration and interface module. The document digitization module includes a scanning sub-module, an image preprocessing sub-module, an optical character recognition sub-module, and a format conversion sub-module; the data and storage module includes a database management system, a data storage architecture, and a storage server; the data processing and parsing module includes a text cleaning and preprocessing module, a semantic parsing sub-module, and a data annotation sub-module; the retrieval and query module includes a retrieval engine, a retrieval interface, and a result display and arrangement sub-module; the user management and permission module includes a user registration and authentication sub-module, a permission management sub-module, and a user behavior monitoring and auditing sub-module. The system integration and interface module includes a user program interface and middleware and message passing;

[0032] The scanning sub-module is the front-end entry of the system, used to convert paper standard documents into electronic images. It needs to be equipped with a high-quality scanner that can handle papers of different sizes and materials. The resolution, color mode, and scanning speed of the scanner are key parameters. For example, a high resolution (such as 1200 dpi) can ensure clear details of text and images. For standard documents with color charts, it is necessary to support the color scanning mode. At the same time, to improve efficiency, the scanning speed should be able to meet the rapid processing requirements of a large number of documents;

[0033] There may be some problems with the scanned images, such as image tilt, noise, poor contrast, etc. The image preprocessing sub-module will preprocess the images, including operations such as image rotation correction, denoising, and contrast enhancement. For example, the edge detection algorithm is used to determine the image tilt angle and perform correction, and filtering techniques are used to remove the noise generated during scanning to make the image quality meet the requirements of subsequent processing;

[0034] The optical character recognition sub-module is the core part, which converts the text information in the preprocessed image into a text format that can be recognized by a computer. It uses complex character recognition algorithms that analyze the characteristics of characters such as shape, structure, and strokes. To improve the recognition accuracy, OCR software usually pre-trains character models and optimizes them for different fonts, font sizes, languages, etc. For example, for standard documents with some ancient or special fonts, it is necessary to ensure accurate recognition by increasing the training samples or adjusting the recognition parameters. After recognition, post-processing operations such as spelling check and text correction are also performed;

[0035] The format conversion sub-module converts the text information after OCR recognition into a suitable digital document format. Common formats include PDF, DOCX, etc. The PDF format has good cross-platform compatibility and document integrity and is suitable as the final storage and publishing format for standard documents. The DOCX format is convenient for users to edit. This sub-module will typeset the text content according to the specifications of the corresponding format based on the system requirements and user choices, including font settings, paragraph formats, chart insertion, etc.

[0036] Database management system: Select a suitable database to store the digitized standard document data. For highly structured data, such as the basic information of standard documents (number, name, issuing agency, release date, etc.), chapter content, clause relationships, etc., relational databases (such as MySQL, Oracle) are good choices because they can handle the associations and queries between data well. For some unstructured data, such as document attachments containing a large number of pictures and complex formats, non-relational databases (such as MongoDB) can provide a more flexible storage method.

[0037] Data storage architecture: Inside the database, a reasonable data storage architecture needs to be designed. It can be classified and stored according to the type of standards (such as industry standards, national standards, local standards), fields (such as mechanical manufacturing, food hygiene, environmental protection), etc. At the same time, to facilitate data retrieval and management, indexes will be established, such as indexes for important information such as standard numbers and keywords to speed up querying. In addition, a data backup and recovery mechanism needs to be considered, and a combination of regular full backups and real-time incremental backups is adopted to ensure the security and integrity of the data.

[0038] Storage server: The storage server is the physical carrier of data storage, and its performance directly affects the storage and access efficiency of data. Factors such as the storage capacity, read and write speed, and reliability of the server need to be considered. To meet the storage requirements of a large amount of standard document data, disk array (RAID) technology is usually adopted. By combining multiple hard disks, the storage capacity and data read and write speed are increased, and a certain degree of data redundancy is provided to prevent data loss.

[0039] Text cleaning and preprocessing: Clean the stored text data to remove interfering information such as extra spaces, line breaks, and special symbols in the text. At the same time, standardize the text, such as unifying the character encoding and case conversion, etc. This helps to improve the accuracy and efficiency of subsequent text processing.

[0040] Semantic parsing sub-module: It uses natural language processing (NLP) technology to perform semantic parsing on the text content of standard documents. It can identify information such as sentence structure, part of speech, entities (such as organization names, technical terms), etc. By constructing syntax trees, named entity recognition and other technologies, it understands the meaning of the text. For example, it can identify key clauses, definitions, scopes, etc. in the standard and establish logical relationships between them;

[0041] Data annotation sub-module: Adds annotation information to the parsed text content for better understanding and retrieval. The annotation content can include the importance level, theme classification, citation relationship of standard clauses. For example, for a clause involving safety critical indicators, it can be annotated as "high-importance safety indicator" to facilitate users to quickly locate important information during query;

[0042] The retrieval engine is the core functional component of the system, responsible for processing user retrieval requests. It can use a full-text retrieval engine (such as Lucene) or the retrieval function built into the database. The retrieval engine will build an index for the stored standard document data. Through techniques such as inverted indexing, it associates the keywords in the document with the document itself. When the user enters a retrieval term, it can quickly locate the documents containing these keywords;

[0043] Retrieval interface: Provides a user-friendly retrieval interface to facilitate users to enter retrieval conditions. The interface design should conform to the user's usage habits, including a simple and clear retrieval box and the selection of multiple retrieval methods (such as simple search, advanced search). Simple search allows users to directly enter keywords for retrieval, while advanced search can provide more filtering conditions, such as combined retrieval according to the standard release time, scope of application, issuing agency, etc.;

[0044] Result display and sorting sub-module: After the retrieval is completed, it displays the results to the user. The display content includes the basic information of the standard document (name, number, issuing agency, etc.) and the fragments related to the retrieval keywords. At the same time, in order to facilitate users to quickly find the most relevant documents, it sorts the retrieval results. The sorting method can be based on various methods such as relevance score (obtained by calculating factors such as the occurrence frequency and position of the retrieval term in the document), release time (the latest released document is ranked first), etc.;

[0045] User registration and authentication sub-module: Responsible for user registration and identity authentication. Users need to provide basic information (such as name, unit, contact information) for registration. After registration, the system can perform identity authentication through methods such as username / password, digital certificate, biometric recognition, etc. For example, for accessing standard documents with a high security level, digital certificate authentication may be required to ensure that only authorized users can access;

[0046] Permission Management Sub-module: Different permissions are assigned according to the roles and responsibilities of users. Common user roles include system administrators, standard formulators, ordinary users, etc. System administrators can manage the entire system, including user management, data update, system maintenance, etc.; standard formulators can edit and review standard documents; ordinary users mainly query and use standard documents. Permission management is achieved through technologies such as access control lists (ACLs) to ensure that each user can only perform operations within their authorized scope;

[0047] User Behavior Monitoring and Auditing Sub-module: Monitors and audits the behavior of users in the system, records information such as the user's login time, retrieved content, download behavior, etc. This helps to detect abnormal behaviors, such as unauthorized access, frequent downloading of sensitive standard documents, etc., and can be used as the basis for system security audits;

[0048] Application Programming Interface (API): Provides an open API to enable the system to be integrated with other external systems (such as an enterprise's quality management system, production management system, etc.). The API defines the interaction methods between the system and external systems, including data transmission formats, call methods, etc. Through the API, external systems can easily obtain standard document data and implement the application of standards in other business processes;

[0049] Middleware and Message Passing: Utilizes middleware to achieve communication and data transfer between the system and external systems. For example, message queue middleware can ensure the reliable transmission of data between different systems. When a standard document is updated, relevant external systems can be notified in a timely manner through message passing to ensure the collaborative work between systems.

[0050] All standard parts used in the present invention can be purchased from the market. Special-shaped parts can be customized according to the descriptions in the specification and the attached drawings. The specific connection methods of each part all adopt conventional means such as bolts, rivets, and welding that are mature in the prior art. Machinery, parts, and equipment all adopt conventional models in the prior art. In addition, the circuit connection adopts conventional connection methods in the prior art, which will not be elaborated here. The control method of the present invention is controlled by a controller, and the control circuit of the controller can be realized by simple programming by those skilled in the art. The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0051] In the description of the present invention, it should be understood that the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0052] Working principle,

[0053] The paper-based standard digital embedding system scans the paper-based standards through the document digitization module, converts the format of the scanned content, stores the converted paper-based standards in the database, embeds different standards into different systems according to different usage requirements, and enables users to retrieve the standards through the retrieval engine during actual use. At the same time, the data processing and analysis module mines and analyzes the data in the digital standards, such as counting the usage frequency of the standards, analyzing the correlation between different standards, etc., providing data support for the formulation, revision, and improvement of the standards, and improving the scientificity and practicality of the standards.

[0054] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A paper standard digital embedding system, comprising a document digitization module, a data and storage module, a data processing and analysis module, a retrieval and query module, a user management and authority module and a system integration and interface module, characterized in that: The document digitization module includes a scanning submodule, an image preprocessing submodule, an optical character recognition submodule and a format conversion submodule; the data and storage module includes a database management system, a data storage architecture and a storage server; the data processing and parsing module includes a text cleaning and preprocessing module, a semantic parsing submodule and a data annotation submodule; The retrieval and query module includes a retrieval engine, a retrieval interface, and a result display and arrangement submodule; The user management and authority module includes a user registration and authentication submodule, an authority management submodule and a user behavior monitoring and auditing submodule. The system integration and interface module includes a user program interface and middleware and message transmission.

2. A paper standard digital embedding system according to claim 1, characterized in that: The scanning submodule is the front-end entrance of the system, which is used to convert paper standard documents into electronic images. It needs to be equipped with a high-quality scanner that can handle papers of different sizes and materials. The resolution, color mode and scanning speed of the scanner are key parameters; The image preprocessing submodule will preprocess the image, including image rotation correction, denoising, contrast enhancement and other operations; The optical character recognition submodule is the core part, which converts the text information in the pre-processed image into a text format that can be recognized by the computer. It uses complex character recognition algorithms that analyze the shape, structure, strokes and other features of the characters. In order to improve the recognition accuracy, OCR software usually pre-trains the character model and optimizes it for different fonts, font sizes, languages, etc. Format conversion submodule: converts the text information recognized by OCR into a suitable digital document format. This submodule will layout the text content according to the corresponding format specifications based on the system requirements and user choices, including font settings, paragraph formats, chart insertion and other operations.

3. A paper standard digital embedding system according to claim 1, characterized in that: Select a suitable database through the database management system to store the digitized standard document data. Within the database, a reasonable data storage architecture needs to be designed. Classified storage can be performed according to the type of standard (such as industry standards, national standards, local standards), field (such as machinery manufacturing, food hygiene, environmental protection), etc. At the same time, in order to facilitate data retrieval and management, indexes will be established, such as indexing important information such as standard numbers and keywords to speed up the query speed. In addition, the data backup and recovery mechanism needs to be considered, and a combination of regular full backup and real-time incremental backup is used to ensure data security and integrity; The storage server is the physical carrier of data storage. Its performance directly affects the storage and access efficiency of data. It is necessary to consider factors such as the server's storage capacity, read and write speed, and reliability.

4. A paper standard digital embedding system according to claim 1, characterized in that: Text cleaning and preprocessing are used to clean the stored text data, remove unnecessary spaces, line breaks, special symbols and other interfering information in the text, and standardize the text; Semantic parsing submodule: It uses natural language processing (NLP) technology to perform semantic parsing on the text content of standard documents. It can identify sentence structure, part of speech, entities (such as organization names, technical terms) and other information, and understand the meaning of the text by building syntax trees and named entity recognition. Data annotation submodule: Add annotation information to the parsed text content for better understanding and retrieval. The annotation content can include the importance level of standard clauses, subject classification, citation relationship, etc.

5. A paper standard digital embedding system according to claim 1, characterized in that: Search engine: It is the core functional component of the system and is responsible for processing user search requests. It can use a full-text search engine (such as Lucene) or the search function of the database itself. Search interface: Provide users with a friendly search interface to facilitate users to enter search conditions. The interface design should be in line with user habits, including a simple and clear search box and multiple search methods (such as simple search and advanced search); Result display and sorting submodule: After the search is completed, the results will be displayed to the user. The displayed content includes the basic information of the standard document (name, number, publishing agency, etc.) and fragments related to the search keywords. At the same time, in order to facilitate users to quickly find the most relevant documents, the search results will be sorted. The sorting method can be based on relevance score (obtained by calculating the frequency and position of the search terms in the document), release time (the latest released documents are ranked first), and other methods.

6. A paper standard digital embedding system according to claim 1, characterized in that: User registration and authentication submodule: responsible for user registration and identity authentication. Users need to provide basic information (such as name, unit, contact information) to register. After registration, the system can authenticate the user through user name / password, digital certificate, biometrics, etc. Permission management submodule: different permissions are assigned according to user roles and responsibilities. Common user roles include system administrator, standard compiler, ordinary user, etc.; User behavior monitoring and auditing submodule: monitors and audits user behavior in the system, and records user login time, search content, download behavior and other information.

7. A paper standard digital embedding system according to claim 1, characterized in that: Application Programming Interface: Provides an open API to enable the system to be integrated with other external systems (such as the company's quality management system, production management system, etc.). Middleware and messaging: Use middleware to implement communication and data transfer between the system and external systems.