Metadata-based document quick labeling retrieval method and system
Through the metadata extraction, annotation and retrieval modules, the problems of slow information processing and low accuracy in document management methods in the existing technology are solved, fast and accurate document retrieval is achieved, the system development and maintenance costs are reduced, and the user experience is improved.
Patent Information
- Application Number
- CN202411991088.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing automated document management methods have shortcomings in information processing speed and data processing accuracy, especially poor adaptability to specific fields, and ignore the importance of metadata.
Through the metadata extraction module, annotation module and retrieval module, metadata is directly used for document annotation and retrieval. By extracting metadata, extracting structured information and building a query index based on the metadata retrieval module, a metadata-based document annotation and retrieval system is realized, and a metadata retrieval system is realized, which realizes the rapid annotation and retrieval of documents based on metadata.
It improves retrieval efficiency, enhances retrieval accuracy, lowers technical barriers, and improves user experience.
Smart Images

Figure CN119719344B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, and in particular to a metadata-based document rapid labeling and retrieval method and system. BACKGROUND
[0002] With the rapid development of information technology, the number of documents generated by enterprises, scientific research institutions and individuals in their work and life has shown an explosive growth. In the face of the explosive growth of document data, the traditional manual labeling and retrieval method has been difficult to meet the demand of quickly and accurately obtaining the required data. At present, although some existing automatic document management methods can realize the automatic classification and retrieval of documents to a certain extent, these methods often need to rely on complex natural language processing algorithms. Once the algorithm is too complex, the data processing speed will become slower, and the adaptability to specific fields will also become poor. In addition, in order to improve the speed as much as possible, the existing methods ignore a large amount of metadata (such as author information, creation time, keywords, topic classification, etc.) contained in the document when performing document retrieval. These metadata play an important role in improving the accuracy of retrieval. It can be seen that the existing automatic document retrieval has the problems of slow information processing speed and low data processing accuracy. SUMMARY
[0003] The present application provides a metadata-based document rapid labeling and retrieval method and system to solve the problem of slow information processing speed and low data processing accuracy in the existing automatic document retrieval.
[0004] In order to achieve the above purpose, the present application realizes the technical scheme as follows:
[0005] In the first aspect, the present application provides a metadata-based document rapid labeling and retrieval method applied to a metadata-based document rapid labeling and retrieval system, the system comprising a metadata extraction module, a metadata labeling module and a metadata retrieval module, and the method comprising:
[0006] The metadata extraction module creates a metadata extraction task, a task execution strategy and an import strategy, and extracts metadata in the target document based on the task execution strategy, the metadata extraction task and the import strategy to obtain a first metadata document;
[0007] The metadata labeling module extracts structured information from the first metadata document to obtain a second metadata document, and labels the second metadata document based on the association between the second metadata document and the target document to obtain a labeled document;
[0008] The metadata retrieval module is configured to construct a query index based on metadata and the labeled document, quickly locate the query index based on a query condition input by a user, and query information in the target document based on the query index.
[0009] Optionally, the metadata extraction task includes a type of metadata.
[0010] The task execution strategy includes a task execution time and a task cycle time.
[0011] The storage strategy includes full storage and incremental storage.
[0012] Optionally, the structured information extraction on the first metadata document includes:
[0013] The first metadata document is preprocessed, segmented, and element extracted to obtain a second metadata document, wherein the preprocessing includes removing useless formats, spaces, and special characters in the first metadata document through data cleaning, and standardizing document formats in the first metadata document.
[0014] The document segmentation includes segmenting document data in the first metadata document into small document units, which are paragraphs or sentences.
[0015] The document element extraction includes extracting specific information elements in the first metadata document through regular expressions, natural language processing tools, deep learning, and template matching to construct a second metadata document.
[0016] Optionally, the regular expression is a string matching a specific pattern.
[0017] The natural language processing tool is a part-of-speech tagging and named entity recognition using an NLP library to identify and extract specific entities in the document.
[0018] The deep learning is sequence labeling using a deep learning model to identify and extract specific elements in the document.
[0019] The template matching is matching a specific template with the document to extract matching formats or structures in the document.
[0020] Optionally, the labeling of the second metadata document based on the association between the second metadata document and the target document includes:
[0021] The existing metadata is read from the second metadata document, and the second metadata document and the read existing metadata are associated in an automatic or manual manner.
[0022] The associated document metadata is stored into a database to generate a corresponding labeled document.
[0023] In a second aspect, the embodiments of the present application provide a metadata-based document rapid labeling and retrieval system, which comprises a metadata extraction module, a metadata labeling module and a metadata retrieval module.
[0024] The metadata extraction module is configured to create a metadata extraction task, a task execution strategy and an import strategy, and is further configured to extract metadata in a target document based on the task execution strategy, the metadata extraction task and the import strategy to obtain a first metadata document.
[0025] The metadata labeling module is configured to extract structured information from the first metadata document to obtain a second metadata document, and is further configured to label the second metadata document based on the association between the second metadata document and the target document to obtain a labeled document.
[0026] The metadata retrieval module is configured to construct a query index based on the association between metadata and the labeled document, quickly locate the query index through a user input query condition, and query information in the target document based on the query index.
[0027] Optionally, the metadata retrieval module comprises an index construction module, a query processing module, a retrieval algorithm module and a user interaction module.
[0028] The index construction module is configured to construct the query index.
[0029] The query processing module is configured to parse user input information to optimize query results.
[0030] The retrieval algorithm module is configured to match a retrieval algorithm according to user input information.
[0031] The user interaction module is configured to obtain user input information, and is further configured to display retrieval results.
[0032] Advantages:
[0033] The metadata-based document rapid labeling and retrieval method provided by the present application improves retrieval efficiency: by directly using metadata for retrieval, the complex natural language processing process is avoided, and the retrieval speed is significantly improved; enhances retrieval accuracy: the accuracy and consistency of metadata ensure the accuracy of retrieval results, reducing the cases of false positives and false negatives; reduces technical threshold: reduces the dependence on complex natural language processing technology, reduces the cost and difficulty of system development and maintenance; improves user experience: a friendly user interface and flexible query method enable users to more conveniently manage and retrieve documents. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 Flow chart of metadata-based document quick labeling retrieval method for preferred embodiment of the present application;
[0035] Figure 2 Flow chart of metadata extraction implementation method for preferred embodiment of the present application;
[0036] Figure 3 Flow chart of metadata labeling implementation method for preferred embodiment of the present application;
[0037] Figure 4 Flow chart of metadata retrieval implementation method for preferred embodiment of the present application. DETAILED DESCRIPTION
[0038] The technical solutions of the present application will be described in detail below. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present application.
[0039] Unless otherwise defined, the technical terms or scientific terms used in the present application shall be understood as the usual meanings understood by those skilled in the art to which the present application belongs. The terms "first", "second" and similar terms used in the present application do not represent any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms "one" or "a" and similar terms do not represent a quantity limitation, but represent the existence of at least one. The terms "connected" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right" and the like are only used to represent relative positional relationships, and when the absolute positions of the described objects are changed, the relative positional relationships are also changed accordingly.
[0040] Please refer to Figure 1 The present application embodiment provides a metadata-based document quick labeling retrieval method, which is applied to a metadata-based document quick labeling retrieval system. The system comprises a metadata extraction module, a metadata labeling module and a metadata retrieval module. The method comprises:
[0041] The metadata extraction module creates metadata extraction tasks, task execution strategies and storage strategies, and extracts metadata in target documents based on the task execution strategies, the metadata extraction tasks and the storage strategies to obtain a first metadata document;
[0042] The metadata annotation module is used to perform structured information extraction on the first metadata document to obtain a second metadata document, and to perform document annotation based on the association between the second metadata document and the target document to obtain an annotated document.
[0043] The metadata retrieval module is used to construct a query index based on the association between the metadata and the annotated document, to quickly locate the query index based on a query condition input by a user, and to query information in the target document based on the query index.
[0044] Optionally, the metadata extraction task includes a type of metadata.
[0045] The task execution strategy includes a task execution time and a task cycle time.
[0046] The storage strategy includes full storage and incremental storage.
[0047] In the above embodiment, as shown in the figure, Figure 2 The metadata extraction module is used to establish an extraction task for a database under the information base, to configure a scheduling plan, to automatically extract structured metadata, and to obtain technical metadata, business metadata, and management metadata.
[0048] 1) New extraction task
[0049] The metadata extraction provides a metadata extraction task addition function, and the function can be used to visually construct a metadata extraction task.
[0050] 2) Set basic information
[0051] The function can be used to set a task name, an execution strategy, and a storage strategy. The execution strategy is mainly an execution time strategy, for example, daily execution at midnight, Monday, Wednesday, and Friday at 18:00, etc. The storage strategy mainly includes full storage and incremental storage.
[0052] 3) Select data source
[0053] The data source of the metadata extraction is set, and the data is derived from a data collection-> data source management-> connected data source.
[0054] 4) Start extraction task
[0055] After the extraction task is started, the extraction task will be started according to the set execution strategy rule. For example, the set support strategy is daily execution at midnight, the storage strategy is full storage, and when the start button is clicked, the task is executed at midnight the next day.
[0056] 5) Execute extraction task
[0057] After the extraction task is executed, the extraction task will be started immediately according to the set execution strategy rule. For example, the support strategy is set to be executed at midnight every day, the storage strategy is full volume, and when the execution button is clicked, the extraction task is executed immediately in the full volume extraction mode.
[0058] Optionally, the structured information extraction on the first metadata document obtains a second metadata document, and the structured information extraction on the first metadata document comprises:
[0059] The second metadata document is obtained after document preprocessing, document segmentation and document element extraction on the first metadata document, wherein the document preprocessing comprises removing useless formats, spaces and special characters in the first metadata document through data cleaning, and standardizing document formats in the first metadata document;
[0060] The document segmentation comprises segmenting document data in the first metadata document into small document units, and the small document unit is a paragraph or a sentence;
[0061] The document element extraction comprises extracting specific information elements in the first metadata document through regular expressions, natural language processing tools, deep learning and template matching to construct the second metadata document.
[0062] Optionally, the regular expression is a string matching a specific pattern.
[0063] The natural language processing tool is a part-of-speech tagging and named entity recognition using an NLP library to identify and extract specific entities in the document.
[0064] The deep learning is sequence labeling using a deep learning model to identify and extract specific elements in the document.
[0065] The template matching is matching a specific template with the document to extract the matched format or structure in the document.
[0066] Optionally, the document annotation based on the association between the second metadata document and the target document obtains an annotated document, and the document annotation based on the association between the second metadata document and the target document comprises:
[0067] The existing metadata is read from the second metadata document, and the second metadata document and the read existing metadata are associated in an automatic or manual manner.
[0068] The associated document metadata is stored in a database to generate a corresponding annotated document.
[0069] In the above embodiments, as Figure 3As shown, through structured information extraction, the title, author, creation date, and keywords of the qualitative document can be obtained. Then, using metadata-based document labeling, technical metadata, salesperson data, and management metadata of the standardized document are generated, providing basic data for metadata retrieval.
[0070] 1) Structured Information Extraction
[0071] Structured information extraction is the process of converting unstructured document data into structured data, including document preprocessing, document segmentation, and document element extraction.
[0072] Document preprocessing: Through data cleaning, remove unnecessary formats, spaces, special characters, etc.; through standardization, unify document formats, such as date, currency, etc.
[0073] Document segmentation: divide the document into paragraphs, sentences or smaller units for further processing.
[0074] Document element extraction: the process of extracting specific information elements from the document, such as text, numbers, dates, names, places, etc. Methods include regular expressions, natural language processing tools, deep learning, and template matching.
[0075] a) Regular expressions, use regular expressions to match strings of specific patterns, such as phone numbers, email addresses, date formats, etc.
[0076] b) Natural language processing tools, use NLP libraries for part-of-speech tagging, named entity recognition, etc. to identify and extract specific entities in the text.
[0077] c) Deep learning, use deep learning models for sequence labeling tasks to identify and extract specific elements in the document.
[0078] d) Template matching, use template matching methods to extract fixed formats or structures contained in the document.
[0079] 2) Metadata-based Document Labeling
[0080] Metadata-based document labeling is a process of attaching structured information to document content to describe the key features and content of the document. Its implementation steps include reading metadata, metadata association, manual review, and metadata storage.
[0081] a) Read metadata, read existing technical metadata, business metadata, and management metadata from the information library;
[0082] b) Metadata association, using automatic or manual way to realize the association between document and metadata, manual way provides visual operation interface, and automatic way uses semantic analysis method to realize the automatic association between document and metadata.
[0083] c) Artificial review, providing visual operation, which can view and review the metadata associated by automatic or manual way, and can modify the metadata associated with the document.
[0084] d) Metadata storage, storing the metadata of the document associated and reviewed into database for subsequent retrieval and management.
[0085] The embodiment of the application further provides a metadata-based document rapid labeling retrieval system, which comprises a metadata extraction module, a metadata labeling module and a metadata retrieval module.
[0086] The metadata extraction module is used for creating a metadata extraction task, a task execution strategy and a storage strategy, and is further used for extracting metadata in a target document based on the task execution strategy, the metadata extraction task and the storage strategy to obtain a first metadata document.
[0087] The metadata labeling module is used for extracting structured information from the first metadata document to obtain a second metadata document, and is further used for labeling the second metadata document based on the association between the second metadata document and the target document to obtain a labeled document.
[0088] The metadata retrieval module is used for constructing a query index based on metadata and the labeled document, quickly positioning the query index through a user input query condition, and querying information in the target document based on the query index.
[0089] Optionally, the metadata retrieval module comprises an index construction module, a query processing module, a retrieval algorithm module and a user interaction module.
[0090] The index construction module is used for constructing the query index.
[0091] The query processing module is used for analyzing user input information to optimize query results.
[0092] The retrieval algorithm module is used for matching retrieval algorithms according to user input information.
[0093] The user interaction module is used for obtaining user input information, and is further used for displaying retrieval results.
[0094] In the above embodiment, as Figure 4As shown, metadata retrieval mainly provides index construction, query service, builds various indexes based on document annotation results, and provides fast query service based on user input and query conditions.
[0095] a) Index construction module, inverted index construction - create metadata to document mapping for fast retrieval; forward index construction - maintain the original content of the document and its attributes for retrieval result display; index update strategy - handle new or modified documents to maintain the timeliness of the index.
[0096] b) Query processing module, query analysis - understand user query intent, including query expansion, spelling correction; query optimization - improve the query to get more accurate or more relevant results.
[0097] c) Retrieval algorithm module, matching algorithm - based on keyword, phrase or sentence matching; similarity calculation - cosine similarity, Jaccard similarity are used to evaluate the relevance of documents and queries; sorting algorithm - sort the retrieval results according to relevance or other standards.
[0098] d) User interaction module, user interface - provide search box, advanced search options; result display - display retrieval results in the form of list, summary, etc.; user feedback - collect user feedback on retrieval results for system improvement.
[0099] The above describes the preferred embodiments of the present application in detail. It should be understood that those skilled in the art can make many modifications and changes without creative labor according to the concept of the present application. Therefore, any technical solution obtained by logical analysis, reasoning or limited experiment based on the existing technology according to the concept of the present application shall be within the protection scope determined by the claims.
Claims
1. A metadata-based document rapid annotation and retrieval method, applied to a metadata-based document rapid annotation and retrieval system, characterized in that: The system includes: a metadata extraction module, a metadata annotation module and a metadata retrieval module, and the method includes: Creating a metadata extraction task, a task execution strategy, and a storage strategy through a metadata extraction module, and extracting metadata from a target document based on the task execution strategy, the metadata extraction task, and the storage strategy to obtain a first metadata document; Extracting structured information from the first metadata document using a metadata annotation module to obtain a second metadata document, and annotating the second metadata document based on an association between the second metadata document and a target document to obtain an annotated document; Using the metadata retrieval module to build a query index based on the metadata and the annotated document, quickly locate the query index through the query conditions input by the user, and query information in the target document based on the query index; The metadata extraction task includes: the type of metadata; The task execution strategy includes: task execution time and task cycle time; The warehousing strategy includes full warehousing and incremental warehousing; The step of annotating the second metadata document based on the association between the second metadata document and the target document to obtain an annotated document includes: Reading existing metadata from the second metadata document, and associating the second metadata document with the read existing metadata in an automatic or manual manner; The associated document metadata is stored in the database to generate the corresponding annotated document.
2. The method for rapid document annotation and retrieval based on metadata according to claim 1, characterized in that: Extracting structured information from the first metadata document to obtain a second metadata document includes: Performing document preprocessing, document segmentation, and document element extraction on the first metadata document to obtain a second metadata document, wherein the document preprocessing includes: removing useless formats, spaces, and special characters in the first metadata document through data cleaning, and then standardizing the document format within the same first metadata document; Document segmentation includes: segmenting the document data in the first metadata document into small document units, the small document units being paragraphs or sentences; Document element extraction includes: extracting specific information elements from the first metadata document through regular expressions, natural language processing tools, deep learning and template matching to construct a second metadata document.
3. The method for rapid document annotation and retrieval based on metadata according to claim 2, characterized in that: The regular expression is: a string that matches a specific pattern; The natural language processing tool is: using the NLP library to perform part-of-speech tagging and named entity recognition to identify and extract specific entities in the document; The deep learning method is: using a deep learning model to perform sequence labeling, identify and extract specific elements in documents; The template matching is to use a specific template to match the document and extract the matching format or structure in the document.
4. A metadata-based document rapid annotation and retrieval system, used to execute the metadata-based document rapid annotation and retrieval method according to any one of claims 1 to 3, characterized in that: The system includes: a metadata extraction module, a metadata annotation module and a metadata retrieval module: A metadata extraction module, configured to create a metadata extraction task, a task execution strategy, and a storage strategy, and further configured to extract metadata from a target document based on the task execution strategy, the metadata extraction task, and the storage strategy to obtain a first metadata document; a metadata annotation module, configured to extract structured information from the first metadata document to obtain a second metadata document, and further configured to annotate the second metadata document based on an association between the second metadata document and a target document to obtain an annotated document; The metadata retrieval module is used to construct a query index based on the metadata and the annotated document, quickly locate the query index through the query conditions input by the user, and query information in the target document based on the query index.
5. The metadata-based document rapid annotation and retrieval system according to claim 4, characterized in that: The metadata retrieval module includes: an index construction module, a query processing module, a retrieval algorithm module and a user interaction module; The index building module is used to build a query index; The query processing module is used to parse user input information and optimize query results; The retrieval algorithm module is used to match the retrieval algorithm according to the user input information; The user interaction module is used to obtain user input information and also to display search results.
Citation Information
Patent Citations
Electronic document screening query method and system
CN116662521A
Generating labels for images associated with a user
US20170185670A1