Basic search platform-oriented special file enhanced retrieval method and device
By enhancing the retrieval method for special files on the basic search platform, classifying and identifying PDF and WORD files, performing enhanced data processing and cleaning, and writing them into the index library, the problem of special files being unable to be retrieved is solved, and search accuracy and system stability are improved.
Patent Information
- Application Number
- CN202510702377.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-23
AI Technical Summary
The existing basic search platform is unable to effectively process and retrieve special electronic documents such as scanned and encrypted documents, resulting in users being unable to accurately locate these files, affecting the practicality of the search platform and user experience.
By classifying special files into PDF and WORD files, developing positioning strategies to identify and associate unique identifiers, obtaining basic information, performing enhanced data processing and cleaning, generating readable text data, and writing it into the index library to improve retrieval accuracy.
It achieves accurate identification and retrieval of special files, maintains file format consistency, reduces implementation costs, ensures system stability, and improves user experience.
Smart Images

Figure CN120687588A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of special file retrieval, and in particular relates to a special file enhanced retrieval method and device oriented to a basic search platform. Background Art
[0002] With the rapid development of enterprise digital transformation and the next generation of digital technologies, major companies are recording and converting all kinds of raw information from their business processes into data through business digitization. They are also integrating data and information technology into existing business processes as new production factors. As data increases, companies are building basic search systems to retrieve and locate massive amounts of data. Electronic document retrieval and acquisition primarily involves locating target files through querying file names, full-text searches, or knowledge graphs generated from layout files. Basic search platforms do not subdivide and label documents within heterogeneous, multi-source raw data. Document parsing generally defaults to unified parsing and text extraction, such as using open-source LibreOffice technology.
[0003] During the digital transformation of enterprises, a large number of business documents exist in special formats, such as scanned and encrypted documents (read-only, no editing allowed), or with primary content inserted as images. These special document types are difficult to effectively process on basic search platforms. Even mainstream file parsing tools like LibreOffice are unable to accurately extract the text content. This results in users being unable to locate these special documents through search, severely impacting the practicality and user experience of the search platform.
[0004] The enterprise basic search system currently solves the problem of special file retrieval mainly in the following two ways.
[0005] 1. System architecture reconstruction: Redesign and develop the system architecture to support special file processing. This approach requires a lot of development resources and takes a long time. 2. Component Replacement and Upgrade: Replace or upgrade the file parsing components of the basic search platform. Different file types often require customized parsing components. This method requires decoupling the file parsing modules of the existing architecture of the basic search platform, and the development approach is limited by upstream and downstream related system modules.
[0006] These solutions require the technical team to have a deep understanding of the system architecture, and implementation can impact the normal operation of existing business processes. Furthermore, because enterprise search platforms often involve data interaction across multiple business systems, any architectural changes can have ripple effects, making project risks and costs difficult to control. Therefore, a specialized file-enhanced retrieval method and device for basic search platforms is needed to address these issues. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a special file enhanced retrieval method and device for a basic search platform, aiming to solve the problem in the prior art that the basic search platform based on open source file parsing tools cannot adapt to the limited content search of special electronic documents.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is: A special file enhanced retrieval method for a basic search platform, comprising: S1, screening special files that need to be search enhanced based on historical information data of the basic search platform, and classifying the special files into PDF special files or WORD special files based on technical characteristics; S2, based on the special file types defined in step S1, formulate a special file location strategy adapted to the basic search platform, and identify special files that meet the definition in step S1 in the multi-source heterogeneous data of the basic search platform; S3, associating the unique identifier of the special file identified in step S2 with the relational database of the basic search platform, obtaining and solidifying the basic information of the special file, and realizing persistence of the basic information of the special file; S4: For PDF special files and WORD special files, based on the special file basic information persisted in step S3, enhanced data processing is performed according to a preset period to generate temporary text data corresponding to the special files; S5, performing data cleaning and processing on the temporary text data formed in step S4 to form pure text enhanced content to be written, so as to ensure the readability of the text content output by the text recognition model and that the text content data can be normally written into the index library; S6, based on the special file basic information persisted in step S3, writing the plain text enhanced content data generated in step S5 into the corresponding index in the basic search platform index library to enhance the indexing capability of the target special file; S7, after completing the indexing capability enhancement in step S6, a search request is made to the index library; the search request is retrieved based on the metadata information and optimized text content of the special file to improve the accuracy and search scope of the search results.
[0009] Preferably, in step S1, the special file is classified as a PDF special file or a WORD special file definition including: Special PDF files include PDF scans, encrypted PDFs, and PDFs with an added OCR layer. A PDF scan is a document that has been scanned into an image format and then saved as a PDF file, from which text cannot be directly extracted or edited. An encrypted PDF is a PDF file that has been set with a blank user password to restrict others from printing, modifying, or annotating. A WORD special file is a WORD document that contains images.
[0010] Preferably, in step S2, the special file location strategy includes: For special PDF files, identification is performed by processing the exported data in the index library: If the field used to record and store the parsed content of the PDF file in the exported index document is empty, it can be determined as a special file, and the file unique identifier is obtained. The file unique identifier is used to associate with the file metadata table to obtain the storage location of the corresponding file; For special WORD files, identify them based on whether they contain pictures: If an image is detected, the document is treated as a special file and the file's unique identifier is used to associate it with the file metadata table to obtain the storage location of the corresponding file.
[0011] Preferably, in step S3, persisting basic information of special files includes: Based on the positioning result of step S2, the unique identifier of the special file is associated with the relational database of the basic search platform to obtain the detailed data of the PDF special file and the WORD special file, and solidified into a fixed special file information table in the structured database; the information table contains at least the following 6 fields: file global unique ID (uniquely identifies each file), file type (identifies the file type, such as PDF or WORD), file storage path (records the storage location of the file in the unstructured data storage system), file processing status (indicates the current document processing progress), creation time (records the time when the data is first entered into the database) and modification time (records the time of the last update).
[0012] Preferably, in step S4, for a PDF special file type, special file enhanced data processing includes: Based on the file storage location field in the persistent basic information in step S3, the physical electronic file to be processed is obtained from the corresponding storage address in the unstructured data storage system; PDF special files: Decrypt the physical electronic file using a PDF tool to generate a temporary PDF file to be processed in the next step; Extract the image objects contained in the temporary PDF file, optimize the image objects, use the text recognition model to perform text recognition on the optimized image data, the recognition results include the text content of the corresponding image data, and integrate the text recognition content of all image objects in the entire document to form temporary text data.
[0013] WORD special files: The image objects contained in the physical electronic file are extracted by an image extraction tool, the image objects are optimized, the optimized image data are subjected to text recognition using a text recognition model, and the recognition contents of all image objects in the entire document are integrated to form temporary text data.
[0014] Furthermore, the key steps of step S4 include: image preprocessing (such as denoising, binarization, tilt correction, etc.), text area detection (such as locating the text area through edge detection and contour analysis), character segmentation (splitting and extracting character features based on character spacing and connectivity), character recognition (using support vector machines or convolutional neural networks to classify and identify character features), and finally restoring the recognized characters into editable coherent text through text integration.
[0015] Preferably, in step S5, performing data cleaning on the temporary text data formed in step S4 includes: The temporary text data generated in step S4 is cleaned, converted and processed, including: special character replacement, adding escape characters, filtering redundant spaces and line breaks, character set conversion and text base64 conversion, so that the text content meets the requirements for being correctly written into the index library.
[0016] Preferably, in step S6, writing the enhanced content into the basic search platform index library includes: Setting a target index, a globally unique identifier of the target special file, plain text enhanced content of the target special file, and a target location; The plain text enhanced content of the target special file is written into the index library of the basic search platform, and operations are performed based on the amount of data: For a small amount of data whose data volume is less than a set threshold, the plain text enhanced content of the target special file is written into the target location of the index library based on the update instruction using the data manipulation language DML; For batch large-scale data with a data volume equal to a set threshold, the plain text enhanced contents of several target special files are written into the target location of the index library through the data manipulation language DML based on the index library batch API.
[0017] Preferably, in step S7, performing a search request on the index library includes: The search query body is generated by performing word segmentation, semantic analysis, and feature word processing (proper noun identification, synonym replacement, error correction word replacement, and sensitive word blocking) on the search content entered by the user. This is combined with the user's desired search object classification scope (such as business category, document type, and time range), recall priority (such as by time or by relevance), and word segmentation method (such as including all keywords, including any keywords, including complete keywords, and not including keywords). In the basic search platform index library after step S6 is completed, the search recall result record is determined through similarity calculation based on the preset threshold, the attribute information and storage information of the data object in the record are extracted, the recall result is generated, and it is returned to the basic search platform front end to be displayed to the search user.
[0018] Preferably, a basic search platform-oriented special file enhanced retrieval device is used to execute the basic search platform-oriented special file enhanced retrieval method; the device includes the following modules: The file recognition module is used to collect, process and identify special files; the file recognition module specifically includes: The document metadata acquisition unit is used to extract key data information of documents from the index library and relational database data storage components of the basic search platform to support subsequent document processing and recognition processes; The file metadata processing unit is used to process the extracted document metadata information, screen out target files using a special file recognition algorithm, and generate structured information in a unified format. The information is then solidified in a relational database, providing a high-quality data foundation for subsequent processing. The data management module is used to record the information of the special file processing process and support the management of the data processing process. The data management module specifically includes: Process data persistence unit, used to record data processing transaction information and its processing status, to ensure the traceability and reliability of the data processing process; Data visualization unit, used to visualize recorded information, facilitate management and monitoring of data processing, and improve the transparency and efficiency of data processing; The file processing module is used to recognize text in special files, generate temporary text data, and optimize the temporary text data to form enhanced text data to be written. The file processing module specifically includes: Special file decryption unit, used to perform unified decryption operations and remove file protection restrictions; An image processing unit, used for extracting image objects contained in the file and performing optimization processing on the image objects; OCR recognition unit, used for OCR recognition of images, extracting original text information, and providing a more accurate data basis for subsequent text processing; The cleaning and conversion unit is used to uniformly process the extracted original text information, including encoding conversion and special symbol replacement, to improve the quality and consistency of text data; The text writing module is used to write the enhanced text data generated by the file processing module into the search engine of the search platform and monitor the writing process. The text writing module specifically includes: The data writing unit is used to write the enhanced text data into the basic search platform index library in a predetermined format. It supports both batch writing and single writing modes to ensure the efficiency and accuracy of data writing. The write monitoring unit is used to monitor the data writing status in real time, record indicators such as the write success rate and write speed, and issue early warnings when anomalies are detected to ensure the stability and reliability of the data writing process; The file query module has the following functions: Perform word segmentation, semantic analysis, and feature word processing based on the search content entered by the user; Generate a search query request body based on the user's desired search object classification scope, recall priority, and word segmentation method; Based on the updated index library data, the search recall result records are determined by similarity calculation and based on the preset threshold; Extract the attribute information and storage information of the data objects in the record, generate the recall results, and return them to the front end of the basic search platform to be displayed to the search user.
[0019] The beneficial effects of the present invention are: 1. This patent enables content search for special electronic documents on an enterprise's basic search platform. By enhancing the basic search platform's technology, it effectively solves the problem of ineffective content retrieval for special types of files, such as scanned and encrypted documents. It also improves the search platform's processing capabilities for special files, making them searchable while preserving the source file format and style. The files users access on the basic search platform front-end remain consistent with the source files, ensuring that the files they obtain are authentic and reliable.
[0020] 2. This patent utilizes a lightweight enhancement solution, providing a unified standard and framework for special file processing. This solution offers excellent scalability and allows for flexible deployment of solutions based on different special file types and formats. This allows for enhanced retrieval of special documents within the basic search platform at a low implementation cost and workload. The enhancement process does not impact the existing business processes of the basic search platform and is imperceptible to platform users, ensuring system stability.
[0021] 3. Without changing the existing architecture of the underlying search platform, this patent achieves precise identification and periodic enhanced processing of special files through a unified special file processing framework and flexible scheduling mechanism. The highly decoupled nature of different modules allows for continuous solidification of process information, providing a sound information foundation for the ongoing technological iteration of the underlying search platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Schematic diagram of the method flow of the present invention; Figure 2 It is a schematic diagram of the device structure block diagram of the present invention. DETAILED DESCRIPTION
[0023] Example 1: like Figure 1 As shown, a special file enhanced retrieval method for a basic search platform includes: S1, screening special files that need to be search enhanced based on historical information data of the basic search platform, and classifying the special files into PDF special files or WORD special files based on technical characteristics; S2, based on the special file types defined in step S1, formulate a special file location strategy adapted to the basic search platform, and identify special files that meet the definition in step S1 in the multi-source heterogeneous data of the basic search platform; S3, associating the unique identifier of the special file identified in step S2 with the relational database of the basic search platform, obtaining and solidifying the basic information of the special file, and realizing persistence of the basic information of the special file; S4: For PDF special files and WORD special files, based on the special file basic information persisted in step S3, enhanced data processing is performed according to a preset period to generate temporary text data corresponding to the special files; S5, performing data cleaning and processing on the temporary text data formed in step S4 to form pure text enhanced content to be written, so as to ensure the readability of the text content output by the text recognition model and that the text content data can be normally written into the index library; S6, based on the special file basic information persisted in step S3, writing the plain text enhanced content data generated in step S5 into the corresponding index in the basic search platform index library to enhance the indexing capability of the target special file; S7, after completing the indexing capability enhancement in step S6, a search request is made to the index library; the search request is retrieved based on the metadata information and optimized text content of the special file to improve the accuracy and search scope of the search results.
[0024] Preferably, in step S1, the special file classification and definition include: Special PDF files include: PDF scans, encrypted PDFs, and PDF files with an added OCR layer. PDF scans include files that are scanned into image format and then saved as PDFs, from which text cannot be directly extracted or edited. Encrypted PDFs include PDF files that have been set with a blank user password to restrict others from printing, modifying, and annotating. A WORD special file is a WORD document that contains images.
[0025] Preferably, in step S2, the special file location strategy includes: For special PDF files, identification is performed by processing the exported data in the index library: If the field used to record and store the parsed content of the PDF file in the exported index document is empty, it can be determined as a special file, and the file unique identifier is obtained. The file unique identifier is used to associate with the file metadata table to obtain the storage location of the corresponding file; For special WORD files, identify them based on whether they contain pictures: If an image is detected, the document is treated as a special file and the file's unique identifier is used to associate it with the file metadata table to obtain the storage location of the corresponding file.
[0026] Preferably, in step S3, persisting basic information of special files includes: Based on the positioning result of step S2, the unique identifier of the special file is associated with the relational database of the basic search platform to obtain the detailed data of the PDF special file and the WORD special file, and solidified into a fixed special file information table in the structured database; the information table contains at least the following 6 fields: file global unique ID (uniquely identifies each file), file type (identifies the file type, such as PDF or WORD), file storage path (records the storage location of the file in the unstructured data storage system), file processing status (indicates the current document processing progress), creation time (records the time when the data is first entered into the database) and modification time (records the time of the last update).
[0027] Preferably, in step S4, for a PDF special file type, special file enhanced data processing includes: Based on the file storage location field in the persistent basic information in step S3, the physical electronic file to be processed is obtained from the corresponding storage address in the unstructured data storage system; PDF special files: Decrypt the physical electronic file using a PDF tool to generate a temporary PDF file to be processed in the next step; Extract the image objects contained in the temporary PDF file, optimize the image objects, use the text recognition model to perform text recognition on the optimized image data, the recognition results include the text content of the corresponding image data, and integrate the text recognition content of all image objects in the entire document to form temporary text data.
[0028] WORD special files: The image objects contained in the physical electronic file are extracted by an image extraction tool, the image objects are optimized, the optimized image data are subjected to text recognition using a text recognition model, and the recognition contents of all image objects in the entire document are integrated to form temporary text data.
[0029] Furthermore, the key steps of step S4 include: image preprocessing (such as denoising, binarization, tilt correction, etc.), text area detection (such as locating the text area through edge detection and contour analysis), character segmentation (splitting and extracting character features based on character spacing and connectivity), character recognition (using support vector machines or convolutional neural networks to classify and identify character features), and finally restoring the recognized characters into editable coherent text through text integration.
[0030] Preferably, in step S5, performing data cleaning on the temporary text data formed in step S4 includes: The temporary text data generated in step S4 is cleaned, converted and processed, including: special character replacement, adding escape characters, filtering redundant spaces and line breaks, character set conversion and text base64 conversion, so that the text content meets the requirements for being correctly written into the index library.
[0031] Preferably, in step S6, writing the enhanced content into the basic search platform index library includes: Setting a target index, a globally unique identifier of the target special file, plain text enhanced content of the target special file, and a target location; The plain text enhanced content of the target special file is written into the index library of the basic search platform, and operations are performed based on the amount of data: For a small amount of data whose data volume is less than a set threshold, the plain text enhanced content of the target special file is written into the target location of the index library based on the update instruction using the data manipulation language DML; For batch large-scale data with a data volume equal to a set threshold, the plain text enhanced contents of several target special files are written into the target location of the index library through the data manipulation language DML based on the index library batch API.
[0032] Preferably, in step S7, performing a search request on the index library includes: The search query body is generated by performing word segmentation, semantic analysis, and feature word processing (proper noun identification, synonym replacement, error correction word replacement, and sensitive word blocking) on the search content entered by the user. This is combined with the user's desired search object classification scope (such as business category, document type, and time range), recall priority (such as by time or by relevance), and word segmentation method (such as including all keywords, including any keywords, including complete keywords, and not including keywords). In the basic search platform index library after step S6 is completed, the search recall result record is determined through similarity calculation based on the preset threshold, the attribute information and storage information of the data object in the record are extracted, the recall result is generated, and it is returned to the basic search platform front end to be displayed to the search user.
[0033] like Figure 2 As shown, preferably, a basic search platform-oriented special file enhanced retrieval device is used to execute the basic search platform-oriented special file enhanced retrieval method; the device includes the following modules: The file recognition module is used to collect, process and identify special files; the file recognition module specifically includes: The document metadata acquisition unit is used to extract key data information of documents from the index library and relational database data storage components of the basic search platform to support subsequent document processing and recognition processes; The file metadata processing unit is used to process the extracted document metadata information, screen out target files using a special file recognition algorithm, and generate structured information in a unified format. The information is then solidified in a relational database, providing a high-quality data foundation for subsequent processing. The data management module is used to record the information of the special file processing process and support the management of the data processing process. The data management module specifically includes: Process data persistence unit, used to record data processing transaction information and its processing status, to ensure the traceability and reliability of the data processing process; Data visualization unit, used to visualize recorded information, facilitate management and monitoring of data processing, and improve the transparency and efficiency of data processing; The file processing module is used to recognize text in special files, generate temporary text data, and optimize the temporary text data to form enhanced text data to be written. The file processing module specifically includes: Special file decryption unit, used to perform unified decryption operations and remove file protection restrictions; An image processing unit, used for extracting image objects contained in the file and performing optimization processing on the image objects; OCR recognition unit, used for OCR recognition of images, extracting original text information, and providing a more accurate data basis for subsequent text processing; The cleaning and conversion unit is used to uniformly process the extracted original text information, including encoding conversion and special symbol replacement, to improve the quality and consistency of text data; The text writing module is used to write the enhanced text data generated by the file processing module into the search engine of the search platform and monitor the writing process. The text writing module specifically includes: The data writing unit is used to write the enhanced text data into the basic search platform index library in a predetermined format. It supports both batch writing and single writing modes to ensure the efficiency and accuracy of data writing. The write monitoring unit is used to monitor the data writing status in real time, record indicators such as the write success rate and write speed, and issue timely warnings when abnormalities are found to ensure the stability and reliability of the data writing process; The file query module has the following functions: Perform word segmentation, semantic analysis, and feature word processing based on the search content entered by the user; Generate a search query request body based on the user's desired search object classification scope, recall priority, and word segmentation method; Based on the updated index library data, the search recall result records are determined by similarity calculation and based on the preset threshold; Extract the attribute information and storage information of the data objects in the record, generate the recall results, and return them to the front end of the basic search platform to be displayed to the search user.
[0034] Example 2: This embodiment provides a method and device for enhancing retrieval of special files on a basic search platform, including: Step 1: Collect file metadata information. By classifying special files, select appropriate metadata processing methods to process the file metadata information, filter out the "target files", and perform data cleaning and conversion in this process to generate standardized structured information.
[0035] Step 2: Special file processing process information management, tracking the status of file processing by recording each stage of file processing.
[0036] Step 3: Extract text from special files to generate temporary text data, and optimize it to form enhanced text data to be written to ensure the accuracy and readability of the text content.
[0037] Step 4: Write the text data generated in step 3 into the index library of the search platform and update the index of the target document.
[0038] In this example, file metadata is collected using the open-source Elasticsearch-dump component. During the data initialization phase, this component fully exports the indexes specified by the Elasticsearch index repository. Subsequently, combined with the special file processing information from step 2, incremental data synchronization is achieved, and the exported file format is selected based on the requirements of subsequent processing steps.
[0039] In this example, file metadata processing is implemented using the open-source tool Kettle, which is used to process index files exported from Elasticsearch. Kettle is a powerful ETL tool that provides a wealth of data processing capabilities, including data cleaning, transformation, and integration, and can efficiently process metadata information for index files.
[0040] In this embodiment, the persistence of information in the special file processing process is implemented using the structured database MySQL. As a transactional database, MySQL can efficiently support data operations (DML) in the special file processing process, ensuring data consistency and integrity.
[0041] In this embodiment, text extraction for special files is performed by calling the Tesseract OCR open source optical character recognition engine to process images in PDF files and extract the text content. The Tesseract OCR engine supports text recognition in multiple languages and can accurately identify text in images and convert it into editable text format. This text extraction method is highly accurate and reliable, can effectively process text content in PDF files, and provides important support for enhanced retrieval of special files.
[0042] In this example, writing data in batches using Elasticsearch's _bulk API significantly improves writing efficiency and performance. The _bulk API's multi-operation processing capabilities allow it to complete multiple indexing, creation, deletion, or update operations simultaneously, making it the preferred method for writing large amounts of text data into indexes.
[0043] The beneficial effects of the second embodiment are as follows: Without changing the underlying search platform architecture, this embodiment establishes an independently operating special file processing framework, enabling accurate identification and timely processing of special files while avoiding the risks associated with system reconstruction. This framework also boasts excellent scalability. It allows for flexible deployment of solutions tailored to specific special file types and formats, effectively reducing the difficulty of system maintenance. Furthermore, through a unified scheduling and management mechanism, it effectively integrates various special file processing components, improving processing efficiency and enhancing the quality of text content, directly enhancing the user search experience.
[0044] Example 3: Based on the second embodiment, step 1 includes: Step 101: Collect file metadata and use the Elasticsearch-dump tool to batch export the data in the Elasticsearch index repository in CSV format. Considering the potential impact of large-scale indexing on search engine performance, to ensure reliable export operations, batch processing is required based on data characteristics and business needs.
[0045] Step 102: File metadata processing. Use Kettle to perform batch data cleansing and conversion on the CSV files exported in Step 101. Using specialized file recognition logic, documents that failed parsing during the original intelligent search platform processing are screened out and further data integration and standardization is performed. Finally, the processed results are inserted into a designated table in the MySQL database for subsequent data analysis and retrieval.
[0046] In this embodiment, when designing a batch export of standardized structures, two key factors need to be considered simultaneously: export efficiency and subsequent processing requirements. To improve export efficiency, some non-critical fields can be selectively filtered out, while retaining key index fields such as "index ID," "document ID," and "source file globally unique ID." Furthermore, to meet subsequent processing requirements, additional fields need to be exported, such as "title" (document title), "context" (document content), and "parsfilePath" (document parsed file path information).
[0047] In this embodiment, the special file recognition logic refers to the fact that in the currently implemented smart search project, the search engine Elasticsearch's retrieval of documents is mainly based on the data information of the three fields "title", "context" and "parsfilePath" in the index library. Among them, the "title" field stores the title information of the document, the "context" field stores the splicing result of the title and content information of the document, and the "parsfilePath" field stores the parsed file path information of the document. When the document is parsed successfully, the title information will be written to the "title" field, the splicing result of the title and content information will be written to the "context" field, and the parsed file path information will be written to the "parsfilePath" field. When the document fails to be parsed, the title information will be written to the "title" field, only the title information will be written to the "context" field, and the "parsfilePath" field will be empty. By analyzing the information of these three fields, PDF documents that have failed to parse can be identified more accurately.
[0048] Beneficial effects of the third embodiment: 1. Improve export efficiency: By using the Elasticsearch-dump tool for batch export and processing data in batches based on data characteristics and business needs, the efficiency and reliability of the export operation are ensured.
[0049] 2. Data cleaning and standardization: Using Kettle to perform batch data cleaning and conversion on exported CSV files can effectively filter out documents that failed parsing and conduct further data integration and standardization, providing a high-quality data foundation for subsequent data analysis and retrieval.
[0050] 3. Key field retention: When designing a batch export of standardized structures, non-key fields are selectively filtered while retaining key index fields (such as "index ID", "document ID" and "source file globally unique ID") to ensure data integrity and meet subsequent processing needs.
[0051] 4. Special file identification logic: By analyzing the "title", "context" and "parsfilePath" fields, it can more accurately identify PDF documents that failed parsing, thereby improving the accuracy and efficiency of document processing.
[0052] 5. Support subsequent data analysis: Inserting the processing results into the specified table in the MySQL library provides good data support for subsequent data analysis and retrieval, enhancing the overall performance of the system and user experience.
[0053] Example 4: Based on the second embodiment, step 2 includes: Step 201: The data of the special file processing process information is persisted and managed using the structured database MySQL. The status information of the special file during the processing is recorded in detail using a flow table record method.
[0054] Step 202: Data visualization: Use BI tools to perform multi-dimensional analysis and display on the data in the flow table to achieve real-time monitoring and statistical analysis of the special file processing process.
[0055] In this embodiment, the creation of the flow table is divided into the following two cases according to the classification of special files: 1. The basic information for PDF special file persistence includes: "Transaction ID" (used to uniquely identify the processing task), "File Globally Unique ID" (used for file tracing), "Index ID" (corresponding to the Elasticsearch index), "Document ID" (corresponding to the document_id in the Elasticsearch index), "File Source System" (identifying the business source of the file), "File Name", "File Storage Path", "File Processing Status" (recording processing progress), "Creation Time" (recording the time when the file information was persisted), and "Modification Time" (recording the time when the status was updated).
[0056] 2. WORD special file persistence includes two association tables: Document Information Table: Contains "Transaction ID" (used to uniquely identify the processing task), "File Globally Unique ID" (used for file tracing), "Index ID" (corresponding to the Elasticsearch index), "Document ID" (corresponding to the document_id in the Elasticsearch index), "File Source System" (identifying the business source of the file), "File Name", "File Storage Path" (recording processing progress), "File Processing Status" (recording processing progress), "Creation Time" (recording the time when file information is persisted), and "Modification Time" (recording the time when status is updated). Document-associated image information table: Contains "Transaction ID," "Image-associated file transaction ID" (associated with the primary document), "Image storage path," "Image processing status," "Creation time" (recording the time when file information is persisted), and "Modification time."
[0057] In this embodiment, the "File Processing Status" field in the flow table uses digital coding to accurately record the various stages of the special file processing process, including the following 13 status values: Initial state: 0 (pending); Decryption stages: 1 (processing), 2 (completed), 3 (failed); OCR stages: 4 (processing), 5 (completed), 6 (failed); Text cleaning stages: 7 (processing), 8 (completed), 9 (failed); Index writing phases: 10 (processing), 11 (completed), 12 (failed).
[0058] In this embodiment, data visualization is implemented using professional BI tools, which mainly include the following functions: Processing progress monitoring: Real-time display of file processing status distribution at each stage; Processing efficiency analysis: statistics on the processing time and success rate of different types of files; Abnormal situation tracking: Summarize and analyze files in failed state; Trend analysis: Displays historical trends of indicators such as processing volume and success rate.
[0059] Beneficial effects of the fourth embodiment: 1. Reasonable data persistence design: Through the carefully designed flow table structure, accurate recording of the entire process of special file processing is achieved, ensuring data integrity and traceability.
[0060] 2. Refined status management: 13 detailed status values are used to accurately track the file processing process, helping to quickly locate and resolve problems during the processing.
[0061] 3. Improved monitoring system: Through the multi-dimensional analysis function of BI tools, real-time monitoring and statistical analysis of the processing process are achieved, providing intuitive data display and decision support.
[0062] 4. Efficient exception handling: Through status coding and visual display, abnormal situations in the processing process can be quickly discovered and located, improving problem-solving efficiency.
[0063] 5. The system is highly scalable: It adopts a classification management approach and designs independent table structures for different types of special files, which facilitates the subsequent expansion of new file types and processing processes.
[0064] Embodiment 5: Based on the second embodiment, step 3 includes: Step 301: Decrypt special files. Use the open source tool PDF-Guru to decrypt the encrypted PDF file and remove the write protection restriction of the file.
[0065] Step 302: Special file OCR recognition, using open source OCR components to perform OCR recognition on the decrypted PDF file or the image in the Word document to extract the original text information.
[0066] Step 303: Text cleaning, uniformly processing the extracted original text information, including encoding conversion and special symbol replacement.
[0067] In this embodiment, steps 301, 302, and 303 are executed in the order of decryption, OCR recognition, and text cleaning. Each step reads the "File Processing Status" field recorded in step 201 of Example 3 to determine the completion status of the current step and whether the next step can be started. This state-based process control ensures the reliability and traceability of the processing process.
[0068] In this embodiment, the special file decryption processing adopts the following strategy: Although not all PDF documents have write protection restrictions, considering that directly performing OCR recognition on encrypted documents will significantly reduce the recognition success rate and affect the overall processing efficiency, and the decryption operation itself consumes less system resources, a unified processing flow of decryption first and then recognition is adopted. This strategy has been proven to be the optimal processing solution in practice.
[0069] In this embodiment, the data cleaning process mainly includes two key steps: 1. Encoding conversion: To address common encoding issues when writing Chinese documents to Elasticsearch, we perform unified character encoding standardization to ensure the accuracy of data writing.
[0070] 2. Special Symbol Processing: Standardize and normalize special symbols in the text after OCR recognition, including but not limited to: Remove invisible characters; Standardize punctuation format; Normalize whitespace characters; Handle special escape characters.
[0071] The beneficial effects of the above embodiment 5 are as follows: 1. Improved processing reliability: The state-based process control mechanism ensures the stability and reliability of special file processing, effectively avoiding abnormal interruptions and data loss during processing.
[0072] 2. Optimized OCR recognition performance: A unified decryption pre-processing strategy is adopted to significantly improve the success rate of OCR recognition, especially for encrypted PDF documents.
[0073] 3. Enhanced data quality: Through a standardized data cleaning process, the consistency and availability of text data are ensured, providing a high-quality data foundation for subsequent retrieval services.
[0074] 4. Improve system efficiency: Through reasonable process design and resource scheduling, while ensuring processing quality, we optimize processing efficiency and reduce unnecessary resource consumption.
[0075] 5. Ensure data integrity: Comprehensive status management and data cleansing mechanisms ensure data integrity and accuracy during the conversion process from original documents to final retrieved data.
[0076] Example 6: Based on the second embodiment, step 4 includes: Step 401: Write text data to a special file. Use the _bulk API of Elasticsearch to write batch data and write the text data generated in step 3 into the Elasticsearch index library.
[0077] Step 402: Data writing monitoring, monitoring the data writing process through the elk monitoring system to ensure the accuracy and completeness of the data writing.
[0078] In this embodiment, the special file text data is written using the Elasticsearch _bulk API for batch data writing. This API can achieve the following functions: 1. Supports multiple indexing, creation, deletion, or update operations at one time.
[0079] 2. Allow users to submit multiple documents in one request, simplifying the data writing process.
[0080] 3. Provide built-in error handling mechanism to identify and report errors in batch operations.
[0081] 4. Support multiple data formats to meet different types of text data writing needs.
[0082] 5. Provides two modes: batch writing and single writing to meet the needs of different scenarios.
[0083] In this embodiment, data writing monitoring is implemented using the elk monitoring system, which mainly includes three core components: 1. Elasticsearch components: Record the real-time status of write operations; Collect key indicators such as write rate, latency, and success rate; Support storage and query of historical data; 2. Logstash components: Collect logs generated during the writing process in real time; Parse exceptions and error messages in logs; Support custom log processing rules; 3. Kibana components: Provide visual monitoring interface; Support multi-dimensional data analysis; Realize real-time alarm function; Provides issue tracking tools.
[0084] The beneficial effects of the above embodiment 6 are as follows: 1. Write performance optimization: The batch processing mechanism of the _bulk API significantly improves data writing efficiency and reduces system resource consumption.
[0085] 2. Comprehensive monitoring system: Based on multi-dimensional system monitoring, all-round monitoring of the data writing process is achieved, supporting rapid problem location and resolution.
[0086] 3. Data quality assurance: A comprehensive monitoring mechanism ensures the accuracy and completeness of data input, providing high-quality data support for retrieval services.
[0087] 4. Improved system reliability: Through reasonable resource scheduling and error handling, the system's stability and reliability are enhanced.
[0088] 5. Improved operation and maintenance efficiency: The visual monitoring interface and alarm mechanism help operation and maintenance personnel quickly discover and solve problems, thereby improving operation and maintenance efficiency.
[0089] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The embodiments and features in the embodiments of this application may be arbitrarily combined with each other unless they conflict. The scope of protection of the present invention shall be the technical solutions described in the claims, including equivalent alternatives to the technical features of the technical solutions described in the claims. Equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.
Claims
1. A special file enhanced retrieval method for a basic search platform, characterized in that: include: S1, screening special files that need to be search enhanced based on historical information data of the basic search platform, and classifying the special files into PDF special files or WORD special files based on technical characteristics; S2, based on the special file types defined in step S1, formulate a special file location strategy adapted to the basic search platform, and identify special files that meet the definition in step S1 in the multi-source heterogeneous data of the basic search platform; S3, associating the unique identifier of the special file identified in step S2 with the relational database of the basic search platform, obtaining and solidifying the basic information of the special file, and realizing persistence of the basic information of the special file; S4: For PDF special files and WORD special files, based on the special file basic information persisted in step S3, enhanced data processing is performed according to a preset period to generate temporary text data corresponding to the special files; S5, performing data cleaning and processing on the temporary text data formed in step S4 to form the plain text enhanced content to be written; S6, based on the special file basic information persisted in step S3, writing the plain text enhanced content data generated in step S5 into the corresponding index in the basic search platform index library to enhance the indexing capability of the target special file; S7, after the indexing capability enhancement is completed in step S6, a search request is made to the index library; the search request is retrieved according to the metadata information and optimized text content of the special file.
2. The method for enhancing the retrieval of special files for a basic search platform according to claim 1, characterized in that: In step S1, the classification and definition of special files include: Special PDF file types are divided into PDF scans, PDF encrypted files, and PDF files with an OCR layer. A PDF scan is a file that is saved as a PDF file after scanning a document into an image format. The text in the file cannot be directly extracted or edited. An encrypted PDF is a PDF file that has a blank user password set to restrict others from printing, modifying, or annotating. An OCR-layered PDF file is a PDF scan that has been OCR-processed and has an OCR layer added. Users can select and copy the text in the OCR layer. WORD special files refer to a type of WORD document that contains pictures.
3. The method for enhancing the retrieval of special files for a basic search platform according to claim 1, characterized in that: In step S2, the special file location strategy includes: For special PDF files, we analyze, identify and locate the exported data in the index library: If the field used to record and store the parsed content of the PDF file in the exported index document is empty, it can be determined as a special file, and the file unique identifier is obtained. The file unique identifier is used to associate with the file metadata table to obtain the storage location of the corresponding file; For special WORD files, identify and locate them based on whether the document contains images: If an image is detected, the document is treated as a special file and the file's unique identifier is used to associate it with the file metadata table to obtain the storage location of the corresponding file.
4. The method for enhancing the retrieval of special files on a basic search platform according to claim 1, characterized in that: In step S3, the persistence of basic information of special files includes: Based on the positioning result of step S2, the unique identifier of the special file is associated with the relational database of the basic search platform, and the detailed data of the PDF special file and the WORD special file are obtained, and solidified into a fixed special file information table in the structured database; the information table contains at least the following 6 key fields: file global unique ID, file type, file storage path, file processing status, creation time and modification time.
5. The method for enhancing the retrieval of special files on a basic search platform according to claim 1, characterized in that: In step S4, enhanced data processing is performed periodically to generate temporary text data corresponding to the special file: Based on the file storage location field in the persistent basic information in step S3, the physical electronic file to be processed is obtained from the corresponding storage address in the unstructured data storage system; PDF special files: Decrypt the physical electronic file using a PDF tool to generate a temporary PDF file to be processed in the next step; Extracting image objects contained in the temporary PDF file, optimizing the image objects, performing text recognition on the optimized image data using a text recognition model, wherein the recognition results include the text content of the corresponding image data, and integrating the text recognition content of all image objects in the entire document to form temporary text data; WORD special files: The image objects contained in the physical electronic file are extracted by an image extraction tool, the image objects are optimized, the optimized image data are subjected to text recognition using a text recognition model, and the recognition contents of all image objects in the entire document are integrated to form temporary text data.
6. The method for enhancing the retrieval of special files on a basic search platform according to claim 5, characterized in that: The key steps of step S4 include image preprocessing, text area detection, character segmentation and character recognition, and finally the recognized characters are restored to editable coherent text through text integration.
7. The method for enhancing the retrieval of special files on a basic search platform according to claim 1, characterized in that: In step S5, data cleaning and processing of the temporary text data generated in step S4 includes: The temporary text data generated in step S4 is cleaned, converted and processed, including: special character replacement, adding escape characters, filtering redundant spaces and line breaks, character set conversion and text base64 conversion, so that the text content meets the requirements for being correctly written into the index library.
8. The method for enhancing special file retrieval for a basic search platform according to claim 1, characterized in that: In step S6, writing the enhanced content into the basic search platform index library includes: Setting a target index, a globally unique identifier of the target special file, plain text enhanced content of the target special file, and a target location; The plain text enhanced content of the target special file is written into the index library of the basic search platform, and operations are performed based on the amount of data: For a small amount of data whose data volume is less than a set threshold, the plain text enhanced content of the target special file is written into the target location of the index library based on the update instruction using the data manipulation language DML; For batch large-scale data with a data volume equal to a set threshold, the plain text enhanced contents of several target special files are written into the target location of the index library through the data manipulation language DML based on the index library batch API.
9. The method for enhancing special file retrieval for a basic search platform according to claim 1, characterized in that: In step S7, performing a search request on the index library includes: Generate a search query body by performing word segmentation, semantic analysis, and feature word processing on the search content entered by the user, combined with the user's desired search object classification range, recall priority, and word segmentation method; In the basic search platform index library after step S6 is completed, the search recall result record is determined through similarity calculation based on the preset threshold, the attribute information and storage information of the data object in the record are extracted, the recall result is generated, and it is returned to the basic search platform front end to be displayed to the search user.
10. A special file enhanced retrieval device for a basic search platform, characterized in that: The device is used to execute the special file enhanced retrieval method for a basic search platform according to claims 1 to 9, and the device includes the following modules: Document recognition module, used to collect, process and identify special documents; The file identification module specifically includes: A file metadata acquisition unit is used to extract key data information of documents from the index library and relational database data storage component of the basic search platform; The file metadata processing unit is used to process the extracted document metadata information, filter out target files using a special file recognition algorithm, and generate structured information in a unified format, which is then solidified in a relational database. The data management module is used to record the information of the special file processing process and support the management of the data processing process. The data management module specifically includes: Process data persistence unit, used to record data processing transaction information and its processing status; Data visualization unit, used to visualize the recorded information; The file processing module is used to recognize text in special files, generate temporary text data, and optimize the temporary text data to form enhanced text data to be written. The file processing module specifically includes: Special file decryption unit, used to perform unified decryption operations and remove file protection restrictions; An image processing unit, used for extracting image objects contained in the file and performing optimization processing on the image objects; OCR recognition unit, used for OCR recognition of images and extraction of original text information; The cleaning and conversion unit is used to uniformly process the extracted original text information, including encoding conversion and special symbol replacement; The text writing module is used to write the enhanced text data generated by the file processing module into the search engine of the search platform and monitor the writing process. The text writing module specifically includes: A data writing unit is used to write the enhanced text data into the basic search platform index library in a predetermined format, supporting both batch writing and single-entry writing modes; The write monitoring unit is used to monitor the data writing status in real time, record indicators such as the write success rate and write speed, and issue timely warnings when anomalies are found; The file query module is used to perform: Perform word segmentation, semantic analysis, and feature word processing based on the search content entered by the user; Generate a search query request body based on the user's desired search object classification scope, recall priority, and word segmentation method; Based on the updated index library data, the search recall result records are determined by similarity calculation and based on the preset threshold; Extract the attribute information and storage information of the data objects in the record, generate the recall results, and return them to the front end of the basic search platform to be displayed to the search user.