Retrieval index generation method and server using the method

By generating a search index in the server, recording the address information and keywords of the file, the cumbersome problem of users retrieving files in multiple cloud storage devices is solved, and the function of quickly locate files is realized.

CN107644049BActive Publication Date: 2025-05-02AVISION
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN201611108711.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-07-21
Filing Date
2016-12-06
Publication Date
2025-05-02
Estimated Expiration
2036-12-06

AI Technical Summary

Technical Problem

When a user uses multiple cloud storage devices at the same time, it becomes very cumbersome to find files with specific keywords, and it needs to be retrieved one by one in each cloud storage device.

Method used

By implementing the search index generation method in the server, receiving the file access instructions, parsing the file to obtain the keyword string, and determining which database the file is written to according to the access instructions, generating the search index, and recording the file address information and keywords.

Benefits of technology

Users can quickly locate the database where files with specific keywords are located through the server, avoiding the tedious process of searching one by one in multiple cloud storage devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN107644049B_ABST
    Figure CN107644049B_ABST
Patent Text Reader

Abstract

A retrieval index generation method is applicable to a database system having a first database and a second database, the method comprising: receiving an access instruction regarding a first file; parsing the first file to obtain a plurality of keyword strings regarding the first file; writing the first file into the first database or the second database according to the access instruction and generating location information regarding the first file; and generating a retrieval index regarding the first file using the location information and the keyword strings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a retrieval index generation method and a server using the method, and in particular to a distributed database and a retrieval index generation method thereof. Background Art

[0002] Cloud storage devices or services are increasingly being used in people's daily lives. For example, Google Drive ), etc. are all commonly used cloud storage devices (services). Nowadays, people also often upload digital files, such as plain text files, Microsoft documents, portable document formats (PDF documents), etc. to their own cloud storage devices.

[0003] However, as the storage space required for current files is getting larger and larger, people tend to use multiple cloud storage devices or services at the same time. This brings a problem. When a user has multiple cloud storage devices at the same time, the user may not classify his files and store them in the corresponding cloud storage devices. When the user wants to find files with specific keywords from multiple cloud storage devices in the future, he needs to spend a lot of effort to search in each cloud storage device one by one. Summary of the invention

[0004] In view of the above problems, the present invention aims to provide a search index generation method and a corresponding server so that users can find desired files more quickly.

[0005] The retrieval index generation method according to the present invention is applicable to a database system having a first database and a second database, and the method comprises: receiving an access instruction of a first file. Parsing the first file to obtain a plurality of keyword strings related to the first file. Parsing the access instruction to determine whether the first file is written into the first database or the second database, so as to generate address information about the first file. Using the address information and the keyword string, a retrieval index about the first file is generated. Writing the first file into the first database or the second database according to the access instruction.

[0006] The server according to the present invention is suitable for being communicatively connected to a first database and a second database, and the server comprises: a processor and an access bus. When the processor receives an access instruction regarding a first file, the first file is parsed to obtain a plurality of first keyword strings regarding the first file. The access bus is communicatively connected to the processor, the first database, and the second database. The access bus writes the first file into the first database or the second database according to the access instruction, and generates first address information regarding the first file. The processor further generates a first search index regarding the first file using the first address information and the first keyword string.

[0007] In summary, when a file is stored in a database by a server, the information of the target database is used as a part of the search index, so the user can quickly know in which database the file with a specific keyword is stored through the server.

[0008] The above description of the content of the present invention and the following description of the embodiments are intended to demonstrate and explain the spirit and principle of the present invention, and to provide a further explanation of the scope of protection of the claims of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a diagram of a database system architecture according to an embodiment of the present invention;

[0010] Figure 2A is a flow chart of a method for generating a retrieval index according to an embodiment of the present invention;

[0011] Figure 2B is a flowchart of step S220 according to an embodiment of the present invention;

[0012] Figure 3 is a diagram of a database system architecture according to another embodiment of the present invention.

[0013] Component number description

[0014] 1000 Database System

[0015] 1100, 1200 database

[0016] 1300 Server

[0017] 1310 processor

[0018] 1320 Access bus

[0019] 1330 Database

[0020] 3100 Terminal Device

[0021] 3200 Server

[0022] 3300~3500 Database DETAILED DESCRIPTION

[0023] The detailed features and advantages of the present invention are described in detail in the following embodiments, and the contents are sufficient to enable any person skilled in the art to understand the technical content of the present invention and implement it, and according to the contents described in this specification, the scope of the patent application and the drawings, any person skilled in the art can easily understand the relevant purposes and advantages of the present invention. The following examples are to further illustrate the viewpoints of the present invention in detail, but do not limit the scope of the present invention in any viewpoint.

[0024] Please refer to Figure 1 , which is a diagram of the database system architecture according to an embodiment of the present invention. Figure 1 As shown, the database system 1000 in this embodiment has a first database 1100, a second database 1200 and a server 1300. The server 1300 is connected to the first server 1100 and the second server 1200. More specifically, the communication connection means that data can be transmitted and received between the server 1300 and the first server 1100 and / or the second server 1200.

[0025] The server 1300 has a processor 1310, an access bus 1320, and a search database 1330. The access bus 1320 is electrically connected to the processor 1310. The access bus 1320 is also connected in communication with the first server 1100 and the second server 1200. The processor 1310 accesses data to the first server 1100 and / or the second server 1200 through the access bus 1320. The search database 1330 is connected in communication with the processor 1310.

[0026] The aforementioned database may be a physical hard disk, hard disk array, magnetic tape, flash storage medium or other non-volatile storage device. The aforementioned processor may be a central processing unit (CPU), a microcontroller (MCU), an advanced RISC machine (ARM) or other circuits with signal processing, logic operation and electronic device control capabilities.

[0027] In one embodiment, when a user wants to store a first file in the first database 1100, the first file is first sent to the server 1300, together with an access instruction indicating that the first file is to be stored in the first database. Therefore, before the processor 1310 writes the first file into the first database 1100 via the access bus 1320, the processor 1310 parses the first file to obtain a keyword string of the first file. In this embodiment, the keyword string includes the file name of the first file.

[0028] Specifically, the processor 1310 determines whether the first file is a text file. When the first file is not a text file, the processor 1310 performs an image recognition program on the first file to generate a text file of the first file, and extracts the content of the text file to obtain one or more first keyword strings related to the first file. When the first file is a text file, the processor 1310 can directly extract the text file content of the first file to obtain the one or more first keyword strings.

[0029] In the process of determining whether the first file is a text file, in one embodiment, the processor 1310 directly determines based on the file name. More specifically, when the extension is doc, xls, ppt (Microsoft document software file), txt (plain text file), etc., it is determined to be a text file. When the extension is portable document format (pdf), tagged image file format (tif / tiff), etc., it is determined to be a non-text file. If the processor 1310 needs to perform image recognition on the first file, optical character recognition (OCR) is used to generate a text file of the first file.

[0030] Thereafter, the processor 1310 stores the first file in the first database 1100 according to the access instruction, and the search index of the first file is stored in the search database 1330 inside the server 1300. The search index includes the keyword string of the first file and the address (first address) where the first file is written, which is the first database 1100 in this embodiment. In some embodiments, the search index may further include the file name of the first file. In other embodiments, when the first file itself is not a text file, the text file of the first file obtained by the aforementioned image recognition can also be used as part of the search index of the first file. In other words, the text file of the first file can be used as a library archive of the first file, which is convenient for users to review in advance when searching for files to confirm whether the first file is the file required by the user.

[0031] In some cases, the first database 1100 does not receive and store files from the server 1300, but the user directly stores the second file in the first database 1100 without going through the server 1300. In this way, the server 1300 cannot establish a search index for the second file. In order to solve such a problem, the present invention provides the following method. In one embodiment, when the first database 1100 writes the second file, it will notify the server 1300 of the writing information of the second file. When the server 1300 receives the writing information about the second file, the server 1300 retrieves the second file from the first database 1100 and executes the aforementioned process of establishing the search index to establish a second search index for the second file.

[0032] In another embodiment, the first database 1100 does not actively transmit the write information of the second file to the server 1300. The server 1300 confirms the file allocation table (FAT) in the first database 1100 regularly or irregularly. For example, when the server 1300 writes a file to the first database 1100 each time, it requests the file allocation table of the first database 1100 from the first database 1100. In this way, the processor 1310 of the server 1300 can compare the currently obtained file allocation table of the first database 1100 (current configuration table) with the file allocation table (recorded configuration table) of the first database 1100 previously obtained and recorded in the search database 1330. If the current configuration table is different from the recorded configuration table, the processor 1310 performs corresponding processing according to the current configuration table.

[0033] For example, if the current configuration table of the first database 1100 indicates that a third file is stored in the first database 1100, and the third file is not recorded in the record configuration table of the first database 1100. The processor 1310 retrieves the third file from the first database 1100, generates a third search index about the third file by the aforementioned process, and updates the record configuration table at the same time. If the record configuration table of the first database 1100 indicates that a fourth file is stored in the first database 1100, and the current configuration table of the first database 1100 does not have information about the fourth file, the processor 1310 deletes the fourth search index about the fourth file in the search database 1330, and updates the record configuration table.

[0034] Therefore, please refer to Figure 2A , which is a flow chart of a method for generating a search index according to an embodiment of the present invention. Figure 2A As shown, the retrieval index generation method according to the present invention includes the following steps: Step S210, receiving an access instruction for a first file. Step S220, parsing the first file to obtain a plurality of first keyword strings for the first file. Step S230, writing the first file into a first database or a second database according to the access instruction, and generating first address information for the first file. Step S240, generating a first retrieval index for the first file using the first address information and the first keyword string.

[0035] Also, please refer to Figure 2B , which is a flowchart of step S220 according to an embodiment of the present invention. Figure 2BAs shown, step S220 includes the following sub-steps: step S221, determining whether the first file is a document file. When the first file is a document file, executing step S223, extracting the content of the text file to obtain one or more first keyword strings. When the first file is not a document file, executing step S225, performing image recognition on the first file to generate the text content of the first file. After executing step S225, returning to step S223. And after completing step S223, proceeding to the aforementioned step S230.

[0036] When the user wants to find a specific file, the user can connect to the server 1300 and search for a specific keyword in the server, and the file with the keyword, its corresponding file name and the data stored in which database (first database 1100 or second database 1200) can be obtained from the search database 1330. Therefore, the user does not need to spend time and effort searching for a specific file in a separate database.

[0037] Please refer to Figure 3 , which is a database system architecture diagram according to another embodiment of the present invention. Figure 3 As shown, the user needs to access the server 3200 through the terminal device 3100. The architecture of the server 3200 is as follows. Figure 1 The server 1300 in the embodiment of the present invention is a server 1300. The server 3200 records the first access key of the user for the first database 3300, the second access key of the user for the second database 3400, and the third access key of the user for the third database 3500. Each access key is, for example, the account password of the user, and the three access keys can be the same or different. Therefore, regardless of whether the user accesses the server 3200, the server 3200 needs to access data from each database regularly or irregularly. Taking the access of the first database 3300 by the server 3200 as an example, the server 3200 requests the file allocation table from the first database 3300, and checks the record configuration table of the server 3200 itself according to the file allocation table of the first database 3300 to determine whether the files stored in the first database 3300 are the same as those recorded by the server 3200. The record configuration table is updated in this way, and the method thereof is not repeated here. In this embodiment, the server 3200 accesses each database through the Internet.

[0038] In summary, when a file is stored in a database by a server, the information of the target database is used as part of the search index, so the user can quickly know which database the file with a specific keyword is stored in through the server. In addition, the server can also update the file allocation table of the database regularly or irregularly to selectively update the search index.

[0039] The specific embodiments proposed in the detailed description of the preferred embodiments are only used to facilitate the explanation of the technical content of the present invention, rather than narrowly limiting the present invention to the above embodiments. Various changes and implementations made without exceeding the spirit of the present invention and the protection scope of the claims all belong to the protection scope of the present invention.

Claims

1. A search index generation method, applicable to a database system having a first database and a second database, characterized in that: The method comprises: Receive an access instruction for a first file; wherein the first file is a file stored by a server; Before the first file is written into the first database or the second database, the server receives the first file, and a processor of the server parses the first file to obtain a plurality of first keyword strings related to the first file; writing the first file into the first database or the second database according to the access instruction, and generating first address information about the first file; and Generate a first search index about the first file using the first address information and the first keyword string; When the second file is written into the first database, the server receives writing information about the second file from the first database; wherein the second file is a file that the user wants to store directly in the first database and is not stored through the server; The server retrieves the second file from the first database according to the written information; Parsing the second file by a processor of the server to obtain a plurality of second keyword strings related to the second file; generating second address information about the second file according to the write information; and A second search index about the second file is generated using the second address information and the second keyword string.

2. The method for generating a search index as claimed in claim 1, wherein: The step of parsing the first file to obtain the first keyword string related to the first file includes: When the first file is a text file, extracting the content of the text file to obtain the first keyword string; and When the first file is not the text file, image recognition is performed on the first file to generate the content of the first file, thereby obtaining the first keyword string.

3. The method for generating a search index as claimed in claim 2, wherein: The image recognition is optical character recognition (OCR).

4. A server, adapted to be communicatively connected to a first database and a second database, characterized in that: The server records the user's access keys for the first database and the second database, and the server includes: a processor, when receiving an access instruction regarding a first file, before the first file is written into the first database or the second database, the first file is sent to the server, and the processor parses the first file to obtain a plurality of first keyword strings regarding the first file; and an access bus, communicatively connected to the processor, the first database, and the second database, for writing the first file into the first database or the second database according to the access instruction, and generating first address information about the first file; wherein the processor generates a first search index about the first file using the first address information and the first keyword string; The first file is a file stored by the server; When the second file is written into the first database, the processor receives writing information about the second file from the first database, retrieves the second file from the first database according to the writing information, and parses the second file to obtain a plurality of second keyword strings about the second file; The access bus generates second address information about the second file according to the write information; wherein the processor generates a second search index about the second file using the second address information and the second keyword string; The second file is a file that the user wants to store directly in the first database without being stored through the server.

5. The server according to claim 4, characterized in that The invention further comprises an image capturing device, which is communicatively connected to the processor and is used for capturing an image of a paper data to generate the first file.

6. The server according to claim 4, characterized in that The processor further determines whether the first file is a text file. When the first file is a non-text file, the processor performs image recognition on the first file to generate a text file of the first file, and captures the content of the text file of the first file to obtain the first keyword string. When the first file is a text file, the processor captures the content of the first file to obtain the first keyword string.

7. The server according to claim 6, characterized in that The processor processes the first document by optical character recognition (OCR).

8. The server according to claim 4, characterized in that When the first database transmits write information about a second file to the server, the processor determines whether the second file is written by the processor based on the write information. When the second file is not written by the processor, the processor retrieves the second file from the first database through the access bus based on the write information, and parses the second file to obtain a plurality of second keyword strings about the second file. The processor also generates second address information about the second file based on the write information. Using the second address information and the second keyword string, the processor generates a second search index about the second file.

9. A search database, characterized in that: include: A non-volatile storage medium is used to store a first search index, wherein the first search index is related to a first file, and the first search index includes: at least one keyword string about the first file, wherein the at least one keyword string about the first file is obtained by a processor of the server parsing the first file when the first file is sent to the server before the first file is written into the first database or the second database; and a first address, wherein the first address is a first database where the first file is stored; The first file is a file stored by the server; The non-volatile storage medium is further used to store a second search index, wherein the second search index is related to a second file, and the second search index includes: At least one keyword string about the second file, wherein the at least one keyword string about the second file is obtained by the server receiving writing information about the second file from the first database, retrieving the second file from the first database according to the writing information, and parsing the second file when the second file is written into the first database; and a second address, wherein the second address is a first database where the second file is stored; The second file is a file that the user wants to store directly in the first database without being stored through the server.

10. The search database according to claim 9, characterized in that The non-volatile storage medium is further used to store a first file allocation table, and the first file allocation table is related to the first database.

Citation Information

Patent Citations

  • Mobile terminal and file browsing method implemented by same

    CN101916164A

  • Method and system for building indexes and method and system for retrieving indexes

    CN103488709A

  • Method for data access and cloud server system

    CN103795696A

  • Big data based index acquisition method and system

    CN105320746A

  • Document and file indexing system

    WO2007070774A2