Data management method, device and system, storage medium and computer program product

By supporting multi-language and multi-scene word segmentation in the data search system and adaptively determining the word segmentation granularity according to the text length, the problem of poor search performance in the prior art is solved, and more efficient data search is achieved.

CN120234299APending Publication Date: 2025-07-01CHENGDU HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311850630.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The word segmentation algorithm of the existing data search system cannot effectively support multi-language, multi-scene and arbitrary length text, resulting in poor search performance.

Method used

Provide a data management method, by obtaining text and determining the granularity of word segmentation based on its length, supporting multilingual and multi-scene word segmentation, and establishing indexes to improve search speed.

Benefits of technology

This method improves the performance of the data search system, ensures search accuracy and speed, and enhances the robustness of the word segmentation algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234299A_ABST
    Figure CN120234299A_ABST
Patent Text Reader

Abstract

The invention discloses a data management method, device and system, a storage medium and a computer program product, and belongs to the technical field of computers. In the method, a first word segmentation granularity is determined based on the length of a first text, word segmentation is performed on the first text according to the first word segmentation granularity, and an index of a first file is created based on a word segmentation result. Wherein the first text comprises any one or more of a file name of the first file and a file path of the first file. Therefore, the word segmentation algorithm supports texts of multiple languages, multiple scenes and any length, and is more universal and stronger in robustness. Wherein multiple languages indicate that the text to be subjected to word segmentation is not limited to comprise any one or multiple languages, multiple scenes indicate that index establishment based on file names and index establishment based on file paths are supported, and any length indicates that the word segmentation granularity is determined based on the text length and can be adaptive to the text length. According to the scheme, the data search speed can be increased on the premise of ensuring the search accuracy, and the performance of a data search system is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data management method, apparatus, system, storage medium, and computer program product. Background Art

[0002] With the development of the big data era, the data storage volume in the data search system is becoming increasingly large. A reasonable data management solution can accelerate the data search speed and accuracy. Currently, the data management solution in the data search system mainly includes a data storage process and a data search process. In the data storage process, the server uses a word segmentation algorithm to create an index of the file to be stored. In the data search process, the server also uses the same word segmentation algorithm to segment the search term, and uses the multiple words obtained by the segmentation as an index to search for the target file. Among them, the word segmentation algorithm is the key to affecting the performance of the data search system. For example, the index obtained based on a reliable word segmentation algorithm can quickly and accurately search for the target file. Therefore, how to use the word segmentation algorithm to improve the performance of the data search system has always been the research focus in the related field. Summary of the Invention

[0003] This application provides a data management method, apparatus, system, storage medium, and computer program product, which can support multiple languages and multiple scenarios, and can adapt to the length of the text to be segmented, that is, the word segmentation is more robust, can improve the data search speed on the premise of ensuring the search accuracy, and effectively improve the performance of the data search system. The technical solutions are as follows:

[0004] In a first aspect, a data management method is provided, and the method includes:

[0005] Obtain a first text, where the first text includes any one or more of the file name of the first file and the file path of the first file; determine at least one first word segmentation granularity based on the first word segmentation parameter and the length of the first text; segment the first text according to the at least one first word segmentation granularity respectively, to obtain at least one first word corresponding to each first word segmentation granularity, and the length of each first word is equal to the corresponding first word segmentation granularity; create at least one first index based on the at least one first word and the identifier of the first file, where the at least one first index is an index of the first file, each first index includes a key-value pair, and the key in the key-value pair included in each first index is a first word and the value includes the identifier of the first file; store the at least one first index in the index table.

[0006] It can be seen that the word segmentation algorithm in this solution supports multi - language, multi - scenario, and texts of any length, and is a more general word segmentation algorithm. Among them, multi - language means that there is no restriction on file names and file paths, including Chinese, English, or any other one or more languages. Multi - scenario means that it supports both indexing based on file names and indexing based on file paths. Any length means that the word segmentation granularity is determined based on the text length, that is, the word segmentation granularity can adapt to the text length. Therefore, this word segmentation algorithm has stronger robustness, can improve the data search speed on the premise of ensuring search accuracy, and effectively improve the performance of the data search system.

[0007] Among them, based on the first word - segmentation parameter and the length of the first text, determining at least one first word - segmentation granularity includes: based on the first word - segmentation parameter and the length of the first text, determining a second word - segmentation parameter, where the second word - segmentation parameter satisfies the first condition, and the first condition includes M r not exceeding the length of the first text, M represents the first word - segmentation parameter, r represents the second word - segmentation parameter, M is an integer not less than 1, and r is an integer not less than 0; based on the first word - segmentation parameter and the second word - segmentation parameter, determining at least one first word - segmentation granularity.

[0008] Optionally, in order to ensure that the subsequent index obtained based on word segmentation can improve search accuracy, the total number of the at least one first word - segmentation granularity can be relatively large. For example, the at least one first word - segmentation granularity includes M i , and the value range of i includes all integers not less than 0 and not greater than r.

[0009] Optionally, the first condition further includes that the length of the first text does not exceed M r+1 . In this way, the value range of r can be relatively large, and then the at least one first word - segmentation granularity includes larger word - segmentation granularities, that is, the at least one first word - segmentation granularity will not all be small, so as to avoid that the words obtained after word segmentation are all very short, resulting in weak representativeness of each word.

[0010] Among them, segmenting the first text according to the at least one first word - segmentation granularity respectively to obtain at least one first word corresponding to each first word - segmentation granularity includes: based on the first word - segmentation unit, segmenting the first text according to the at least one first word - segmentation granularity respectively to obtain at least one first word corresponding to each first word - segmentation granularity.

[0011] Optionally, the method further includes: providing a file storage interface for instructing a user to determine a first scenario from multiple scenarios, where the multiple scenarios include a file name scenario and a file path scenario, and the multiple scenarios correspond to multiple word segmentation units one by one; obtaining indication information of the first scenario from the file storage interface; and determining, based on the indication information of the first scenario, the word segmentation unit corresponding to the first scenario as a first word segmentation unit. That is, the user can choose whether to store the first file by file name or file path currently, and the server can determine a suitable word segmentation unit based on the user's choice.

[0012] Wherein, when the first text is the file name of the first file, the first word segmentation unit is one character. And / or, when the first text is the file path of the first file, the first word segmentation unit is one layer of path.

[0013] Optionally, when the first text includes the file name of the first file and the file path of the first file, the above-mentioned first scenario includes a file name scenario and a file path scenario, and the first word segmentation unit includes two word segmentation units, that is, the word segmentation unit corresponding to the file name scenario and the word segmentation unit corresponding to the file path. The server performs word segmentation on the first text respectively according to the at least one first word segmentation granularity based on the first word segmentation unit, including: performing word segmentation on the file name of the first file respectively according to the at least one first word segmentation granularity based on the word segmentation unit corresponding to the file name scenario, and performing word segmentation on the file path of the first file respectively according to the at least one first word segmentation granularity based on the word segmentation unit corresponding to the file path. That is, this solution supports word segmentation of both the file name and file path of the same file, so as to establish an index of the first file.

[0014] Optionally, the method further includes: providing a parameter setting interface for instructing a user to input first word segmentation parameters; and obtaining the first word segmentation parameters from the parameter setting interface. That is, this solution supports users to customize word segmentation parameters, thereby improving the flexibility of the word segmentation algorithm.

[0015] Optionally, the method further includes: obtaining a second text for searching a second file; determining at least one second word segmentation granularity based on the first word segmentation parameters and the length of the second text; performing word segmentation on the second text respectively according to the at least one second word segmentation granularity to obtain at least one second word corresponding to each second word segmentation granularity, and the length of each second word is equal to the corresponding second word segmentation granularity; obtaining the values of each second index in at least one second index from an index table, where the at least one second index includes an index in the index table whose key is the same as the second word, and the at least one file identifier set corresponds to the at least one second index one by one; and searching for the second file based on the at least one file identifier set. That is, the server performs data search through a similar word segmentation algorithm.

[0016] Among them, at least one second word segmentation granularity is determined based on the first word segmentation parameter and the length of the second text, including: determining a third word segmentation parameter based on the first word segmentation parameter and the length of the second text, where the third word segmentation parameter satisfies the second condition, and the second condition includes that M k does not exceed the length of the second text, M represents the first word segmentation parameter, k represents the third word segmentation parameter, M is an integer not less than 1, and k is an integer not less than 0; based on the first word segmentation parameter and the third word segmentation parameter, the at least one second word segmentation granularity is determined.

[0017] Optionally, the at least one second word segmentation granularity includes M k . That is, in order to improve the search speed, the total number of the second word segmentation granularities can be smaller, such as one.

[0018] In a second aspect, a data management device is provided, and the data management device has a function of implementing the behavior of the data management method in the first aspect above. The data management device includes one or more modules, and the one or more modules are used to implement the data management method provided in the first aspect above.

[0019] In a third aspect, a data management system is provided, and the system includes a server and a client; the client is used to provide data to be managed to the server; the server is used to implement the data management method provided in the first aspect above.

[0020] In a fourth aspect, a computer device is provided, and the computer device includes a processor and a memory. The memory is used to store a program for executing the data management method provided in the first aspect above, and store data involved in implementing the data management method provided in the first aspect above. The processor is configured to execute the program stored in the memory. The computing device may further include a communication bus, and the communication bus is used to establish a connection between the processor and the memory.

[0021] In a fifth aspect, a computer-readable storage medium is provided, and instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the data management method described in the first aspect above.

[0022] In a sixth aspect, a computer program product including instructions is provided. When the computer program product runs on a computer, the computer is caused to execute the data management method described in the first aspect above.

[0023] The technical effects obtained in the second to sixth aspects above are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be elaborated here. Description of the Drawings

[0024] Figure 1It is an architecture diagram of a data search system involved in a data management method provided by an embodiment of the present application;

[0025] Figure 2 It is a schematic structural diagram of a computer device provided by an embodiment of the present application;

[0026] Figure 3 It is a flowchart of a data management method provided by an embodiment of the present application;

[0027] Figure 4 It is a flowchart of an index creation process provided by an embodiment of the present application;

[0028] Figure 5 It is a flowchart of a data search process provided by an embodiment of the present application;

[0029] Figure 6 It is a schematic diagram of data management of a search system provided by an embodiment of the present application;

[0030] Figure 7 It is a schematic structural diagram of a data management device provided by an embodiment of the present application. Detailed implementation manners

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0032] First, some terms related to the embodiments of the present application are explained.

[0033] Word segmentation algorithm: It is a process of splitting a continuous text into individual words (abbreviated as words) according to certain rules.

[0034] Index table: It is a data structure used to quickly search for documents or records containing specific words. For example, if the index table includes the index of the first file, then the index of the first file is used to quickly search for the first file. The index table in the embodiments of the present application is an inverted index table. In an inverted index, each index includes a key-value pair, that is, it includes a key and a value. During the data storage process, the words obtained by word segmentation are used as the keys of the index, and the identifier of the file to be stored is used as a value of the index. That is, the keys in the index table are the words obtained by word segmentation, and the values are the file identifier sets.

[0035] Metadata search: Metadata search is a search technology that searches for and organizes data by searching and analyzing the metadata of data (that is, the data describing data). In the embodiments of the present application, the metadata of a file includes an index. By matching the retrieval term with the key of the index, the value of the matched index is obtained, that is, the file identifier set included in the matched index is obtained, and relevant files are searched based on the obtained file identifier set.

[0036] Next, a brief introduction to the word segmentation algorithms in the related technologies will be given.

[0037] The word segmentation algorithm is an essential part of the search system. By reformulating the text into a sequence of words, the word segmentation algorithm reveals the meaning of the text. The word segmentation algorithm is widely used in fields such as search engines, natural language processing (NLP), and intelligent recommendation. In the metadata search system, the goal of the word segmentation algorithm is to split the query statement input by the user into a set of keywords so that the search system can match relevant content based on these keywords. Applying the word segmentation algorithm to the metadata search system can improve the accuracy and speed of search results, thereby meeting the requirements of search performance.

[0038] Currently, the main word segmentation algorithms in the metadata search system are as follows:

[0039] 1. Rule-based word segmentation algorithm: Segment the text according to certain rules, such as segmenting according to dictionaries, part-of-speech tagging, etc. This algorithm requires a large number of artificial rules and dictionary support.

[0040] 2. Statistics-based word segmentation algorithm: Segment the text according to statistical models, such as the maximum matching, maximum probability, etc. algorithms. Although this algorithm has a relatively high degree of automation, its accuracy is relatively low.

[0041] 3. Understanding-based word segmentation algorithm: This algorithm does not simply segment according to the dictionary, but through the understanding and analysis of language, combined with context information, to perform more accurate word segmentation on the text. The word segmentation speed of this algorithm is relatively low.

[0042] 4. elasticsearch word segmentation algorithm: Segment according to simple rules, such as: standard - segment by word, simple - segment according to non-alphabetic characters, whitespace - segment according to spaces, etc. Although this algorithm has a simple configuration, it cannot be used as a general word segmentation algorithm.

[0043] The disadvantages of the related technologies at least include the following three points:

[0044] 1. Some word segmentation algorithms only segment according to a single rule or a simple concatenation of rules, without special treatment for texts of different lengths, that is, they cannot adaptively select the word segmentation granularity according to the length of the search term, resulting in too many search word segments for long texts, thus affecting the performance of the search system.

[0045] 2. Do not consider general support for multiple scenarios (file names, file paths, etc.) and multiple languages (Chinese, English, etc.).

[0046] 3. For different scenarios, it is difficult to flexibly configure the word segmentation granularity according to the actual scenario requirements, that is, the word segmentation granularity cannot be customized, which affects the robustness of the search system.

[0047] In view of the above problems, the embodiments of the present application provide a general word segmentation algorithm and a data search system for multiple languages and multiple scenarios, which support users to customize the word segmentation granularity and can adapt to the length of the text to be segmented. The management of data in the data search system is mainly divided into two processes: the data storage process and the data search process. Among them, the key to the data storage process lies in index construction, and the data search process is, for example, online file search.

[0048] Next, the data search system provided by the embodiments of the present application will be introduced.

[0049] Figure 1 is the architecture diagram of the data search system involved in a data management method provided by the embodiments of the present application. Refer to Figure 1 This system architecture includes a client and a server, and the server includes a storage device. A wireless or wired communication connection is established between the client and the server.

[0050] The client is used to submit a file to be stored to the server, and the server is used to store the file submitted by the client according to the data management method provided by the embodiments of the present application.

[0051] The client is also used to send a data search request to the server and receive the file returned by the server. The server is also used to respond to the data search request of the client and return the corresponding file to the client.

[0052] Optionally, the client includes a user interface through which the user can store and search for files, and can also be used to customize word segmentation parameters, etc. Among them, the user interface includes a file storage interface, a file retrieval interface, a parameter setting interface, etc. The file storage interface is used to instruct the user to input the file to be stored. The file retrieval interface is used to instruct the user to input a search term. The parameter setting interface is used to instruct the user to input the first word segmentation parameter.

[0053] In the embodiments of the present application, the client is any terminal device, such as a mobile phone, a computer, etc., and the client can also refer to an application installed on the terminal device. The server is any server or service device, such as a cloud server or a server cluster, etc.

[0054] It should be understood that the system architecture and business scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0055] Next, a computer device provided by an embodiment of the present application will be introduced. This computer device can be part or all of a client or a server.

[0056] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a computer device shown according to an embodiment of the present application. Optionally, this computer device is the Figure 1 client or server shown in , and this computer device includes one or more processors 201, a communication bus 202, a memory 203, and one or more communication interfaces 204.

[0057] The processor 201 is a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solution of the present application. For example, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. Optionally, the above PLD is a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0058] The communication bus 202 is used to transfer information between the above components. Optionally, the communication bus 202 is divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0059] Optionally, the memory 203 is a read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), optical disc (including compact disc read-only memory (CD-ROM), compressed optical disc, laser disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage medium, or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited thereto. The memory 203 exists independently and is connected to the processor 201 through the communication bus 202, or the memory 203 is integrated with the processor 201.

[0060] The communication interface 204 uses any device such as a transceiver for communicating with other devices or communication networks. The communication interface 204 includes a wired communication interface, and optionally, also includes a wireless communication interface. Among them, the wired communication interface is, for example, an Ethernet interface, etc. Optionally, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.

[0061] Optionally, in some embodiments, the computer device includes multiple processors, such as Figure 2 the processor 201 and the processor 205 shown in. Each of these processors is a single-core processor or a multi-core processor. Optionally, the processor here refers to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0062] In a specific implementation, as an embodiment, the computer device further includes an output device 206 and an input device 207. The output device 206 communicates with the processor 201 and can display information in various ways. For example, the output device 206 is a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 207 communicates with the processor 201 and can receive user input in various ways. For example, the input device 207 is a mouse, a keyboard, a touch screen device, or a sensing device, etc.

[0063] In some embodiments, the memory 203 is used to store the program code 210 for executing the solution of this application, and the processor 201 can execute the program code 210 stored in the memory 203. The program code includes one or more software modules. Taking this computer device as an example of a server, this computer device can implement the data management method provided in the following Figure 3 embodiments through the processor 201 and the program code 210 in the memory 203.

[0064] Next, the data management method provided in the embodiments of this application will be introduced.

[0065] Figure 3 is a flowchart of a data management method provided in the embodiments of this application. This method is applied to a data search system, which can be referred to as a metadata search system. This method is specifically applied to the server in the data search system. Please refer to Figure 3 , and this method includes the following steps.

[0066] Step 301: Obtain a first text, where the first text includes any one or more of the file name of the first file and the file path of the first file.

[0067] Among them, the first file is a file to be stored. The first text can be the file name of the first file, or the file path of the first file, or can also include both the file name of the first file and the file path of the first file. The user can choose to establish an index for the first file according to the file name of the first file, or choose to establish an index for the first file according to the file path of the first file, or choose to establish an index for the first file according to both the file name of the first file and the file path of the first file according to the actual situation.

[0068] It should be understood that the user submits the first file to the server through the client, and the attribute information of the submitted first file includes the first text. After receiving the first file, the server obtains the file name of the first file and / or the file path of the first file from the attribute information of the first file to obtain the first text, and uses the first text as the text to be segmented.

[0069] In the embodiments of this application, the server provides a data storage interface, and the data storage interface is used to instruct the user to submit the first file. That is, the data storage interface is displayed on the client, and the user can submit the first file through the data storage interface.

[0070] The user can submit one file at a time, and the server obtains the file name and / or file path of this file, so as to establish an index for this one file. The user can also submit a batch of files at a time, and the server obtains the file names and / or file paths of the individual files in this batch of files, so as to establish an index for the individual files in this batch of files. In the embodiment of the present application, the process of establishing an index is introduced by taking the first file as an example. The process of the server establishing an index for other files except the first file is similar to the process of establishing an index for the first file.

[0071] Step 302: Determine at least one first word segmentation granularity based on the first word segmentation parameter and the length of the first text.

[0072] After the server obtains the first text, it performs word segmentation on the first text according to the word segmentation algorithm provided in the embodiment of the present application.

[0073] Among them, the server first determines at least one first word segmentation granularity based on the first word segmentation parameter and the length of the first text. That is, the word segmentation algorithm can adapt to the length of the text to be segmented, and the determined word segmentation granularity matches the length of the text to be segmented.

[0074] One implementation manner for the server to determine at least one first word segmentation granularity based on the first word segmentation parameter and the length of the first text is: determine a second word segmentation parameter based on the first word segmentation parameter and the length of the first text, and the second word segmentation parameter satisfies the first condition, and the first condition includes M r not exceeding the length of the first text, where M represents the first word segmentation parameter, r represents the second word segmentation parameter, M is an integer not less than 1, and r is an integer not less than 0; determine the at least one first word segmentation granularity based on the first word segmentation parameter and the second word segmentation parameter.

[0075] Exemplarily, if the length of the first text is 8 and M = 2, then r cannot exceed 3, and r can be 1, 2, or 3.

[0076] Optionally, in order to ensure that the subsequent index obtained based on word segmentation can improve the search accuracy, the total number of the at least one first word segmentation granularity can be relatively large. For example, the at least one first word segmentation granularity includes M i , and the value range of i includes all integers not less than 0 and not greater than r. That is, the total number of the at least one first word segmentation granularity is multiple.

[0077] Exemplarily, if the length of the first text is 8, M = 2, and r = 3, then the server determines 1, 2, 4, and 8 as the first word segmentation granularities, that is, the at least one first word segmentation granularity includes 1, 2, 3, and 4.

[0078] Optionally, to ensure that the total number of indexes obtained based on the at least one first word segmentation granularity is moderate later, the total number of the at least one first word segmentation granularity can also be smaller. For example, the at least one first word segmentation granularity includes M q , where q takes an integer not less than 0 and not greater than r. That is, the total number of the at least one word segmentation granularity is one.

[0079] Exemplarily, the length of the first text is 8, M = 2, r = 2, and q takes a value of 0 or 1 or 2. That is, the server determines 1 or 2 or 4 as the first word segmentation granularity.

[0080] Optionally, the first condition further includes that the length of the first text does not exceed M r+1 . For example, if the length of the first text is 8 and M = 2, then r = 3. Another example is that if the length of the first text is 50 and M = 2, then r = 5. In this way, the value of r can be made larger, so that the at least one first word segmentation granularity includes larger word segmentation granularities, that is, it will not make all of the at least one first word segmentation granularities smaller, so as to avoid that the words obtained after word segmentation are all very short, resulting in weak representativeness of each word.

[0081] As can be seen from the above, the embodiments of the present application support user-defined word segmentation parameters, and the word segmentation parameters are used to determine the word segmentation granularity. Therefore, the user-defined word segmentation parameters are equivalent to the user-defined word segmentation granularity.

[0082] Optionally, the server provides a parameter setting interface, which is used to instruct the user to input the first word segmentation parameters, and the server obtains the first word segmentation parameters from the parameter setting interface. Among them, the parameter setting interface may include some prompt information about parameter setting, and the prompt information may include some word segmentation examples corresponding to the word segmentation parameters. For example, the word segmentation result of the text "Zhang San's Autobiography" when the first word segmentation parameter is 2. The user can refer to the prompt information to set the first word segmentation parameters more reasonably.

[0083] In some other embodiments, the server is default-configured with the first word segmentation granularity and / or the first word segmentation parameters, and the default-configured first word segmentation granularity and / or the first word segmentation parameters can be determined according to experience or data statistics or other means.

[0084] Step 303: Segment the first text according to the at least one first word segmentation granularity respectively, to obtain at least one first word corresponding to each first word segmentation granularity, and the length of each first word is equal to the corresponding first word segmentation granularity.

[0085] After determining the at least one first word segmentation granularity, the server segments the first text according to each first word segmentation granularity, to obtain at least one first word corresponding to each first word segmentation granularity.

[0086] Among them, a way for the server to segment the first text according to at least one first word segmentation granularity to obtain at least one first word corresponding to each first word segmentation granularity is as follows: Based on the first word segmentation unit, segment the first text according to the at least one first word segmentation granularity respectively to obtain at least one first word corresponding to each first word segmentation granularity. Among them, the first word segmentation unit refers to the smallest unit for segmenting the first text.

[0087] A way for the server to determine the first word segmentation unit is as follows: Provide a file storage interface, which is used to instruct the user to determine the first scenario from the multiple scenarios, and the multiple scenarios include the file name scenario and the file path scenario, and the multiple scenarios correspond to multiple word segmentation units one by one; Obtain the indication information of the first scenario from the file storage interface; Based on the indication information of the first scenario, determine the word segmentation unit corresponding to the first scenario as the first word segmentation unit. That is to say, the user can choose whether to establish the index of the first file by file name or file path currently, and the server can determine the appropriate word segmentation unit based on the user's selection.

[0088] Optionally, when the first text includes the file name of the first file and the file path of the first file, the above-mentioned first scenario includes the file name scenario and the file path scenario, and the first word segmentation unit includes two word segmentation units, that is, the word segmentation unit corresponding to the file name scenario and the word segmentation unit corresponding to the file path. The server segments the first text according to the at least one first word segmentation granularity based on the first word segmentation unit, including: Segment the file name of the first file according to the at least one first word segmentation granularity based on the word segmentation unit corresponding to the file name scenario, and segment the file path of the first file according to the at least one first word segmentation granularity based on the word segmentation unit corresponding to the file path. That is to say, this solution supports the user to select the above two scenarios at the same time to instruct the server to segment both the file name and the file path of the same file, so as to establish the index of the file.

[0089] Another way for the server to determine the first word segmentation unit is as follows: After the server obtains the first text, determine whether the first text is a file name or a file path by parsing the first text. When it is determined that the first text is a file name, the server determines the word segmentation unit corresponding to the file name scenario as the first word segmentation unit. When it is determined that the first text is a file path, the server determines the word segmentation unit corresponding to the file path scenario as the first word segmentation unit.

[0090] Optionally, the server default configuration indexes files according to the file name and / or file path. If the server default configuration indexes files according to the file name, then the first file obtained by the server is the file name of the first file, and the first tokenization unit is the tokenization unit corresponding to the file name scenario. If the server default configuration indexes files according to the file path, then the first file obtained by the server is the file path of the first file, and the first tokenization unit is the tokenization unit corresponding to the file path scenario. If the server default configuration indexes files according to the file name and the file path, then the first file obtained by the server includes the file name of the first file and the file path of the first file, and the first tokenization unit includes the tokenization unit corresponding to the file name scenario and the tokenization unit corresponding to the file path scenario.

[0091] Among them, when the first text is the file name of the first file, the first tokenization unit is one character, that is, tokenization is performed in units of characters. When the first text is the file path of the first file, the first tokenization unit is one layer of path, that is, tokenization is performed in units of one layer of path. When the first text includes the file name of the first file and the file path of the first file, the first tokenization unit includes the above two tokenization units, and the server tokenizes the file name of the first file in units of characters and tokenizes the file path of the first file in units of one layer of path.

[0092] Exemplarily, Table 1 shows the results of tokenizing various texts according to different first tokenization parameters provided by the embodiments of the present application. When tokenizing the file names "X 1 X 2 X 3 X 4 X 5 X 6 X 7 " and "day_hours", it is in units of characters, and when tokenizing the path " / a / b / c / d / e / f", it is in units of one layer of path. Among them, one "X 1 X 2 X 3 X 4 X 5 X 6 X 7 " in the file name "X i " represents one Chinese character.

[0093] Table 1

[0094]

[0095]

[0096] For spaces in file names, the server can automatically remove them, that is, spaces are not regarded as characters during word segmentation. Of course, for specified symbols in file names (such as dashes, punctuation marks, etc.), the server can also automatically remove them, and the specified characters can be determined through configuration.

[0097] Step 304: Create at least one first index based on the at least one first word and the identifier of the first file.

[0098] Among them, the at least one first index is the index of the first file, each first index includes a key-value pair, and the key in the key-value pair included in each first index is a first word and the value includes the identifier of the first file.

[0099] That is to say, after word segmentation by the server, each first word obtained by word segmentation is used as the key of the index of the first file, the keys of the at least one first index are obtained, and the identifiers of the first file are used as the values of the at least one first index.

[0100] Exemplarily, the server uses each word in [a, b, c, d, e, f, abc, bcd, cde, def] obtained by word segmentation as the key of the index of the file with the path " / a / b / c / d / e / f", and uses the identifier of this file as the value of the index of this file.

[0101] Step 305: Store the at least one first index in the index table.

[0102] The index table is used to store the key-value pairs included in each index. Among them, the keys in the index table are the words obtained by word segmentation, the multiple keys in the index table are different words, the values in the index table are a set of file identifiers, a set of file identifiers includes one or more file identifiers, and the values corresponding to different keys in the index table may include the same file identifier.

[0103] After the server creates the above at least one first index, it stores the at least one first index in the index table. For example, it stores the identifier of the first file in the set of file identifiers included in the index with the key being the first word.

[0104] Among them, if the server determines that there is already an index in the index table whose key is equal to a certain or certain first words after obtaining at least one first word by word segmentation, then the server adds the identifier of the first file to the set of file identifiers included in this or these indexes. If there is no index in the index table whose key is equal to a certain first word, then the server creates an index corresponding to this first word in the index table, the key of the created index is equal to this first word, and the value of the created index includes the identifier of the first file.

[0105] Exemplarily, after the server uses each word in [a, b, c, d, e, f, abc, bcd, cde, def] as the key of the index of the file with the path " / a / b / c / d / e / f", taking the file identifier of this file as "File 4" as an example, the identifier "File 4" is added to the set of file identifiers corresponding to the 10 index keys of "a", "b", "c"... "def".

[0106] It can be seen that the index table in the embodiment of the present application is an inverted index table. The specific method of the inverted index can also refer to the related art.

[0107] It should be understood that the server also stores the first file according to the identifier of the first file. For example, the first file is stored in a database. Among them, the database is part of the server, or the database and the server are two independent devices. The database can be implemented based on a disk or an optical disc, etc.

[0108] Figure 4 is a flowchart of an index creation process provided by an embodiment of the present application. Refer to Figure 4 , the user customizes the first word segmentation parameter M and selects a scenario, uploads the file to be stored (including the file name and / or file path), and calls the file name and / or file path the first text. The server determines the first word segmentation granularity T based on the user-customized first word segmentation parameter M and the length of the first text, performs word segmentation on the first text according to the first word segmentation granularity T, uses the words after word segmentation as the keys of the index of this file, uses the identifier of this file as the value of the index of this file, stores the index of this file in the inverted index table, and stores the file according to the file identifier.

[0109] Exemplarily, assuming that the user-customized word segmentation parameter is M, then the word segmentation granularity T is: {M 0 , M 1 , M 2 , …, M r}, where r satisfies the first condition introduced above. Word segmentation is performed on the first text input by the user with the word segmentation granularities in T respectively. For the file name, word segmentation is performed in units of characters, and for the file path, one layer of the path is used as the basic word segmentation unit, as shown in Table 1 above.

[0110] This solution can improve the linear long-character search speed with the growth of non-linear storage requirements. Assuming that the length of the text to be word-segmented is L, the data volume of the file to be stored is N, and the total number of word segmentation granularities is K. Taking the text to be word-segmented as the file name as an example, the number of different characters in the text to be word-segmented is n, then there are at most n K combinations in the word segmentation result. There are the following several assumptions: the probability of obtaining each combination by word segmentation is the same; since N is usually much larger than n K, assume that the index contains all different combinations; the probability that multiple words obtained by segmenting the same text are the same is very low. Let's assume that the multiple words obtained by segmenting the text are all different. Then, according to the related art, the size of the inverted index segmented only at one q granularity is: Σ(L_q)≈N(Σ(L)-q + 1). Similarly, the size of the inverted index segmented by this solution is: Σ(L_d)≈N(Σ(L)-T + 1)log_M[Σ(L)]. Among them, Σ(L) represents the length of the text to be segmented. In the worst case, this solution can increase the segmentation storage space to about log M Σ(L), and improve the search speed by Σ(L) times. Among them, Σ represents the expected value.

[0111] Next, the data search process in the embodiments of the present application will be introduced.

[0112] The server obtains a second text, which is used to search for a second file. Based on the first segmentation parameter and the length of the second text, at least one second segmentation granularity is determined. The server segments the second text according to the at least one second segmentation granularity respectively, and obtains at least one second word corresponding to each second segmentation granularity. The length of each second word is equal to the corresponding second segmentation granularity. The server obtains the values of each second index in at least one second index from the index table, and obtains at least one file identifier set. The at least one second index includes the index in the index table whose key is the same as the second word. The at least one file identifier set corresponds to the at least one second index one by one. The server searches for the second file based on the at least one file identifier set.

[0113] The second text refers to the retrieval term input by the user. The retrieval term can also be called a keyword, a search term, etc. For the same user, after the user sets the first segmentation parameter for the first time, there is no need to repeat inputting the first segmentation parameter when retrieving files, so as to ensure the stability of the segmentation algorithm. Of course, the user can also set the first segmentation parameter again when storing the next or a batch of files. Optionally, the user can modify the first segmentation parameter at any time.

[0114] Among them, one implementation manner for the server to determine at least one second segmentation granularity based on the first segmentation parameter and the length of the second text is: based on the first segmentation parameter and the length of the second text, a third segmentation parameter is determined. The third segmentation parameter satisfies a second condition, and the second condition includes that M k does not exceed the length of the second text, M represents the first segmentation parameter, k represents the third segmentation parameter, M is an integer not less than 1, and k is an integer not less than 0. The server determines at least one second segmentation granularity based on the first segmentation parameter and the third segmentation parameter.

[0115] Optionally, the at least one second segmentation granularity includes M kThat is, to improve the search speed, the total number of the second word segmentation granularities can be smaller, for example, it can be one.

[0116] Of course, the total number of the at least one second word segmentation granularity can also be multiple. For example, the at least one second word segmentation granularity includes M j , where the value range of j includes all integers not less than 0 and not greater than k.

[0117] The specific implementation process of the server for word segmentation of the second text is similar to the process of word segmentation of the first text, and will not be introduced in detail here.

[0118] One implementation manner for the server to search for the second file based on the at least one file identifier set is as follows: The server determines the intersection of the at least one file identifier set to obtain the identifiers of multiple candidate files; obtains the multiple candidate files based on the identifiers of the multiple candidate files; and uses the multiple candidate files as the second file.

[0119] Another implementation manner for the server to search for the second file based on the at least one file identifier set is as follows: The server determines the intersection of the at least one file identifier set to obtain the identifiers of multiple candidate files; scores the multiple candidate files based on the identifiers of the multiple candidate files, and determines the second file from the multiple candidate files according to the scoring results. Among them, there can be multiple scoring manners. For example, the scoring can be performed in a manner combining one or more manners such as storage time, file popularity, etc. The present application does not limit this.

[0120] Figure 5 is a flowchart of a data search process provided by an embodiment of the present application. Refer to Figure 5 , the user customizes the first word segmentation parameter M, the user submits the search term of the file to be retrieved, and the server determines the second word segmentation granularity based on the first word segmentation parameter M and the length of the search term The search term is segmented according to the second word segmentation granularity T, the corresponding file identifier is found from the index table based on the segmented words, and then the corresponding file is obtained from the storage to obtain the search result and return it to the user.

[0121] Exemplarily, assume the search term is Query and the length is L. Then the word segmentation granularity T can be determined according to the following logic:

[0122] if L == 1 or L < M:

[0123] T = 1

[0124] Else:

[0125]

[0126] The word segmentation results Tokens obtained using the word segmentation granularity T are as follows:

[0127] if L == T:

[0128] Tokens = {Query}

[0129] Else:

[0130] Tokens = {Query[0:T], … Query[L-T:L]}

[0131] In the related art, the word segmentation results Tokens segmented only by one q granularity are {Query[0:q], Query[1:q+1] … Query[L-q:L]}. Compared with this technology, in this solution, the number of segmented words is reduced from L-q+1 to 1 to M through a word segmentation algorithm with adaptive word segmentation granularity interval. Through the interval adaptive word segmentation algorithm in this solution, the word segmentation granularity interval can be dynamically determined, the number of matching phrases in the index can be reduced, and a faster matching speed can be obtained.

[0132] Figure 6 It is a schematic diagram of data management of a search system provided by an embodiment of the present application. Refer to Figure 6 , during the index creation process, the user can customize the word segmentation parameters, and the server creates the index according to this solution. During the data search process, the server still segments words based on the word segmentation parameters customized by the user, and then queries the identifiers of the corresponding files from the index table according to the word segmentation results.

[0133] In summary, in the embodiment of the present application, the first word segmentation granularity can be determined based on the first word segmentation parameter and the length of the first text. By segmenting the first text according to the first word segmentation granularity, the index of the first file can be obtained, and the length of each first word obtained by word segmentation is equal to the first word segmentation granularity. Among them, the first text includes any one or more of the file name of the first file and the file path of the first file. It can be seen that the word segmentation algorithm in this solution supports multi-language, multi-scene, and texts of any length, and is a more general word segmentation algorithm. Among them, multi-language means that there is no limit to whether the file name and file path include Chinese, English, or any other one or more languages. Multi-scene means that it supports creating indexes based on file names and also supports creating indexes based on file paths. Any length means that the word segmentation granularity is determined based on the text length, that is, the word segmentation granularity can adapt to the text length. Therefore, the robustness of this word segmentation algorithm is stronger, and it can improve the data search speed on the premise of ensuring the search accuracy, effectively improving the performance of the data search system.

[0134] Figure 7It is a schematic structural diagram of a data management device provided by an embodiment of the present application. The data management device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be Figure 1 the server shown. Refer to Figure 7 , the device includes: a first acquisition module 701, a first determination module 702, a first word segmentation module 703, an index creation module 704, and an index storage module 705.

[0135] The first acquisition module 701 is configured to acquire a first text, where the first text includes any one or more of the file name of the first file and the file path of the first file;

[0136] The first determination module 702 is configured to determine at least one first word segmentation granularity based on the first word segmentation parameter and the length of the first text;

[0137] The first word segmentation module 703 is configured to segment the first text according to the at least one first word segmentation granularity respectively, to obtain at least one first word corresponding to each first word segmentation granularity, and the length of each first word is equal to the corresponding first word segmentation granularity;

[0138] The index creation module 704 is configured to create at least one first index based on the at least one first word and the identifier of the first file. The at least one first index is an index of the first file, and each first index includes a key-value pair. The key in the key-value pair included in each first index is a first word, and the value includes the identifier of the first file;

[0139] The index storage module 705 is configured to store the at least one first index in an index table.

[0140] Optionally, the first determination module 702 includes:

[0141] A first determination sub-module is configured to determine a second word segmentation parameter based on the first word segmentation parameter and the length of the first text. The second word segmentation parameter satisfies a first condition, and the first condition includes that M r does not exceed the length of the first text, M represents the first word segmentation parameter, r represents the second word segmentation parameter, M is an integer not less than 1, and r is an integer not less than 0;

[0142] A second determination sub-module is configured to determine the at least one first word segmentation granularity based on the first word segmentation parameter and the second word segmentation parameter.

[0143] Optionally, the at least one first word segmentation granularity includes M i , and the value range of i includes all integers not less than 0 and not greater than r.

[0144] Optionally, the first condition further includes that the length of the first text does not exceed M r+1 .

[0145] Optionally, the first word segmentation module 703 includes:

[0146] A first word segmentation sub-module, configured to segment the first text respectively according to the at least one first word segmentation granularity based on the first word segmentation unit, so as to obtain at least one first word corresponding to each first word segmentation granularity.

[0147] Optionally, the apparatus further includes:

[0148] A first providing module, configured to provide a file storage interface, where the file storage interface is used to instruct a user to determine a first scenario from multiple scenarios, the multiple scenarios include a file name scenario and a file path scenario, and the multiple scenarios correspond to multiple word segmentation units one by one;

[0149] A second obtaining module, configured to obtain indication information of the first scenario from the file storage interface;

[0150] A second determining module, configured to determine the word segmentation unit corresponding to the first scenario as the first word segmentation unit based on the indication information of the first scenario.

[0151] Optionally, when the first text is the file name of a first file, the first word segmentation unit is one character.

[0152] Optionally, when the first text is the file path of a first file, the first word segmentation unit is one layer of path.

[0153] Optionally, the apparatus further includes:

[0154] A second providing module, configured to provide a parameter setting interface, where the parameter setting interface is used to instruct a user to input first word segmentation parameters;

[0155] A third obtaining module, configured to obtain the first word segmentation parameters from the parameter setting interface.

[0156] Optionally, the apparatus further includes:

[0157] A fourth obtaining module, configured to obtain a second text, where the second text is used to search for a second file;

[0158] A third determining module, configured to determine at least one second word segmentation granularity based on the first word segmentation parameters and the length of the second text;

[0159] A second word segmentation module, configured to segment the second text respectively according to the at least one second word segmentation granularity, so as to obtain at least one second word corresponding to each second word segmentation granularity, and the length of each second word is equal to the corresponding second word segmentation granularity.

[0160] A fifth acquisition module, configured to acquire values of each of at least one second index from an index table, to obtain at least one set of file identifiers, where the at least one second index includes indexes in the index table with keys the same as a second word, and the at least one set of file identifiers corresponds to the at least one second index one by one;

[0161] A file search module, configured to search for a second file based on the at least one set of file identifiers.

[0162] Optionally, the third determination module includes:

[0163] A third determination sub-module, configured to determine a third word segmentation parameter based on a first word segmentation parameter and the length of a second text, where the third word segmentation parameter satisfies a second condition, and the second condition includes that M k does not exceed the length of the second text, M represents the first word segmentation parameter, k represents the third word segmentation parameter, M is an integer not less than 1, and k is an integer not less than 0;

[0164] A fourth determination sub-module, configured to determine the at least one second word segmentation granularity based on the first word segmentation parameter and the third word segmentation parameter.

[0165] Optionally, the at least one second word segmentation granularity includes M k .

[0166] In an embodiment of the present application, based on the first word segmentation parameter and the length of the first text, the first word segmentation granularity can be determined, and by segmenting the first text according to the first word segmentation granularity, the index of the first file can be obtained, and the length of each first word obtained by segmentation is equal to the first word segmentation granularity. Among them, the first text includes any one of the file name of the first file and the file path of the first file. It can be seen that the word segmentation algorithm in this solution supports multi-languages, multi-scenarios, and texts of any length, and is a more general word segmentation algorithm. Among them, multi-languages means that there is no limitation on whether the file name and file path include Chinese, English, or any other one or more languages, multi-scenarios means that it supports both establishing indexes based on file names and establishing indexes based on file paths, and any length means that the word segmentation granularity is determined based on the text length, that is, the word segmentation granularity can adapt to the text length. Therefore, the word segmentation algorithm has stronger robustness, can improve the data search speed on the premise of ensuring the search accuracy, and effectively improve the performance of the data search system.

[0167] It should be noted that: when the data management device provided in the above embodiments manages data, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data management device provided in the above embodiments and the embodiments of the data management method belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0168] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)). It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.

[0169] It should be understood that the "at least one" mentioned herein refers to one or more, and the "multiple" refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; the "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, for the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the words such as "first" and "second" do not limit the quantity and execution order, and the words such as "first" and "second" do not necessarily limit to be different.

[0170] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the first file, the second file, etc. involved in the embodiments of the present application are all obtained under sufficient authorization.

[0171] The above are the embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A data management method, characterized in that, The method includes: Obtaining a first text, where the first text includes any one or more of the file name of a first file and the file path of the first file; Determining at least one first word segmentation granularity based on a first word segmentation parameter and the length of the first text; Performing word segmentation on the first text respectively according to the at least one first word segmentation granularity to obtain at least one first word corresponding to each first word segmentation granularity, and the length of each first word is equal to the corresponding first word segmentation granularity; Creating at least one first index based on the at least one first word and the identifier of the first file, where the at least one first index is an index of the first file, and each first index includes a key-value pair, the key in the key-value pair is one of the first words and the value includes the identifier of the first file; Storing the at least one first index in an index table.

2. The method according to claim 1, characterized in that, The determining at least one first word segmentation granularity based on a first word segmentation parameter and the length of the first text includes: Determine a second word segmentation parameter based on the first word segmentation parameter and the length of the first text, where the second word segmentation parameter satisfies a first condition, and the first condition includes that M r does not exceed the length of the first text, where M represents the first word segmentation parameter, r represents the second word segmentation parameter, M is an integer not less than 1, and r is an integer not less than 0; Determining the at least one first word segmentation granularity based on the first word segmentation parameter and a second word segmentation parameter.

3. The method according to claim 2, characterized in that, The at least one first word segmentation granularity includes M i , where the value of i includes all integers not less than 0 and not greater than r.

4. The method according to claim 2 or 3, characterized in that, The first condition further includes that the length of the first text does not exceed the M r+1 .

5. The method according to any one of claims 1-4, characterized in that, The performing word segmentation on the first text respectively according to the at least one first word segmentation granularity to obtain at least one first word corresponding to each first word segmentation granularity includes: Performing word segmentation on the first text respectively according to the at least one first word segmentation granularity based on a first word segmentation unit to obtain at least one first word corresponding to each first word segmentation granularity.

6. The method according to claim 5, wherein The method further includes: Providing a file storage interface, where the file storage interface is used to instruct a user to determine a first scenario from multiple scenarios, the multiple scenarios include a file name scenario and a file path scenario, and the multiple scenarios correspond to multiple word segmentation units one by one; Obtaining indication information of the first scenario from the file storage interface; Determining the word segmentation unit corresponding to the first scenario as the first word segmentation unit based on the indication information of the first scenario.

7. The method according to claim 5 or 6, characterized in that, When the first text is the file name of the first file, the first word segmentation unit is one character.

8. The method according to any one of claims 5 to 7, characterized in that, When the first text is the file path of the first file, the first word segmentation unit is one layer of path.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: Providing a parameter setting interface, where the parameter setting interface is used to instruct a user to input the first word segmentation parameter; Obtaining the first word segmentation parameter from the parameter setting interface.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: Obtaining a second text, where the second text is used to search for a second file; Determining at least one second word segmentation granularity based on the first word segmentation parameter and the length of the second text; Performing word segmentation on the second text respectively according to the at least one second word segmentation granularity to obtain at least one second word corresponding to each second word segmentation granularity, and the length of each second word is equal to the corresponding second word segmentation granularity; Obtaining the values of each second index in at least one second index from the index table to obtain at least one file identifier set, where the at least one second index includes the indexes in the index table with the same key as the second word, and the at least one file identifier set corresponds to the at least one second index one by one; Searching for the second file based on the at least one file identifier set.

11. The method according to claim 10, characterized in that Determining at least one second word segmentation granularity based on the first word segmentation parameter and the length of the second text includes: Determine a third word segmentation parameter based on the first word segmentation parameter and the length of the second text, where the third word segmentation parameter satisfies a second condition, and the second condition includes M k not exceeding the length of the second text, where M represents the first word segmentation parameter, k represents the third word segmentation parameter, M is an integer not less than 1, and k is an integer not less than 0; Determining the at least one second word segmentation granularity based on the first word segmentation parameter and the third word segmentation parameter.

12. The method according to claim 11, wherein, The at least one second participle granularity includes the M k .

13. A data management device, characterized in that, The apparatus includes: A first acquisition module, configured to acquire a first text, where the first text includes any one or more of a file name of a first file and a file path of the first file; A first determination module, configured to determine at least one first word segmentation granularity based on a first word segmentation parameter and the length of the first text; A first word segmentation module, configured to perform word segmentation on the first text respectively according to the at least one first word segmentation granularity to obtain at least one first word corresponding to each first word segmentation granularity, and the length of each first word is equal to the corresponding first word segmentation granularity; An index creation module, configured to create at least one first index based on the at least one first word and an identifier of the first file, where the at least one first index is an index of the first file, each first index includes a key-value pair, and the key in the key-value pair included in each first index is one of the first words and the value includes the identifier of the first file; An index storage module, configured to store the at least one first index in an index table.

14. The device according to claim 13, wherein The first determination module includes: The first determination sub-module is configured to determine a second word segmentation parameter based on the first word segmentation parameter and the length of the first text, where the second word segmentation parameter satisfies a first condition, and the first condition includes that M r does not exceed the length of the first text, where M represents the first word segmentation parameter, r represents the second word segmentation parameter, M is an integer not less than 1, and r is an integer not less than 0; A second determination sub-module, configured to determine the at least one first word segmentation granularity based on the first word segmentation parameter and the second word segmentation parameter.

15. The device according to claim 14, wherein The at least one first word segmentation granularity includes M i , where the value range of i includes all integers not less than 0 and not greater than r.

16. The device according to claim 14 or 15, characterized in that, The first condition further includes that the length of the first text does not exceed the M r+1 .

17. The device according to any one of claims 13 to 16, characterized in that The first word segmentation module includes: A first word segmentation sub-module, configured to perform word segmentation on the first text respectively according to the at least one first word segmentation granularity based on a first word segmentation unit to obtain at least one first word corresponding to each first word segmentation granularity.

18. The device according to claim 17, characterized in that, The apparatus further includes: A first providing module, configured to provide a file storage interface, where the file storage interface is used to instruct a user to determine a first scenario from multiple scenarios, the multiple scenarios include a file name scenario and a file path scenario, and the multiple scenarios correspond to multiple word segmentation units one by one; A second acquisition module, configured to acquire indication information of the first scenario from the file storage interface; A second determination module, configured to determine the word segmentation unit corresponding to the first scenario as the first word segmentation unit based on the indication information of the first scenario.

19. The device according to claim 17 or 18, characterized in that, When the first text is the file name of the first file, the first word segmentation unit is one character.

20. The device according to any one of claims 17-19, characterized in that, When the first text is the file path of the first file, the first word segmentation unit is one layer of path.

21. The device according to any one of claims 13-20, characterized in that, The apparatus further includes: A second providing module, configured to provide a parameter setting interface, where the parameter setting interface is used to instruct a user to input the first word segmentation parameter; A third acquisition module, configured to acquire the first word segmentation parameter from the parameter setting interface.

22. The device according to any one of claims 13-21, characterized in that, The apparatus further includes: A fourth acquisition module, configured to acquire a second text, where the second text is used to search for a second file; A third determination module, configured to determine at least one second word segmentation granularity based on the first word segmentation parameter and the length of the second text; A second word segmentation module, configured to perform word segmentation on the second text respectively according to the at least one second word segmentation granularity, to obtain at least one second word corresponding to each second word segmentation granularity, and the length of each second word is equal to the corresponding second word segmentation granularity; A fifth obtaining module, configured to obtain values of each of the at least one second index from the index table, to obtain at least one file identifier set, where the at least one second index includes indexes in the index table whose keys are the same as the second word, and the at least one file identifier set corresponds to the at least one second index one by one; A file search module, configured to search the second file based on the at least one file identifier set.

23. The device according to claim 22, characterized in that, The third determination module includes: A third determination sub-module, configured to determine a third word segmentation parameter based on the first word segmentation parameter and the length of the second text, where the third word segmentation parameter satisfies a second condition, and the second condition includes that M k does not exceed the length of the second text, where M represents the first word segmentation parameter, k represents the third word segmentation parameter, M is an integer not less than 1, and k is an integer not less than 0; A fourth determination sub-module, configured to determine the at least one second word segmentation granularity based on the first word segmentation parameter and the third word segmentation parameter.

24. The device according to claim 23, wherein The at least one second participle granularity includes the M k .

25. A data management system, characterized in that, The system includes a server and a client; The client is configured to provide data to be managed to the server; The server is configured to execute the steps of the method according to any one of claims 1-12.

26. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1-12 are implemented.

27. A computer program product, characterized in that, A computer instruction is stored in the computer program product, and when the computer instruction is executed by a processor, the steps of the method according to any one of claims 1-12 are implemented.

Citation Information

Cited By

  • Full-text retrieval method and device, medium, equipment and computer program product

    CN121434324A