A method, device, equipment and storage medium for text data set management
By storing a brief description of text data sets in the index library and storing metadata in the database, and storing samples in the distributed file system, the problem of long-term retrieval of text data sets in the prior art is solved, and faster retrieval and matching calculation are achieved.
Patent Information
- Application Number
- CN202210026588.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-01-11
AI Technical Summary
The prior art takes a long time to search and process text data sets, mainly because the database is a structured database and the text data set is unstructured data, resulting in inefficient retrieval.
By storing a brief description of the text data set in the index library, storing samples of the data set using a distributed file system, and storing metadata in combination with the database, it realizes rapid retrieval and matching degree calculation, so as to select candidate text data sets that meet the query needs.
It improves the search speed of text data sets, facilitates remote use and maintenance of data sets, and achieves faster data set retrieval compared to the existing technology.
Smart Images

Figure CN114328844B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data, and in particular, to a method, device, equipment and storage medium for managing a text data set. Background Art
[0002] With the application and development of natural language processing (NLP) technology in the field of artificial intelligence, the scale of relevant text data sets is getting larger and larger. Currently, for the storage and management of text data sets, they are usually stored in databases such as MySQL, and a certain operation process is used to implement the management of the data sets.
[0003] However, since the database belongs to a structured database, while the text data set is usually unstructured data, the time consumption is relatively serious when retrieving and processing the text data set. Summary of the Invention
[0004] An object of the present invention is to provide a method, device, equipment and storage medium for managing a text data set, which is proposed for the deficiencies of the above-mentioned existing technologies, and this object is achieved through the following technical solutions.
[0005] A first aspect of the present invention provides a method for managing a text data set, the method comprising:
[0006] Searching for the most relevant preset number of candidate text data sets from an index library according to a query requirement input by a user;
[0007] Reading metadata of each candidate text data set from a database, and determining a matching degree between each candidate text data set and the query requirement by using the metadata;
[0008] Selecting at least one candidate text data set that meets the query requirement from the preset number of candidate text data sets according to the matching degree;
[0009] Reading samples included in the selected candidate text data sets from a distributed file system.
[0010] In some embodiments of the present application, a mapping relationship between a text data set and a brief description word is recorded in the index library; the searching for the most relevant preset number of candidate text data sets from the index library according to a query requirement input by a user includes:
[0011] Extracting a retrieval keyword from the query requirement; determining a first similarity between the retrieval keyword and a brief description word in the index library, sorting the text data sets corresponding to the brief description words in descending order according to the first similarity, and sequentially obtaining a preset number of text data sets from the first text data set in the sorting result as candidate text data sets.
[0012] In some embodiments of the present application, the determining the matching degree between each candidate text dataset and the query requirement by using the metadata includes:
[0013] For each candidate text dataset, determining a second similarity between the metadata of the candidate text dataset and the query requirement; and determining the matching degree by using the second similarity and the first similarity corresponding to the candidate text dataset.
[0014] In some embodiments of the present application, the determining the second similarity between the metadata of the candidate text dataset and the query requirement includes:
[0015] Determining a task type similarity between the description information of the task type field in the metadata and the task type in the query requirement; determining a domain similarity between the description information of the domain field in the metadata and the domain in the query requirement; determining a category similarity between the description information of the category field in the metadata and the category name in the query requirement; and obtaining the second similarity by using the task type similarity, the domain similarity, and the category similarity.
[0016] In some embodiments of the present application, the selecting at least one candidate text dataset that meets the query requirement from a preset number of candidate text datasets according to the matching degree includes:
[0017] Sorting the preset number of candidate text datasets in reverse order according to the matching degree; and selecting at least one candidate text dataset that meets the query requirement according to the sorting result of the candidate text datasets.
[0018] In some embodiments of the present application, the selecting at least one candidate text dataset that meets the query requirement according to the sorting result of the candidate text datasets includes:
[0019] Outputting and displaying the sorting result of the candidate text datasets in a list manner; when receiving a first viewing request from the user for a candidate text dataset, outputting and displaying the metadata of the candidate text dataset selected by the user; when receiving a second viewing request from the user for multiple candidate text datasets, summarizing and outputting and displaying the metadata of the multiple candidate text datasets selected by the user; and when receiving the candidate text dataset selected by the user, using the candidate text dataset selected by the user as the candidate text dataset that meets the query requirement.
[0020] In some embodiments of the present application, the method further includes a storage process of the text dataset:
[0021] Receive an addition request for a text data set; store the samples included in the text data set in the distributed file system; obtain the metadata of the text data set and store the metadata in the database; extract brief description words using the description information in the summary field of the metadata, and store the mapping relationship between the extracted brief description words and the identifier of the text data set in the index library.
[0022] The second aspect of the present invention proposes a text data set management device, which includes:
[0023] A search module, configured to search for the most relevant preset number of candidate text data sets from the index library according to the query requirements input by the user;
[0024] A matching determination module, configured to read the metadata of each candidate text data set from the database and determine the matching degree between each candidate text data set and the query requirements using the metadata;
[0025] A selection module, configured to select at least one candidate text data set that meets the query requirements from the preset number of candidate text data sets according to the matching degree;
[0026] A reading module, configured to read the samples included in the selected candidate text data set from the distributed file system.
[0027] The third aspect of the present invention proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in the first aspect above are implemented.
[0028] The fourth aspect of the present invention proposes a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0029] Based on the text data set management method and device described in the first and second aspects above, the present invention has at least the following beneficial effects or advantages:
[0030] Store the brief description of the text data set in the index library that is convenient for quick retrieval, so that the candidate text data set most relevant to the query requirements input by the user can be quickly retrieved from the index library. Since the metadata has a high query frequency, small occupied space, and high degree of structuring, the metadata used to describe the text data set is stored in a structured database, which is convenient for quickly querying the metadata of the candidate text data set. Further, use the metadata to calculate the matching degree between the candidate text data set and the query requirements, select the candidate text data set that meets the requirements according to the matching degree, and read the samples included in the candidate text data set that meets the requirements from the distributed file system that is convenient for storage and reading.
[0031] As can be seen, by retrieving the candidate text dataset through the index library, querying the database to obtain the metadata of the candidate text dataset for further calculation of the matching degree, and reading the distributed file system to obtain the samples of the final text dataset, compared with using a database to manage the text dataset in the prior art, the dataset retrieval can be made faster and it is convenient to use the dataset remotely.
[0032] The present invention can be applied to the technical field of smart cities, thereby promoting the construction of smart cities. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0034] Figure 1A is a storage flowchart of a text dataset shown according to an exemplary embodiment of the present invention;
[0035] Figure 1B is according to the present invention Figure 1A is a schematic diagram of an input interface shown according to the illustrated embodiment;
[0036] Figure 2A is a flowchart of an embodiment of a text dataset management method shown according to an exemplary embodiment of the present invention;
[0037] Figure 2B is according to the present invention Figure 2A is a schematic diagram of a search interface shown according to the illustrated embodiment;
[0038] Figure 2C is according to the present invention Figure 2A is a schematic diagram of a data export interface shown according to the illustrated embodiment;
[0039] Figure 3 is a schematic structural diagram of a text dataset management device shown according to an exemplary embodiment of the present invention;
[0040] Figure 4 is a schematic hardware structure diagram of an electronic device shown according to an exemplary embodiment of the present invention;
[0041] Figure 5 is a schematic structural diagram of a storage medium shown according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0043] The terms used in the present invention are for the purpose of describing particular embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0044] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0045] Since the scale of text data sets in the NLP field varies greatly, for example, a supervised data set may have only a few hundred samples, while a pre-trained data set requires several gigabytes of storage space. Therefore, in the prior art, based on a database, a certain operation process is used to manage text data sets, which is very time-consuming for users in terms of retrieval and use.
[0046] To solve the above technical problems, the present invention proposes a method for managing text data sets to achieve fast retrieval of text data sets and facilitate remote maintenance and use of text data sets.
[0047] To enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.
[0048] Embodiment 1:
[0049] Figure 1A This is a storage flowchart of a text data set shown according to an exemplary embodiment of the present invention. In this embodiment, the storage architecture of the text data set includes a distributed file system for storing samples, a database for storing metadata of the text data set, and an index library for storing a brief description of the text data set. As Figure 1AAs shown, the storage process of the text dataset includes the following steps:
[0050] Step 101: Receive an addition request for the text dataset.
[0051] In a possible implementation, the user loads the text dataset to be entered into the entry interface and enters the descriptions related to the text dataset according to the input prompts for each field on the entry interface.
[0052] As Figure 1B shown in the entry interface, after the user finishes operating according to the entry interface and clicks the "Enter Data" button, an addition request is triggered.
[0053] From Figure 1B it can be seen that the descriptions that need to be manually entered by the user are: the description of the task type field, the description of the field field, the description of the abstract field, and the descriptions of other field types.
[0054] Step 102: Store the samples included in the text dataset in a distributed file system.
[0055] Among them, the distributed file system can realize the storage and reading of large-scale data. The samples (including text and labels) in the text dataset are the main part of the dataset. Storing all the samples included in the dataset in the distributed file system facilitates subsequent direct and rapid reading.
[0056] Step 103: Obtain the metadata of the text dataset and store the metadata in a database.
[0057] Among them, the metadata used to describe the text dataset is characterized by high structuralization, small occupied space, and high query frequency. Therefore, storing the metadata in a structured database facilitates high-frequency query.
[0058] It should be noted that part of the metadata of the text dataset comes from manual entry by the user, and part comes from system analysis. The data manually entered by the user is carried in the addition request.
[0059] Based on this, in a possible implementation, for the process of obtaining the metadata of the text dataset, by extracting the description information of the task type field, the description information of the field field, and the description information of the abstract field from the addition request, and analyzing the description information of the category field according to the samples included in the text dataset, and then using the description information of the task type field, the description information of the field field, the description information of the abstract field, and the description information of the category field as the metadata of the text dataset.
[0060] Among them, the description information of the category field refers to the category name and category label ID comparison table and the number of categories of the text dataset.
[0061] Those skilled in the art can understand that the above-described metadata fields are only for illustrative purposes, and of course, they may also include description information of other metadata fields. As shown in Table 1 below, the metadata format includes description information of 15 metadata fields.
[0062]
[0063]
[0064] Table 1
[0065] Step 104: Extract brief description words using the description information in the summary field of the metadata, and store the mapping relationship between the extracted brief description words and the identifier of the text dataset in the index library.
[0066] Among them, the summary field is the summary field name shown in Table 1 above. Since the description information of this field belongs to unstructured data, after tokenizing the description information of this field and removing stop words, an inverted index mapping relationship with the tokenization result as the key and the identifier of the text dataset as the value is stored in the index library for fast retrieval.
[0067] So far, the above Figure 1A shown storage process is completed. By storing the main part of the text dataset (i.e., text and labels) in a distributed file system, large-scale data storage can be achieved. By storing the metadata used to describe the text dataset in a database, high-frequency queries are facilitated. By storing the unstructured brief description as structured brief description words in the index library, fast retrieval is convenient.
[0068] Embodiment 2:
[0069] Figure 2A This is a flowchart of an embodiment of a text dataset management method shown according to an exemplary embodiment of the present invention. Based on the above Figure 1A shown embodiment, as Figure 2A shown, the text dataset management method includes the following steps:
[0070] Step 201: Search for the most relevant preset number of candidate text datasets from the index library according to the query requirements input by the user.
[0071] Among them, the user can input various query conditions on the search interface according to the requirements of the task to be processed. Therefore, the query requirements include at least one query condition.
[0072] Such as Figure 2BIn the shown search interface, the user uses a query statement in the form of natural language, as well as information such as task type and field to describe their own needs. After the user clicks the "Search Data" button, a query requirement is triggered.
[0073] In a possible implementation, by extracting retrieval keywords from the query requirement, and determining the first similarity between the retrieval keywords and the brief description words in the index library, arranging the text data sets corresponding to the brief description words in reverse order according to the first similarity, and then sequentially obtaining a preset number of text data sets from the first text data set in the arrangement result as candidate text data sets.
[0074] Among them, the first similarity belongs to the matching degree between the main keywords in the query requirement and the brief description words.
[0075] Optionally, by segmenting the requirement description in the form of natural language in the query requirement and filtering stop words, and then retrieving a certain number of text data sets with the highest first similarity from the index library according to a preset algorithm (such as the BM25 algorithm) as candidate text data sets.
[0076] Step 202: Read the metadata of each candidate text data set from the database, and use the metadata to determine the matching degree between each candidate text data set and the query requirement.
[0077] Among them, the matching degree between the candidate text data set and the query requirement is the comprehensive matching degree between each query condition in the query requirement and the metadata.
[0078] In a possible implementation, for the determination process of the matching degree, for each candidate text data set, by determining the second similarity between the metadata of the candidate text data set and the query requirement, and using the second similarity and the first similarity corresponding to the candidate text data set to determine the matching degree.
[0079] Optionally, the query requirement may include a requirement description of the task type, a requirement description of the field, and a requirement description of the category name, etc.
[0080] In specific implementation, by determining the task type similarity between the description information of the task type field in the metadata and the task type in the query requirement, determining the field similarity between the description information of the field field in the metadata and the field in the query requirement, and determining the category similarity between the description information of the category field in the metadata and the category name in the query requirement, and finally obtaining the second similarity by using the task type similarity, the field similarity, and the category similarity.
[0081] Based on the above-described process, the calculation formula for the matching degree between the candidate text data set and the query requirement is as follows:
[0082] score=score BM25 *score task_type *score domain *score class (Formula 1)
[0083] In the above formula 1, score BM25 Indicates the first similarity, score task_type Indicates the similarity of task types, score domain Indicates domain similarity, score class Indicates category similarity.
[0084] Specifically, the task type similarity and domain similarity can be Jaccard similarity, and the calculation formula for Jaccard similarity is as follows:
[0085]
[0086] In the above formula 2, set 1 represents a description set of query conditions in the query requirement, and set 2 is a description information set of corresponding fields in the metadata.
[0087] The calculation formula for category similarity is as follows:
[0088]
[0089] In the above formula 3, accardSimilarity(label query,n ,label 数据集,m ) is the Jaccard similarity between the n-th category name in the query requirement and the m-th category name in the metadata in terms of character granularity.
[0090] Step 203: selecting at least one candidate text data set that meets the query requirement from a preset number of candidate text data sets according to the matching degree.
[0091] In a possible implementation, a preset number of candidate text data sets are arranged in reverse order according to the matching degree, and then at least one candidate text data set that meets the query requirement is selected according to the arrangement result of the candidate text data sets.
[0092] Among them, the candidate text datasets that are ranked higher are datasets that are highly relevant to the query requirements.
[0093] Considering that the description of the query requirements input by the user and the metadata description of the text dataset are relatively simple, the above ranking of the text datasets is relatively rough. The user may need to conduct a more detailed analysis of the situation of the text dataset to make the best decision.
[0094] Based on this, in specific implementation, the arrangement result of the candidate text dataset can be output and displayed in a list. When receiving a first viewing request from the user for a candidate text dataset, the metadata of the candidate text dataset selected by the user is output and displayed for the user to view its details. When receiving a second viewing request from the user for multiple candidate text datasets, the metadata of the multiple candidate text datasets selected by the user is summarized and then output and displayed. When receiving the candidate text dataset selected by the user, the candidate text dataset selected by the user is used as the candidate text dataset that meets the query requirements.
[0095] Among them, the summarization of multiple candidate text datasets can be in the form of union summarization. Based on the metadata format given in Table 1 above, as shown in Table 2, it is the summarization method of the description information of each field in the metadata.
[0096] Field Name Summary Method task_type Union of the task type lists of each dataset domain Union of the domain lists of each dataset word_cloud Union of the word cloud words of each dataset; Sum the weights of the same words. mean_text_len Calculate the weighted mean according to the corpus scale. max_text_len Select the maximum value. size Sum. detail_metrcis Union.
[0097] Table 2
[0098] Step 204: Read the samples included in the selected candidate text dataset from the distributed file system.
[0099] Among them, the samples in the selected candidate text dataset can be exported from the distributed system to a specified path. As Figure 2C shown, through this export interface, the selected candidate text dataset can be downloaded to a local specified path or to a remote specified path.
[0100] So far, the above-mentioned Figure 2A shown management process is completed. The brief description of the text dataset is stored in an index library that is convenient for rapid retrieval, so that the candidate text dataset most relevant to the query requirements input by the user can be quickly retrieved from the index library. Since the metadata has a high query frequency, occupies little space, and has a high degree of structuring, the metadata used to describe the text dataset is stored in a structured database, which is convenient for quickly querying the metadata of the candidate text dataset. Further, the matching degree between the candidate text dataset and the query requirements is calculated using the metadata, and the candidate text dataset that meets the requirements is selected according to the matching degree, and the samples included in the candidate text dataset that meets the requirements are read from the distributed file system that is convenient for storage and reading.
[0101] It can be seen that by retrieving the candidate text dataset through the index library, querying the database to obtain the metadata of the candidate text dataset for further calculation of the matching degree, and reading the distributed file system to obtain the samples of the final text dataset, compared with using a database to manage text datasets in the prior art, it can make the dataset retrieval faster and facilitate the remote use of the dataset.
[0102] Corresponding to the embodiments of the foregoing text dataset management method, the present invention also provides embodiments of a text dataset management device.
[0103] Figure 3 FIG. is a schematic structural diagram of a text dataset management device according to an exemplary embodiment of the present invention. The device is used to execute the text dataset management method provided in any of the foregoing embodiments, such as Figure 3 shown, the text dataset management device includes:
[0104] A search module 310, configured to search for a preset number of most relevant candidate text datasets from an index library according to a query requirement input by a user;
[0105] A matching determination module 320, configured to read metadata of each candidate text dataset from a database, and use the metadata to determine a matching degree between each candidate text dataset and the query requirement;
[0106] A selection module 330, configured to select at least one candidate text dataset that meets the query requirement from a preset number of candidate text datasets according to the matching degree;
[0107] A reading module 340, configured to read samples included in the selected candidate text dataset from a distributed file system.
[0108] For the functions and actions of each unit in the above device, the specific implementation processes are detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0109] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0110] The embodiments of the present invention also provide an electronic device corresponding to the text dataset management method provided in the foregoing embodiments to execute the above text dataset management method.
[0111] Figure 4The following is a hardware structure diagram of an electronic device according to an exemplary embodiment of the present invention. The electronic device includes: a communication interface 601, a processor 602, a memory 603, and a bus 604. Among them, the communication interface 601, the processor 602, and the memory 603 complete mutual communication through the bus 604. The processor 602 can execute the text dataset management method described above by reading and executing machine-executable instructions corresponding to the control logic of the text dataset management method in the memory 603. For the specific content of this method, please refer to the above embodiments and will not be repeated here.
[0112] The memory 603 mentioned in the present invention can be any electronic, magnetic, optical, or other physical storage device, and can store information such as executable instructions, data, etc. Specifically, the memory 603 can be RAM (Random Access Memory), flash memory, a storage drive (such as a hard disk drive), any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 601 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0113] The bus 604 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 603 is used to store a program, and the processor 602 executes the program after receiving an execution instruction.
[0114] The processor 602 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 602 or instructions in the form of software. The above-mentioned processor 602 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor.
[0115] The electronic device provided by the embodiment of the present application and the text data set management method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by them.
[0116] The embodiment of the present application also provides a computer-readable storage medium corresponding to the text data set management method provided by the foregoing embodiment. Please refer to Figure 5 As shown, the computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the text data set management method provided by any of the foregoing embodiments.
[0117] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical or magnetic storage media, which will not be elaborated here one by one.
[0118] The computer-readable storage medium provided by the above embodiment of the present application and the text data set management method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0119] Those skilled in the art will readily think of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0120] It should also be noted that the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, commodity, or device comprising the element.
[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for managing a text data set, characterized in that, The method includes: Searching for the most relevant preset number of candidate text data sets from the index library according to the query requirements input by the user; Reading the metadata of each candidate text data set from the database, and using the metadata to determine the matching degree between each candidate text data set and the query requirements; Selecting at least one candidate text data set that meets the query requirements from the preset number of candidate text data sets according to the matching degree; Reading the samples included in the selected candidate text data sets from the distributed file system; The determining the matching degree between each candidate text data set and the query requirements by using the metadata includes: For each candidate text data set, determining a second similarity between the metadata of the candidate text data set and the query requirements; using the second similarity and the first similarity corresponding to the candidate text data set to determine the matching degree, where the first similarity is the matching degree between the main keywords in the query requirements and the brief description words in the index library; The selecting at least one candidate text data set that meets the query requirements from the preset number of candidate text data sets according to the matching degree includes: Sorting the preset number of candidate text data sets in reverse order according to the matching degree, and outputting and displaying the sorting result of the candidate text data sets in a list manner; when receiving a first viewing request from the user for a candidate text data set, outputting and displaying the metadata of the candidate text data set selected by the user; when receiving a second viewing request from the user for multiple candidate text data sets, summarizing and outputting the metadata of the multiple candidate text data sets selected by the user; when receiving the candidate text data set selected by the user, using the candidate text data set selected by the user as the candidate text data set that meets the query requirements.
2. The method according to claim 1, characterized in that, The mapping relationship between the text data set and the brief description words is recorded in the index library; The searching for the most relevant preset number of candidate text data sets from the index library according to the query requirements input by the user includes: Extracting retrieval keywords from the query requirements; Determining a first similarity between the retrieval keywords and the brief description words in the index library; Sorting the text data sets corresponding to the brief description words in reverse order according to the first similarity, and sequentially obtaining a preset number of text data sets from the first text data set in the sorting result as candidate text data sets.
3. The method according to claim 1, characterized in that, The determining the second similarity between the metadata of the candidate text data set and the query requirements includes: Determining a task type similarity between the description information of the task type field in the metadata and the task type in the query requirements; Determining a domain similarity between the description information of the domain field in the metadata and the domain in the query requirements; Determining a category similarity between the description information of the category field in the metadata and the category name in the query requirements; Obtaining the second similarity by using the task type similarity, the domain similarity, and the category similarity.
4. The method according to any one of claims 1-3, characterized in that, The method further includes a storage process for text data sets: Receiving an addition request for a text data set; Store the samples included in the text data set in the distributed file system; Obtain the metadata of the text data set and store the metadata in the database; Extract brief description words using the description information in the abstract field of the metadata, and store the mapping relationship between the extracted brief description words and the identifier of the text data set in the index library.
5. A text data set management device, characterized in that, The device includes: A search module for searching the most relevant preset number of candidate text data sets from the index library according to the query requirements input by the user; A matching determination module for reading the metadata of each candidate text data set from the database and determining the matching degree between each candidate text data set and the query requirements using the metadata; A selection module for selecting at least one candidate text data set that meets the query requirements from the preset number of candidate text data sets according to the matching degree; A reading module for reading the samples included in the selected candidate text data set from the distributed file system; The matching determination module is specifically configured to, in the process of determining the matching degree between each candidate text data set and the query requirements using the metadata, for each candidate text data set, determine the second similarity between the metadata of the candidate text data set and the query requirements; determine the matching degree using the second similarity and the first similarity corresponding to the candidate text data set, where the first similarity is the matching degree between the main keywords in the query requirements and the brief description words in the index library; The selection module is specifically configured to sort the preset number of candidate text data sets in reverse order according to the matching degree, and output and display the sorting result of the candidate text data sets in a list; when receiving a first viewing request from the user for a candidate text data set, output and display the metadata of the candidate text data set selected by the user; when receiving a second viewing request from the user for multiple candidate text data sets, summarize and output the metadata of the multiple candidate text data sets selected by the user; when receiving the candidate text data set selected by the user, use the candidate text data set selected by the user as the candidate text data set that meets the query requirements.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
Full-text retrieval system based on two-level semantic analysis
CN103136352A
Data set retrieval method and system
CN111026710A