A document sorting method, system, electronic device and storage medium
Through entity recognition and knowledge graph construction, combined with similarity calculation, the problem that existing document sorting methods cannot calculate correlation from a deep semantic perspective is solved, and higher-quality document sorting is achieved, helping users quickly find relevant documents.
Patent Information
- Application Number
- CN202110423451.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-04-20
AI Technical Summary
The existing document sorting methods can only start from shallow features and cannot calculate the correlation between documents and queries from a deep semantic perspective, resulting in the unoptimized sorting results and it is difficult to quickly query the most relevant documents.
Through the entity recognition and extraction step, the entities in the query statement and document are extracted, and relationship recognition is performed to build a knowledge graph. Then, using the similarity calculation step, the similarity between documents and queries is calculated from a deep semantic perspective based on the knowledge graph, and the documents are sorted.
It improves the quality of document sorting, enables users to query the desired document results more easily and quickly, and improves the accuracy of correlation calculation through deep semantic representation.
Smart Images

Figure CN113139383B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information retrieval, and particularly relates to a document sorting method, system, electronic device and storage medium. Background Art
[0002] In the existing document sorting method, some inherent features are first extracted, such as text relevance, content quality classification, content timeliness, click type, user portrait type and other features. Then these features are input into a machine learning model to perform feature combination and weight learning in different ways to obtain a learner. Finally, the learner is used to predict the relevance between the query and the document, so that the documents can be sorted according to the relevance.
[0003] However, the existing method only uses simple and inherent features, and does not find a good way to represent the document, nor does it mine the deep semantic information of the document. As a result, when calculating the relevance between the query and the document, it can only start from the shallow features and cannot calculate from the deep semantic perspective. Therefore, the returned document sorting may not be optimal, and it is impossible for users to conveniently and quickly query the most relevant document results.
[0004] Sorting learning is an important part of the search algorithm. The essence of document search is to first recall a group of documents with high relevance to the keyword query (query) input by the user, and then use the sorting algorithm to accurately sort according to the relevance between the query and each document. A knowledge graph is a graph-based data structure composed of nodes and edges. Each node represents an entity, such as an employee, a product, a company, etc. Each edge is a relationship between entities. Essentially, it is a semantic network that reveals the relationships between entities and can connect all information together. Summary of the Invention
[0005] The embodiments of the present application provide a document sorting method, system, electronic device and storage medium to at least solve the problem that the existing document sorting method can only start from shallow features and cannot calculate from the deep semantic perspective.
[0006] In a first aspect, the embodiments of the present application provide a document sorting method, including: an entity recognition and extraction step of recognizing the entities of a query statement and multiple documents to be sorted, and extracting the query statement entity of the query statement and the document entities of the documents; a document relationship recognition step of recognizing the relationships of the document entities and obtaining the knowledge subgraph of the documents based on the relationships; a similarity calculation step of calculating the similarity between the documents and the query statement according to the query statement entity and the knowledge subgraph; and a document sorting calculation step of sorting the documents according to the similarity between the documents and the query statement.
[0007] Preferably, the similarity calculation step further includes: a first similarity calculation step of calculating the similarity between the query statement and each knowledge subgraph in each document; a second similarity calculation step of calculating the overall similarity between the query statement and each document according to the similarity between the query statement and each knowledge subgraph in each document.
[0008] Preferably, the first similarity calculation step further includes: an entity - to - entity calculation step of calculating the similarity between each query statement entity and each document entity; an entity sub - graph calculation step of calculating the similarity between each query statement entity and the knowledge subgraph according to the similarity between each query statement entity and each document entity; a query statement sub - graph calculation step of calculating the similarity between the query statement and each knowledge subgraph in the document according to the similarity between each query statement entity and the knowledge subgraph.
[0009] Preferably, the entity - to - entity calculation step further includes: calculating the TF - IDF value of the document entity, mapping the query statement entity and the document entity into a query statement entity vector and a document entity vector respectively, and calculating the cosine similarity between the query statement entity and the document entity according to the TF - IDF value, the query statement entity vector and the document entity vector.
[0010] In a second aspect, an embodiment of the present application provides a document sorting system applicable to the above - mentioned document sorting method, including: an entity recognition and extraction module for recognizing the entities of a query statement and multiple documents to be sorted, and extracting the query statement entities of the query statement and the document entities of the documents; a document relationship recognition module for recognizing the relationships of the document entities and obtaining the knowledge subgraphs of the documents based on the relationships; a similarity calculation module for calculating the similarity between the documents and the query statement according to the query statement entities and the knowledge subgraphs; a document sorting calculation module for sorting the documents according to the similarity between the documents and the query statement.
[0011] In a specific implementation, the similarity calculation module further includes: a first similarity calculation module for calculating the similarity between the query statement and each knowledge subgraph in each document; a second similarity calculation module for calculating the overall similarity between the query statement and each document according to the similarity between the query statement and each knowledge subgraph in each document.
[0012] In a specific implementation, the first similarity calculation module further includes: an entity - to - entity calculation unit that calculates the similarity between each query statement entity and each document entity; an entity sub - graph calculation unit that calculates the similarity between each query statement entity and the knowledge sub - graph according to the similarity between each query statement entity and each document entity; and a query statement sub - graph calculation unit that calculates the similarity between the query statement and each knowledge sub - graph in the document according to the similarity between each query statement entity and the knowledge sub - graph.
[0013] In a specific implementation, the entity - to - entity calculation unit further includes: calculating the TF - IDF value of the document entity, mapping the query statement entity and the document entity into a query statement entity vector and a document entity vector respectively, and calculating the cosine similarity between the query statement entity and the document entity according to the TF - IDF value, the query statement entity vector, and the document entity vector.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a document sorting method as described in the first aspect above.
[0015] In a fourth aspect, an embodiment of the present application provides a computer - readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a document sorting method as described in the first aspect above.
[0016] The present invention can be applied to the field of information retrieval technology. Compared with the related technology, a document sorting method provided by an embodiment of the present application uses entity and relationship extraction to obtain a series of sub - graphs to form a knowledge graph, uses the knowledge graph to represent document information from a deep semantic perspective, and finally calculates the similarity between the query and the knowledge - graph representation using tf - idf. This similarity can be directly used as a sorting basis or as one of the features to train a model, thereby improving the quality of document sorting and enabling users to more conveniently and quickly query the desired document results. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0018] Figure 1 is a flowchart of the document sorting method of the present invention;
[0019] Figure 2 is Figure 1 a sub - step flowchart of step S3 in
[0020] Figure 3 is Figure 2 The flowchart of the sub - steps of step S31 in
[0021] Figure 4 The framework diagram of the document sorting system of the present invention;
[0022] Figure 5 The framework diagram of the first similarity calculation module of the present invention;
[0023] Figure 6 The framework diagram of the electronic device of the present invention;
[0024] In the above figures:
[0025] 1. Entity recognition and extraction module; 2. Document relationship recognition module; 3. Similarity calculation module; 4. Document sorting calculation module; 31. First similarity calculation module; 32. Second similarity calculation module; 311. Inter - entity calculation unit; 312. Entity sub - graph calculation unit; 313. Query sentence sub - graph calculation unit; 60. Bus; 61. Processor; 62. Memory; 63. Communication interface. Detailed implementation manners
[0026] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0027] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without making creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood that the content disclosed in the present application is insufficient.
[0028] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment each time, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art will explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.
[0029] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meaning understood by those of ordinary skill in the technical field to which this application belongs. The words "a", "one", "kind", "the" and similar words involved in this application do not indicate a quantity limitation and can mean singular or plural. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices.
[0030] The embodiments of the present invention are described in detail below with reference to the accompanying drawings:
[0031] Figure 1 For the flowchart of the document sorting method of the present invention, please refer to Figure 1 , the document sorting method of the present invention includes the following steps:
[0032] S1: Identify the entities of a query statement and multiple documents to be sorted, and extract the query statement entity of the query statement and the document entities of the documents.
[0033] In specific implementation, extract the entities in the query statement and multiple documents to be sorted, such as products, departments, employees, projects, etc. Optionally, the entity recognition methods used include, but are not limited to, dictionary-based methods and deep learning neural network-based methods.
[0034] In specific implementation, if there is no large amount of labeled data, use the dictionary-based method to sort out and summarize various types of entities to be extracted to form a dictionary, and use dictionary matching to extract entities from the documents; if there is an available labeled data set, use sequence labeling models such as CRF, LSTM+CRF, Bert+CRF, etc. for entity recognition and extraction.
[0035] S2: Identify the relationships of the document entities, and obtain the knowledge subgraph of the documents based on the relationships.
[0036] In a specific implementation, the document entities identified in step S1 are grouped in pairs for relationship identification. Optionally, the relationship identification method used can be to query a pre-defined entity relationship table or a method based on a deep learning neural network.
[0037] In a specific implementation, if there is no large amount of labeled data and the entity types are few, the relationship categories between two entity types can be pre-defined, a "entity - entity - relationship" table can be constructed, and the relationship between two entities can be directly obtained by querying this table; if there is an available labeled data set, a relationship classification model can be trained, with two entity types as inputs and the relationship type as the output for training.
[0038] After the relationship identification of the entities in the document in pairs is completed, multiple knowledge sub-graphs can be obtained (a sub-graph means that any two entities are connected through a relationship, that is, an undirected connected graph). All the knowledge sub-graphs constitute the entire knowledge graph, which can be used to represent the deep semantic information of the document.
[0039] S3: Calculate the similarity between the document and the query statement according to the query statement entity and the knowledge sub-graph.
[0040] Optionally, Figure 2 For Figure 1 the sub-step flowchart of step S3, please further refer to Figure 2 :
[0041] S31: Calculate the similarity between the query statement and each knowledge sub-graph in each document respectively.
[0042] In a specific implementation, to calculate the similarity between the query statement and each sub-graph in each document, multiple entities are identified from the query statement, and then for each entity in the query statement, the similarity with the document sub-graph is calculated respectively.
[0043] Optionally, Figure 3 For Figure 2 the sub-step flowchart of step S31, please further refer to Figure 3 :
[0044] S311: Calculate the similarity between each query statement entity and each document entity respectively; optionally, calculate the TF-IDF value of the document entity, map the query statement entity and the document entity to a query statement entity vector and a document entity vector respectively, and calculate the cosine similarity between the query statement entity and the document entity according to the TF-IDF value, the query statement entity vector and the document entity vector.
[0045] In a specific implementation, the document can obtain a knowledge graph representation using entity recognition - relationship recognition methods. Treating entities as ordinary words, the tf-idf values of all entities in the document can be calculated. The specific calculation formula is as follows:
[0046]
[0047]
[0048] TF-IDF = TF * IDF
[0049] From the above formula, the TF-IDF values of the entities in the document can be calculated. Then, the similarity between the query statement entity and the document entity is calculated using embedding. The pre-trained word vectors can be fine-tuned with the current document library corpus to obtain a word vector table suitable for the current corpus. Then, according to the query statement entity and the document entity, the word vector table is queried, and the entity is mapped to a vector (if the entity word is long, such as "knowledge graph", but there are only "knowledge" and "graph" in the word vector table, then the average of the word vectors mapped respectively can be used as the word vector of the entire entity word). Then, the cosine similarity between the word vectors of the query statement entity and the document entity is calculated. Finally, the following formula can be used to obtain the similarity between the final query statement entity and the document entity. sim(query statement entity, document entity)
[0050] = TFIDF * (1 + cos(emb(query statement entity), emb(document entity)))
[0051] According to the above steps, the similarity between each entity in the query statement and each entity in the document can be calculated and obtained.
[0052] Optionally, the Jaccard similarity algorithm can also be used for similarity calculation; in a specific implementation, the embodiments of the present application can also use any other similarity algorithm that can implement the above similarity calculation.
[0053] S312: According to the similarity between each query statement entity and each document entity, calculate the similarity between each query statement entity and the knowledge subgraph respectively.
[0054] In a specific implementation, after calculating the similarity between each entity in the query statement and each entity in the document, for the similarity between a single entity in the query statement and all entities in the document, the method of taking the average or the maximum value can be used to obtain the similarity between each entity in the query statement and the document subgraph.
[0055] S313: According to the similarity between each query statement entity and the knowledge subgraph, calculate the similarity between the query statement and each knowledge subgraph in the document respectively.
[0056] In a specific implementation, after calculating the similarity between each entity in the query statement and the document subgraph, for the similarities between all entities in the query statement and the document subgraph, the similarity between the query statement and the document subgraph can be obtained by taking the average or the maximum value.
[0057] Please continue to refer to Figure 2 :
[0058] S32: Calculate the overall similarity between the query statement and each document respectively according to the similarity between the query statement and each knowledge subgraph in each document.
[0059] In a specific implementation, after obtaining the similarity between the query statement and each subgraph of the document, finally, for the similarities between the query statement and all subgraphs of the document, by taking the average or the maximum value, the similarity value between the query statement and the document can be obtained.
[0060] Please continue to refer to Figure 1 :
[0061] S4: Sort the documents according to the similarity between the documents and the query statement.
[0062] In a specific implementation, the similarity value can be directly used as the sorting basis, or a sorting model can be trained by combining common inherent features such as text relevance, content quality classification, content timeliness, click type, and user portrait type.
[0063] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0064] The embodiment of the present application provides a document sorting system applicable to the above-mentioned document sorting method. As used below, terms such as "unit" and "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementations in hardware, or a combination of software and hardware are also possible and contemplated.
[0065] Figure 4 For the framework diagram of the document sorting system according to the present invention, please refer to Figure 4 , including:
[0066] Entity recognition and extraction module 1: Recognize the entities in a query statement and a document to be sorted, and extract the query statement entities of the query statement and the document entities of the document.
[0067] In specific implementation, entities in the query statement and the document are extracted, such as products, departments, employees, projects, etc. Optionally, the entity recognition methods adopted include but are not limited to dictionary-based methods and deep learning neural network-based methods.
[0068] In specific implementation, if there is no large amount of labeled data, the dictionary-based method is used to sort out and summarize various types of entities to be extracted to form a dictionary, and the dictionary matching is used to extract entities from the document; if there is an available labeled data set, sequence labeling models such as CRF, LSTM+CRF, Bert+CRF, etc. are used for entity recognition and extraction.
[0069] Document relationship recognition module 2: Recognize the relationships of the document entities and obtain the knowledge subgraph of the document based on the relationships.
[0070] In specific implementation, the document entities recognized in the entity recognition and extraction module 1 are grouped in pairs for relationship recognition. Optionally, the relationship recognition method adopted can be to query the predefined entity relationship table or a deep learning neural network-based method.
[0071] In specific implementation, if there is no large amount of labeled data and the entity types are few, the relationship categories between two entity types can be predefined, a "entity-entity-relationship" table is constructed, and the relationship between two entities can be directly obtained by querying this table; if there is an available labeled data set, a relationship classification model can be trained, with two entity types as the input and the relationship type as the output for training.
[0072] After the relationship recognition of the entities in the document in pairs is completed, multiple knowledge subgraphs can be obtained (a subgraph means that any two entities are connected through a relationship, that is, an undirected connected graph). All the knowledge subgraphs constitute the entire knowledge graph, which can be used to represent the deep semantic information of the document.
[0073] Similarity calculation module 3: Calculate the similarity between the document and the query statement according to the query statement entity and the knowledge subgraph.
[0074] Optionally, the similarity calculation module 3 further includes:
[0075] First similarity calculation module 31: Calculate the similarity between the query statement and each knowledge subgraph in each document respectively.
[0076] In specific implementation, the similarity between the query statement and each subgraph in each document is calculated. Multiple entities are recognized from the query statement, and then for each entity in the query statement, the similarity with the document subgraph is calculated respectively.
[0077] Optionally, Figure 5The framework diagram of the first similarity calculation module of the present invention is shown below. Please refer to Figure 5 :
[0078] Entity - to - entity calculation unit 311: Calculate the similarity between each of the query statement entities and each of the document entities respectively; Optionally, calculate the TF - IDF value of the document entity, map the query statement entity and the document entity to a query statement entity vector and a document entity vector respectively, and calculate the cosine similarity between the query statement entity and the document entity according to the TF - IDF value, the query statement entity vector, and the document entity vector.
[0079] In a specific implementation, the document can use the entity recognition - relationship recognition method to obtain the knowledge graph representation, treat the entity as an ordinary word, and calculate the tf - idf value of all entities in the document. The specific calculation formula is as follows:
[0080]
[0081] TF - IDF = TF * IDF
[0082] From the above formula, the TF - IDF value of the entity in the document can be calculated. Then, use the embedding to calculate the similarity between the query statement entity and the document entity. The pre - trained word vectors can be fine - tuned with the current document library corpus to obtain a word vector table suitable for the current corpus. Then, according to the query statement entity and the document entity, query the word vector table, and map the entity to a vector (if the entity word is long, such as "knowledge graph", but there are only "knowledge" and "graph" in the word vector table, then map them to word vectors respectively and take the average as the word vector of the entire entity word). Then, calculate the cosine similarity between the word vectors of the query statement entity and the document entity. Finally, the similarity between the final query statement entity and the document entity can be obtained using the following formula. sim(query statement entity, document entity)
[0083] = TFIDF * (1 + cos(emb(query statement entity), emb(document entity)))
[0084] According to the above steps, the similarity between each entity in the query statement and each entity in the document can be calculated.
[0085] Optionally, the Jaccard similarity algorithm can also be used for similarity calculation; In a specific implementation, the embodiments of the present application can also use any other similarity algorithm that can implement the above - mentioned similarity calculation.
[0086] Entity sub - graph calculation unit 312: Calculate the similarity between each of the query statement entities and the knowledge sub - graph respectively according to the similarity between each of the query statement entities and each of the document entities.
[0087] In a specific implementation, after calculating the similarity between each entity in the query statement and each entity in the document, for the similarity between a single entity in the query statement and all entities in the document, the average or maximum value can be taken to obtain the similarity between each entity in the query statement and the document sub-graph.
[0088] Query statement sub-graph calculation unit 313: Calculate the similarity between the query statement and each knowledge sub-graph in the document respectively according to the similarity between each query statement entity and the knowledge sub-graph.
[0089] In a specific implementation, after calculating the similarity between each entity in the query statement and the document sub-graph, then for the similarity between all entities in the query statement and the document sub-graph, the average or maximum value can be taken to obtain the similarity between the query statement and the document sub-graph.
[0090] The similarity calculation module 3 further includes a second similarity calculation module 32: Calculate the overall similarity between the query statement and each document respectively according to the similarity between the query statement and each knowledge sub-graph in each document.
[0091] In a specific implementation, after obtaining the similarity between the query statement and each sub-graph of the document, finally for the similarity between the query statement and all sub-graphs of the document, the average or maximum value can be taken to obtain the similarity value between the query statement and the document.
[0092] Please continue to refer to Figure 4 :
[0093] Document sorting calculation module 4: Sort the documents according to the similarity between the document and the query statement.
[0094] In a specific implementation, the similarity value can be directly used as the sorting basis, or a sorting model can be trained by combining common inherent features such as text relevance, content quality classification, content timeliness, click type, user portrait type, etc.
[0095] In addition, a document sorting method described in combination with Figure 1 、 Figure 2 、 Figure 3 can be implemented by an electronic device. Figure 6 This is the framework diagram of the electronic device of the present invention.
[0096] The electronic device may include a processor 61 and a memory 62 storing computer program instructions.
[0097] Specifically, the above-mentioned processor 61 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.
[0098] Among them, the memory 62 may include a mass memory for data or instructions. By way of example and not limitation, the memory 62 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 62 may include removable or non-removable (or fixed) media. Where appropriate, the memory 62 may be internal or external to the data processing device. In a particular embodiment, the memory 62 is a non-volatile memory. In a particular embodiment, the memory 62 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0099] The memory 62 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 61.
[0100] The processor 61 reads and executes the computer program instructions stored in the memory 62 to implement any one of the document sorting methods in the above embodiments.
[0101] In some of the embodiments, the electronic device may further include a communication interface 63 and a bus 60. Among them, as Figure 6 shown, the processor 61, the memory 62, and the communication interface 63 are connected through the bus 60 and complete communication with each other.
[0102] The communication interface 63 can implement data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.
[0103] The bus 60 includes hardware, software, or both, and couples components of the electronic device to each other. The bus 60 includes at least one of the following, without limitation: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, the bus 60 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In a suitable case, the bus 60 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0104] The electronic device can execute a document sorting method in an embodiment of the present application.
[0105] In addition, in combination with a document sorting method in the above embodiments, an embodiment of the present application can be implemented by providing a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the document sorting methods in the above embodiments is implemented.
[0106] The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs.
[0107] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0108] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A document sorting method, characterized in that, it includes: An entity recognition and extraction step of recognizing entities in a query statement and multiple documents to be sorted, and extracting the query statement entity of the query statement and the document entities of the documents; A document relationship recognition step of recognizing relationships of the document entities and obtaining a knowledge subgraph of the documents based on the relationships; A similarity calculation step of calculating the similarity between the document and the query statement according to the query statement entity and the knowledge subgraph; A document sorting calculation step of sorting the documents according to the similarity between the document and the query statement; wherein, the similarity calculation step further includes: A first similarity calculation step of calculating the similarity between the query statement and each knowledge subgraph in each document respectively; A second similarity calculation step of calculating the overall similarity between the query statement and each document respectively according to the similarity between the query statement and each knowledge subgraph in each document; wherein, the first similarity calculation step further includes: An entity - to - entity calculation step of calculating the similarity between each query statement entity and each document entity respectively; An entity subgraph calculation step of calculating the similarity between each query statement entity and the knowledge subgraph respectively according to the similarity between each query statement entity and each document entity; A query statement subgraph calculation step of calculating the similarity between the query statement and each knowledge subgraph in the document respectively according to the similarity between each query statement entity and the knowledge subgraph.
2. The document sorting method according to claim 1, characterized in that, the entity - to - entity calculation step further includes: calculating the TF - IDF value of the document entity, mapping the query statement entity and the document entity into a query statement entity vector and a document entity vector respectively, and calculating the cosine similarity between the query statement entity and the document entity according to the TF - IDF value, the query statement entity vector and the document entity vector.
3. A document sorting system, characterized in that, it includes: An entity recognition and extraction module for recognizing entities in a query statement and multiple documents to be sorted, and extracting the query statement entity of the query statement and the document entities of the documents; A document relationship recognition module for recognizing relationships of the document entities and obtaining a knowledge subgraph of the documents based on the relationships; A similarity calculation module for calculating the similarity between the document and the query statement according to the query statement entity and the knowledge subgraph; A document sorting calculation module for sorting the documents according to the similarity between the document and the query statement; wherein, the similarity calculation module further includes: A first similarity calculation module for calculating the similarity between the query statement and each knowledge subgraph in each document respectively; A second similarity calculation module for calculating the overall similarity between the query statement and each document respectively according to the similarity between the query statement and each knowledge subgraph in each document; The first similarity calculation module further includes: An inter-entity calculation unit that calculates the similarity between each of the query statement entities and each of the document entities respectively; An entity sub-graph calculation unit that calculates the similarity between each of the query statement entities and the knowledge sub-graph respectively according to the similarity between each of the query statement entities and each of the document entities; A query statement sub-graph calculation unit that calculates the similarity between the query statement and each of the knowledge sub-graphs in the document respectively according to the similarity between each of the query statement entities and the knowledge sub-graph; 4. The document ranking system according to claim 3, wherein, the inter-entity calculation unit further includes: calculating the TF-IDF value of the document entity, mapping the query statement entity and the document entity into a query statement entity vector and a document entity vector respectively, and calculating the cosine similarity between the query statement entity and the document entity according to the TF-IDF value, the query statement entity vector and the document entity vector.
5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, it implements the document ranking method according to any one of claims 1 to 2.
6. A computer-readable storage medium, on which a computer program is stored, wherein, when the program is executed by the processor, it implements the document ranking method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Document retrieval method and device and computer readable storage medium
CN112347223A
Maintenance work order-oriented document retrieval method and system, computer and storage medium
CN112417175A