A method, system, device and storage medium for tag query of similar paragraphs in a document
By adopting strategies such as full-word matching, embedding word vectors and LDA theme models based on the length of the mark text, the problem of difficulty in calculating the similarity of shorter text paragraphs in the existing technology is solved, and more accurate and efficient similarity calculation is achieved, improving the user experience.
Patent Information
- Application Number
- CN202110388914.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-04-12
AI Technical Summary
The prior art is difficult to effectively calculate the similarity of shorter text paragraphs, and does not consider the semantic information of the text. Especially when dealing with flexible language environments such as Chinese, similarity results cannot be accurately obtained.
Different similarity calculation strategies are used according to the length of the tagged text. For short text, use full word matching; for medium and long text, calculate similarity by embedding word vectors; for long text, use the LDA theme model to calculate similarity.
By adopting multiple matching strategies for texts of different lengths, the accuracy and effectiveness of similarity calculations are significantly improved, especially in terms of considering semantic information, and the user experience is improved.
Smart Images

Figure CN113139374B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular to a tag query method, system, device and storage medium for similar paragraphs in a document. Background Art
[0002] Nowadays, many companies have a large amount of document text data, including product manuals, business contracts, deployment documents and other highly professional documents. In order to facilitate unified management, many companies will centralize these document data and provide intelligent services such as query, reading, and recommendation. Providing automatic query and matching services for text-similar paragraphs can help users better utilize the text resources in the document library and enhance the value of document resources. The basic function of the automatic query and matching service for text-similar paragraphs is: the user manually marks a paragraph of text when reading. After marking, the system automatically matches paragraphs with similar content to the marked paragraph from all documents in the document library in the background by using NLP and other related technologies and returns them to the user. Users can find paragraphs or texts with similar content as references based on the matching results.
[0003] Most of the existing technologies are solutions similar to text duplication detection. For example, SimHash, the general calculation process is as follows:
[0004] 1. Extract features from documents and their corresponding weights;
[0005] 2. Hash the features and generate the corresponding hash values;
[0006] 3. Hash value weighting: loop through each bit of the feature hash value: if the bit value is 1, replace it with weight, otherwise, replace it with -weight;
[0007] 4. Sum: Sum the weighted results of the feature hash bit by bit, and then binarize the result bit by bit: if it is greater than 0, it is 1, otherwise it is 0, that is, the final SimHash value is obtained.
[0008] After obtaining the SimHash value of the document, the Hamming distance between the SimHash values of the two documents is calculated as the similarity between the two documents.
[0009] However, SimHash itself is an algorithm used by Google to remove duplicates from massive web pages, and is suitable for calculating the similarity of entire documents. However, for shorter text paragraphs, SimHash often cannot achieve good results. In addition, SimHash does not take the semantic information of the text into account. For a language environment such as Chinese, which has a very flexible way of expression, and only involves one or several concepts, rather than large sections of similar content, SimHash cannot obtain accurate similarity results. Summary of the invention
[0010] In view of the technical problems that the above-mentioned prior art cannot perform similarity calculation for shorter text paragraphs and does not consider the semantic information of the text, the present invention proposes a tag query method, system, device and storage medium for similar paragraphs of a document.
[0011] In a first aspect, an embodiment of the present application provides a tag query method for similar paragraphs of a document, comprising:
[0012] Length determination step S1: determining whether the length of the marked text is greater than a first length threshold;
[0013] Query result obtaining step S2: if the length of the marked text is less than the first length threshold, matching the documents in the document library according to the marked text to obtain the query result and output it; or;
[0014] Query result obtaining step S2': if the length of the marked text is greater than the first length threshold, the documents in the document library are segmented into paragraphs, and query results are obtained and outputted through similarity comparison.
[0015] The above-mentioned method for marking similar paragraphs in the document, wherein the query result obtaining step S2 includes: if the length of the marked text is less than the first length threshold, searching for the marked text in all documents in the document library, and taking the sentence where the marked text is located, the position of the sentence in the document and the corresponding document name as the query result and outputting them.
[0016] In the above-mentioned method for marking and querying similar paragraphs of a document, the query result obtaining step S2' comprises:
[0017] Segmentation step S21': segmenting the document into paragraphs according to the length of the marked text to obtain a plurality of segmented text paragraphs;
[0018] Similarity calculation step S22': calculating the similarity between the marked text and the segmented text paragraphs according to the length of the marked text to obtain multiple similarities;
[0019] Similarity comparison step S23': after comparing the multiple similarities with a similarity threshold, the segmented text paragraphs with similarities higher than the similarity threshold, the positions of the segmented text paragraphs in the document and the corresponding document names are taken as query results and output.
[0020] In the above-mentioned method for querying similar paragraphs of documents, the similarity calculation step S22' comprises:
[0021] Medium-length text similarity calculation step S221': if the length of the marked text is greater than the first length threshold and less than the second length threshold, the similarity between the marked text and the segmented text paragraph is obtained by calculating the embedding word vectors of the marked text and the segmented text paragraph; or;
[0022] Long text similarity calculation step S222 ′: if the length of the marked text is greater than the second length threshold, the similarity between the marked text and the segmented text paragraph is obtained through the LAD topic model.
[0023] In a second aspect, an embodiment of the present application provides a tag query system for similar paragraphs of a document, including:
[0024] Length judgment unit: judging whether the length of the marked text is greater than a first length threshold;
[0025] Query result obtaining unit: If the length of the marked text is less than the first length threshold, the query result obtaining unit matches the documents in the document library according to the marked text to obtain the query result and outputs it; if the length of the marked text is greater than the first length threshold, the query result obtaining unit divides the documents in the document library into paragraphs and obtains the query result through similarity comparison and outputs it.
[0026] The above-mentioned marking query system for similar paragraphs in the document, wherein, if the length of the marking text is less than the first length threshold, the query result obtaining unit searches for the marking text in all documents in the document library, and outputs the sentence where the marking text is located, the position of the sentence in the document and the corresponding document name as the query result.
[0027] In the above-mentioned document similar paragraph tag query system, the query result obtaining unit includes:
[0028] Segmentation module: segmenting the document into paragraphs according to the length of the marked text to obtain a plurality of segmented text paragraphs;
[0029] Similarity calculation module: calculates the similarity between the marked text and the segmented text paragraphs according to the length of the marked text to obtain multiple similarities;
[0030] Similarity comparison module: after comparing the multiple similarities with a similarity threshold, the segmented text paragraphs with similarities higher than the similarity threshold, the positions of the segmented text paragraphs in the document and the corresponding document names are taken as query results and output.
[0031] The above-mentioned marking query system for similar paragraphs of the document, wherein, if the length of the marked text is greater than the first length threshold and less than the second length threshold, the similarity calculation module obtains the similarity between the marked text and the segmented text paragraph by calculating the embedding word vectors of the marked text and the segmented text paragraph; if the length of the marked text is greater than the second length threshold, the similarity calculation module obtains the similarity between the marked text and the segmented text paragraph through the LAD topic model.
[0032] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for marking and querying similar paragraphs of a document as described in the first aspect above is implemented.
[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for marking and querying similar paragraphs of a document as described in the first aspect above.
[0034] Compared with the prior art, the advantages and positive effects of the present invention are:
[0035] 1. The present invention divides the marked text into different types of marked text according to their lengths, and adopts different matching strategies for marked texts of different lengths, so that the query results are more accurate;
[0036] 2. The present invention belongs to the field of deep learning technology. For medium and long texts and long texts, semantic information is fully considered when calculating similarity, which greatly improves the matching effect and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of the steps of a method for marking and querying similar paragraphs of a document provided by the present invention;
[0038] Figure 2 The present invention provides Figure 1 Flow chart of step S2';
[0039] Figure 3 The present invention provides Figure 2 Flowchart of step S22';
[0040] Figure 4 A schematic flow chart of an embodiment of a method for tagging and querying similar paragraphs of a document provided by the present invention;
[0041] Figure 5 A framework diagram of a marking query system for similar paragraphs of a document provided by the present invention;
[0042] Figure 6 A framework diagram of a computer device according to an embodiment of the present application.
[0043] Wherein, the accompanying drawings are marked as follows:
[0044] 1. Length judgment unit; 2. Query result acquisition unit; 21. Segmentation module; 22. Similarity calculation module; 23. Similarity comparison module; 81. Processor; 82. Memory; 83. Communication interface; 80. Bus. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0046] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.
[0047] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0048] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0049] The present invention is described in detail below in conjunction with the various embodiments shown in the accompanying drawings, but it should be noted that these embodiments are not limitations of the present invention, and any equivalent transformations or substitutions in functions, methods, or structures made by ordinary technicians in the field based on these embodiments are all within the scope of protection of the present invention.
[0050] Before elaborating on each embodiment of the present invention in detail, the core inventive concept of the present invention is summarized and elaborated in detail through the following several embodiments.
[0051] The present invention divides the paragraphs marked by users into three different situations according to their lengths. Different similarity calculation strategies are adopted for marked texts of different lengths. The whole word matching method is adopted for short texts, the embedding word vector method is adopted for medium and long texts, and the LDA topic model method is adopted for long texts to calculate the similarity. Finally, the paragraphs with similarity higher than the threshold are returned as the results.
[0052] Embodiment 1:
[0053] Figure 1 A schematic diagram of the steps of a tag query method for similar paragraphs of a document provided by the present invention, such as Figure 1 As shown, this embodiment discloses a specific implementation method of a tag query method for similar paragraphs in a document (hereinafter referred to as "method").
[0054] Since the user may mark only a few words or a long text, the algorithm should consider the similarity calculation problem of text paragraphs of different lengths at the same time. Therefore, the present invention proposes an automatic query matching method for text similar paragraphs suitable for enterprise-level document libraries.
[0055] Specifically, the method disclosed in this embodiment mainly includes the following steps:
[0056] Step S1: Determine whether the length of the marked text is greater than a first length threshold; specifically, if the text length is less than the first length threshold, the marked text is regarded as a short text; if it is greater, it is regarded as a long text or a medium-long text.
[0057] Step S2: If the length of the marked text is less than the first length threshold, the documents in the document library are matched according to the marked text to obtain the query result and output it.
[0058] Specifically, the short text adopts a whole-word matching strategy, that is, the marked text is searched in all documents in the document library, and the sentence where the marked text is located, the position of the sentence in the document and the corresponding document name are taken as query results and output.
[0059] Step S2': if the length of the marked text is greater than the first length threshold, the documents in the document library are segmented into paragraphs, and query results are obtained and outputted through similarity comparison.
[0060] Reference Figure 2 As shown, step S2' specifically includes the following contents:
[0061] Step S21': segmenting the document into paragraphs according to the length of the marked text to obtain a plurality of segmented text paragraphs;
[0062] When calculating the similarity between a medium-length text and a long text, the document is first segmented into several segments with similar lengths to the marked text, so that the similarity calculation can be performed with the marked text in sequence. The length of the segmented text paragraphs should be as close as possible to the marked text to make the similarity calculation result more accurate. Paragraph segmentation usually takes a sentence as the smallest unit, and generally the segmentation will not split the sentence. However, if the length of a sentence is much longer than the marked text length, the sentence will be split according to the marked text length.
[0063] Step S22': calculating the similarity between the marked text and the segmented text paragraphs according to the length of the marked text to obtain multiple similarities;
[0064] Reference Figure 3 As shown, step S22' specifically includes the following contents:
[0065] Step S221': if the length of the marked text is greater than the first length threshold and less than the second length threshold, the similarity between the marked text and the segmented text paragraph is obtained by calculating the embedding word vectors of the marked text and the segmented text paragraph;
[0066] Specifically, the marked text whose length is greater than the first length threshold and less than the second length threshold is regarded as medium-length text, and the embedding word vector method is used to calculate the similarity. First, the marked text and the segmented text paragraphs are segmented, stop words are filtered out, and then the pre-trained word vector model is used to obtain the embedding word vector of each word after segmentation. The average value of all the embedding word vectors in the marked text and the segmented text paragraphs is calculated according to the dimension, and the average word vector is obtained as the vector representation of the marked text and the segmented text paragraph, and then the cosine distance of the two vectors is calculated as the similarity between the marked text and the segmented text paragraph. If there are a large number of professional words in the documents in the document library, you can use the document library to build a corpus and use word2vec or glove to retrain the word vector model.
[0067] Step S222': if the length of the marked text is greater than the second length threshold, the similarity between the marked text and the segmented text paragraph is obtained through the LAD topic model.
[0068] Specifically, if the length of the marked text is greater than the second length threshold, it is considered a long text. The embedding word vector method is not effective in calculating the similarity of long texts because the characteristics of the word vector are weakened when the word vector mean is calculated because the text is too long. Therefore, the LDA topic model method is used to calculate the similarity. First, a corpus is constructed through the documents in the document library, and the LDA topic model is trained using the corpus; the topic distribution of the marked text and the segmented text paragraphs is obtained through the trained LDA topic model; the Hellinger distance between the topic distribution of the marked text and the segmented text paragraph is calculated as the similarity between the marked text and the segmented text paragraph.
[0069] Step S23': after comparing the plurality of similarities with a similarity threshold, the segmented text paragraphs with similarities higher than the similarity threshold, the positions of the segmented text paragraphs in the document and the corresponding document names are taken as query results and output.
[0070] Specifically, when the marked text is medium-length text or long text, a similarity threshold needs to be set. Paragraphs below the threshold are considered dissimilar and are directly filtered out. For paragraphs above the threshold, results can be returned as needed based on the similarity from high to low, and the final similar paragraphs, the position of the paragraph in the document, and the corresponding document name are returned as results.
[0071] Please refer to the following Figure 4 . Figure 4 A flowchart of an embodiment of a method for tagging and querying similar paragraphs of a document provided by the present invention, combined with Figure 4 The specific application process of this method is as follows:
[0072] The present invention divides the paragraph length marked by the user into three types: short, medium and long. Different similarity calculation strategies are used for markers of different lengths. The overall process is as follows:
[0073] 1. First, determine the length of the user's marked text. If the length is less than 6 characters, it is considered a short text. Short text uses the strategy of full-word matching, that is, searching for marked text in all documents in the document library. If the marked text appears in a document, the entire sentence at the location where it appears will be returned as the result. The matching method of short text is similar to global search.
[0074] 2. If the length of the user's marked text is greater than 6 characters, similarity calculation is required.
[0075] First, paragraph segmentation is performed. The purpose of paragraph segmentation is to divide the document into several segments with similar lengths to the marked text, so as to calculate the similarity with the marked text in turn. Controlling the length of the segmented text paragraphs to be as close as possible to the marked text paragraphs can make the similarity calculation result more accurate. Paragraph segmentation is based on a sentence as the smallest unit, that is, in general, segmentation will not split the sentence. However, if the length of a sentence is much longer than the length of the marked paragraph, the long sentence will be split according to the length of the marked paragraph. For example, if the length of the marked text is 20 characters, and the length of a sentence in the document is 45 characters, then the sentence will be split into two paragraphs, the former is 20 characters, and the latter is 25 characters.
[0076] 3. After segmentation, select different text similarity calculation methods according to the length of the marked text.
[0077] If the length of the marked text is less than 25 characters, it is considered as medium-long text. The medium-long text uses the embedding word vector method to calculate the similarity. The specific process is as follows: Segment the text paragraph and remove stop words. Then use the pre-trained word vector model to obtain the embedding word vector of each word after segmentation. The word vectors of all words are averaged according to the dimension, and the average word vector obtained is used as the vector representation of the text paragraph. The vector representations of the marked text paragraph and the segmented text paragraph are obtained respectively, and then the cosine distance of the two vectors is calculated as the similarity of the two paragraphs.
[0078] 4. If the length of the marked text is greater than 25 characters, it is considered a long text. The embedding word vector method is not effective in calculating the similarity of long texts because the word vector characteristics are weakened when the mean is calculated due to the length of the text. Therefore, the LDA topic model method is used to calculate the similarity. The documents in the document library are used to construct the corpus and train the LDA topic model. The topic distribution of the marked text paragraph and the segmented text paragraph is obtained through the trained LDA model. The Hellinger distance of the two topic distributions is calculated as the similarity of the two paragraphs.
[0079] 5. Threshold filtering: Set a similarity threshold, and paragraphs below the threshold are considered dissimilar and filtered out directly. For all paragraphs in a document that are above the threshold, the topK with the highest scores can be returned as the final result.
[0080] 6. Return results: The final similar sentences or paragraphs, the position (offset) of the sentences or paragraphs in the document, and the corresponding document name (document ID) are returned as results.
[0081] In the absence of labeled training data, the present invention uses an unsupervised model, fully considers semantic information, and proposes an idea and specific fusion method of integrating multiple matching strategies for different text lengths to improve the matching effect of situations with different expressions but similar contents.
[0082] Embodiment 2:
[0083] In combination with the method for tag query of similar paragraphs in a document disclosed in the first embodiment, this embodiment discloses a specific implementation example of a tag query system for similar paragraphs in a document (hereinafter referred to as the “system”).
[0084] Reference Figure 5 As shown, the system comprises:
[0085] Length judgment unit 1: judging whether the length of the marked text is greater than a first length threshold;
[0086] Query result obtaining unit 2: If the length of the marked text is less than the first length threshold, the query result obtaining unit matches the documents in the document library according to the marked text to obtain the query result and outputs it; if the length of the marked text is greater than the first length threshold, the query result obtaining unit divides the documents in the document library into paragraphs and obtains the query result through similarity comparison and outputs it.
[0087] Specifically, if the length of the marked text is less than the first length threshold, the query result obtaining unit 2 searches for the marked text in all documents in the document library, and outputs the sentence where the marked text is located, the position of the sentence in the document, and the corresponding document name as the query result.
[0088] Specifically, the query result obtaining unit 2 includes:
[0089] Segmentation module 21: segmenting the document into paragraphs according to the length of the marked text to obtain a plurality of segmented text paragraphs;
[0090] Similarity calculation module 22: calculating the similarity between the marked text and the segmented text paragraphs according to the length of the marked text to obtain multiple similarities;
[0091] Similarity comparison module 23: after comparing the plurality of similarities with a similarity threshold, outputs the segmented text paragraphs with similarities higher than the similarity threshold, the positions of the segmented text paragraphs in the document and the corresponding document names as query results.
[0092] Specifically, if the length of the marked text is greater than the first length threshold and less than the second length threshold, the similarity calculation module 22 obtains the similarity between the marked text and the segmented text paragraph by calculating the embedding word vectors of the marked text and the segmented text paragraph; if the length of the marked text is greater than the second length threshold, the similarity calculation module 22 obtains the similarity between the marked text and the segmented text paragraph through the LAD topic model.
[0093] For the technical solutions of the other parts of the document similar paragraph tag query system disclosed in this embodiment and the document similar paragraph tag query method disclosed in the first embodiment, please refer to the description of the first embodiment, which will not be repeated here.
[0094] Embodiment three:
[0095] Combination Figure 6 As shown, this embodiment discloses a specific implementation of a computer device. The computer device may include a processor 81 and a memory 82 storing computer program instructions.
[0096] Specifically, the processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0097] Among them, the memory 82 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 82 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 82 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 82 may be inside or outside the data processing device. In a specific embodiment, the memory 82 is a non-volatile memory. In a specific embodiment, the memory 82 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (Programmable Read-Only Memory, PROM for short), an erasable PROM (Erasable ProgrammableRead-Only Memory, EPROM for short), an electrically erasable PROM (Electrically Erasable ProgrammableRead-Only Memory, EEPROM for short), an electrically alterable ROM (Electrically Alterable Read-Only Memory, EAROM for short) or a flash memory (FLASH) or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0098] The memory 82 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81 .
[0099] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the marking query methods for similar paragraphs of a document in the above embodiments.
[0100] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Figure 6 As shown, the processor 81, the memory 82, and the communication interface 83 are connected via a bus 80 and communicate with each other.
[0101] The communication interface 83 is used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. The communication interface 83 can also implement data communication with other components such as: external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.
[0102] The bus 80 includes hardware, software or both, and couples the components of the computer device to each other. The bus 80 includes but is not limited to at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, bus 80 may include one or more buses. Although embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.
[0103] In addition, in combination with the tag query method for similar paragraphs of a document in the above embodiment, the embodiment of the present application can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the tag query methods for similar paragraphs of a document in the above embodiment is implemented.
[0104] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0105] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A tag query method for similar paragraphs of a document, characterized in that: include: Length determination step S1: determining whether the length of the marked text is greater than a first length threshold; Query result obtaining step S2: if the length of the marked text is less than the first length threshold, matching the documents in the document library according to the marked text to obtain the query result and output it; or; Query result obtaining step S2': if the length of the marked text is greater than the first length threshold, the documents in the document library are segmented into paragraphs, and query results are obtained and outputted through similarity comparison; The query result obtaining step S2' includes: Segmentation step S21': segmenting the document into paragraphs according to the length of the marked text to obtain a plurality of segmented text paragraphs; Similarity calculation step S22': calculating the similarity between the marked text and the segmented text paragraphs according to the length of the marked text to obtain multiple similarities; Similarity comparison step S23': after comparing the plurality of similarities with a similarity threshold, the segmented text paragraphs having the similarity higher than the similarity threshold, the positions of the segmented text paragraphs in the document and the corresponding document names are taken as query results and output; The similarity calculation step S22' comprises: Medium-length text similarity calculation step S221': if the length of the marked text is greater than the first length threshold and less than the second length threshold, the similarity between the marked text and the segmented text paragraph is obtained by calculating the embedding word vectors of the marked text and the segmented text paragraph; or; Long text similarity calculation step S222 ′: if the length of the marked text is greater than the second length threshold, the similarity between the marked text and the segmented text paragraph is obtained through the LAD topic model.
2. A tag query method for similar paragraphs of a document according to claim 1, characterized in that: The query result obtaining step S2 includes: if the length of the marked text is less than the first length threshold, searching for the marked text in all documents in the document library, and taking the sentence where the marked text is located, the position of the sentence in the document and the corresponding document name as the query result and outputting them.
3. A tag query system for similar paragraphs in a document, characterized in that: include: Length judgment unit: judging whether the length of the marked text is greater than a first length threshold; Query result obtaining unit: if the length of the marked text is less than the first length threshold, the query result obtaining unit matches the documents in the document library according to the marked text to obtain the query result and outputs it; if the length of the marked text is greater than the first length threshold, the query result obtaining unit divides the documents in the document library into paragraphs and obtains the query result by similarity comparison and outputs it; Wherein, the query result obtaining unit includes: Segmentation module: segmenting the document into paragraphs according to the length of the marked text to obtain a plurality of segmented text paragraphs; Similarity calculation module: calculates the similarity between the marked text and the segmented text paragraphs according to the length of the marked text to obtain multiple similarities; Similarity comparison module: after comparing the plurality of similarities with a similarity threshold, the segmented text paragraphs with similarities higher than the similarity threshold, the positions of the segmented text paragraphs in the document and the corresponding document names are taken as query results and output; If the length of the marked text is greater than the first length threshold and less than the second length threshold, the similarity calculation module obtains the similarity between the marked text and the segmented text paragraph by calculating the embedding word vectors of the marked text and the segmented text paragraph; if the length of the marked text is greater than the second length threshold, the similarity calculation module obtains the similarity between the marked text and the segmented text paragraph through the LAD topic model.
4. A document similar paragraph tag query system according to claim 3, characterized in that: If the length of the marked text is less than the first length threshold, the query result obtaining unit searches for the marked text in all documents in the document library, and outputs the sentence where the marked text is located, the position of the sentence in the document, and the corresponding document name as the query result.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for tagging and querying similar paragraphs of a document according to any one of claims 1 to 2 is implemented.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the tag query method for similar paragraphs of a document as described in any one of claims 1 to 2 is implemented.
Citation Information
Patent Citations
Method and device for pre-selecting and determining similar documents
CA3151834A1
Science and technology project similarity analysis method, computer equipment and storage medium
CN112199938A