Text comparison method and device, storage medium and electronic equipment
By segmenting paragraphs and text fragments, identifying feature word segments and matching similar texts on patent application texts, the problem of low accuracy in the recognition of abnormal patent application texts in the prior art is solved, and more accurate text comparison and identification of abnormal patent application texts are achieved.
Patent Information
- Application Number
- CN202510549015.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-30
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
The lack of special text comparison technology in the prior art has led to low accuracy in the identification of abnormal patent application texts and is unable to effectively identify patent application texts with significantly the same content of invention or simple combinations.
By segmenting paragraphs and text fragments for the detected text, identifying feature word segments, constructing a search formula and seizing similar text fragments from similar texts, determining the similarity of paragraph text based on similarity, and finally obtaining the comparison result.
It realizes a more accurate comparison of patent application texts, can identify substantially the same technical content, improves the accuracy of identification of abnormal patent application texts, and saves the energy of manual review.
Smart Images

Figure CN120448521A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on April 30, 2024, with application number 202410541834.7 and invention name “Text comparison method, device, storage medium and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the technical field of text comparison, and in particular to a text comparison method, device, storage medium, and electronic device. Background Art
[0003] With the development of computer technology, a variety of texts have emerged. In the field of patent applications, the text of a patent application at the time of submission is called a patent application text. This text will describe the content of the invention in detail and make a formal request for patent rights. When applying for a patent, for various reasons, an applicant may submit multiple patent application texts simultaneously or successively, which are obviously identical in content or are essentially composed of simple combinations and variations of different inventive features or elements. However, such patent application texts are considered abnormal patent application texts, and therefore need to be identified.
[0004] In the existing technology, there is no specialized text comparison technology to identify these types of anomalous patent application texts. Therefore, character matching technology, which is used in the field of academic papers, is used. However, this technology only detects and compares the direct correspondence between character sequences in the text, focusing primarily on the surface form of the text without deeply analyzing the underlying semantics. As a result, the accuracy of identifying anomalous patent application texts using this technology for text comparison is low.
[0005] Therefore, providing a text comparison technology with high accuracy for identifying abnormal patent application texts is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0006] Based on this, the purpose of this application is to provide a text comparison method that can achieve more accurate comparison of patent application texts.
[0007] The present invention discloses a text comparison method, comprising the following steps: Obtaining the text to be detected and the innovation subject information of the text to be detected; Segmenting the text to be detected into paragraphs to obtain a plurality of paragraph texts; Segmenting each paragraph text into text segments to obtain a plurality of text segments corresponding to each paragraph text; Performing feature segmentation recognition on each of the text segments to obtain a feature segmentation set corresponding to each of the text segments; Based on a preset search formula construction template, several characteristic segmentation words in each characteristic segmentation word set and the innovation subject information, a search formula is constructed, and based on the search formula, several similar texts corresponding to each of the text fragments are obtained; According to the feature segmentation set of each text segment, a number of similar text segments matching the text segment are intercepted from the corresponding similar text; wherein the similar text segments include a preset number of feature segmentations in the feature segmentation set; Determining similar texts for each of the paragraph texts based on similarities between each of the text segments and the corresponding similar text segments; A comparison result of the text to be detected is obtained based on the similar text of each paragraph text.
[0008] The present application also discloses a text comparison device, comprising: A module for acquiring text to be detected, used to acquire the text to be detected and the innovation subject information of the text to be detected; A paragraph segmentation module for the text to be detected is used to segment the text to be detected into paragraphs to obtain a plurality of paragraph texts; A paragraph text segmentation module is used to segment each paragraph text into text segments to obtain a plurality of text segments corresponding to each paragraph text; A text segmentation feature recognition module is used to perform feature segmentation recognition on each text segment to obtain a feature segmentation set corresponding to each text segment; A text segment similar text matching module is used to construct a search formula based on a preset search formula construction template, a number of characteristic segmentation words in each characteristic segmentation word set, and the innovation subject information, and obtain a number of similar texts corresponding to each text segment based on the search formula; A similar text similar segment interception module is used to intercept a number of similar text segments that match the text segment from the corresponding similar text according to the feature segmentation set of each text segment; wherein the similar text segment includes a preset number of feature segmentations in the feature segmentation set; A paragraph similar text determination module, configured to determine similar texts of each paragraph text based on similarities between each text segment and corresponding similar text segments; The comparison result generating module is used to obtain the comparison result of the text to be detected based on the similar text of each paragraph text.
[0009] An embodiment of the present application further discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the computer program, when executed by the processor, implements the text comparison method as described in any one of the embodiments of the present application.
[0010] An embodiment of the present application further discloses a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the text comparison method as described in any one of the embodiments of the present application.
[0011] The embodiment of the present application first divides the text to be detected into several paragraphs, and further divides each paragraph into several text fragments. Then, feature segmentation is performed on each text fragment to obtain a corresponding feature segmentation set. A search formula is constructed using these feature segmentations, innovation subject information, and a preset search formula construction template to retrieve several similar texts that match each text fragment. Then, based on the feature segmentation set of each text fragment, similar text fragments that match the text fragment are intercepted from the corresponding similar text. Finally, based on the similarity between each text fragment and its corresponding similar text fragment, the similar text of each paragraph text is determined, thereby obtaining a comparison result for the text to be detected. Through the above method, the embodiment of the present application can effectively identify and compare substantially identical technical content and can more accurately identify abnormal patent application texts. In addition, considering the textual expression characteristics of technical content in the patent field, the embodiment of the present application accurately determines the similar text of the paragraph by identifying the feature segmentations of the text fragments and retrieving similar texts and similar text fragments matched therein through the feature segmentations, making the comparison result more accurate.
[0012] For better understanding and implementation, the present application is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A flowchart of a text comparison method according to an embodiment of the present application is shown; Figure 2 Schematic diagram of the steps for calculating the similarity of text segments in an embodiment of the present application; Figure 3 Schematic diagram of the steps including branching steps in the text segment similarity calculation according to an embodiment of the present application; Figure 4 A schematic diagram of a text comparison device according to an embodiment of the present application; Figure 5 This is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0014] To make the objectives, technical solutions, and advantages of this application more clear, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.
[0015] It should be understood that the embodiments described in the following examples do not represent all embodiments consistent with this application. Rather, they are merely examples of devices and methods consistent with certain aspects of this application, as detailed in the appended claims. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this application without inventive effort are intended to fall within the scope of protection of this application.
[0016] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "the" and "the" used in this application are also intended to include plural forms, unless the context clearly indicates otherwise. In addition, in the description of this application, unless otherwise stated, "a plurality" refers to two or more. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone; the character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0017] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, this information should not be limited to these terms. Moreover, these terms are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood to indicate or imply relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances. Depending on the context, the words "if" / "if" used in this application can be interpreted as "at the time of" or "when" or "in response to determining".
[0018] With the development of computer technology, a variety of texts have emerged. In the field of patent applications, the text of a patent application at the time of submission is called a patent application text. This text will describe the content of the invention in detail and make a formal request for patent rights. When applying for a patent, for various reasons, an applicant may submit multiple patent application texts simultaneously or successively, which are obviously identical in content or are essentially composed of simple combinations and variations of different inventive features or elements. However, such patent application texts are considered abnormal patent application texts, and therefore need to be identified.
[0019] In the existing technology, there is no specialized text comparison technology to identify these types of anomalous patent application texts. Therefore, character matching technology, which is used in the field of academic papers, is used. However, this technology only detects and compares the direct correspondence between character sequences in the text, focusing primarily on the surface form of the text without deeply analyzing the underlying semantics. As a result, the accuracy of identifying anomalous patent application texts using this technology for text comparison is low.
[0020] This application aims to provide a text comparison method that can achieve more accurate comparison of patent application texts.
[0021] Please refer to Figure 1 The first aspect of the embodiment of the present application discloses a text comparison method, comprising the following steps: S101: Acquire a text to be detected and information about the creative subject of the text to be detected; S102: Segment the text to be detected into paragraphs to obtain a plurality of paragraphs of text; S103: Segment each paragraph text into text segments to obtain a plurality of text segments corresponding to each paragraph text; S104: performing feature segmentation recognition on each of the text segments to obtain a feature segmentation set corresponding to each of the text segments; S105: constructing a search formula based on a preset search formula construction template, a number of characteristic segmentation words in each characteristic segmentation word set, and the innovation subject information, and obtaining a number of similar texts corresponding to each of the text segments based on the search formula; S106: Based on the feature segmentation set of each text segment, extract a number of similar text segments that match the text segment from the corresponding similar text; wherein the similar text segments include a preset number of feature segmentations in the feature segmentation set; S107: Determining similar texts of each paragraph text according to similarities between each text segment and corresponding similar text segments; S108: Obtaining a comparison result of the text to be detected based on similar texts of each paragraph text.
[0022] The embodiment of the present application first divides the text to be detected into several paragraphs, and further divides each paragraph into several text fragments. Then, feature segmentation is performed on each text fragment to obtain a corresponding feature segmentation set. A search formula is constructed using these feature segmentations, innovation subject information, and a preset search formula construction template to retrieve several similar texts that match each text fragment. Then, based on the feature segmentation set of each text fragment, similar text fragments that match the text fragment are intercepted from the corresponding similar text. Finally, based on the similarity between each text fragment and its corresponding similar text fragment, the similar text of each paragraph text is determined, thereby obtaining a comparison result for the text to be detected. Through the above method, the embodiment of the present application can effectively identify and compare substantially identical technical content and can more accurately identify abnormal patent application texts. In addition, considering the textual expression characteristics of technical content in the patent field, the embodiment of the present application accurately determines the similar text of the paragraph by identifying the feature segmentations of the text fragments and retrieving similar texts and similar text fragments matched therein through the feature segmentations, making the comparison result more accurate.
[0023] The text comparison method of the present embodiment is primarily used in the context of comparing patent application texts. Unlike general text comparison scenarios, the comparison of patent application texts focuses on the technical content of the technical solutions. The purpose of the comparison is to identify patent application texts that are significantly identical in content or that are essentially formed by simple combinations and variations of different inventive features or elements. The comparison results of the present embodiment can be used to accurately determine whether such abnormal patent application texts exist, significantly reducing the effort required to manually review abnormal patent applications.
[0024] The text comparison method of the embodiments of the present application can be executed by a text comparison device. The text comparison device can be implemented through software and / or hardware. The text comparison device can be composed of two or more physical entities or a single physical entity. The hardware referred to by the text comparison device is essentially a computer device. For example, the text comparison device can be a computer, a mobile phone, a tablet, or an intelligent device such as an interactive tablet.
[0025] In step S101, the text to be detected and the innovation subject information of the text to be detected are obtained.
[0026] The text to be detected includes, but is not limited to, a patent application to be submitted. In some embodiments, it may also be a paper to be submitted or a self-media text to be published. The text to be detected may be encrypted or unencrypted.
[0027] In the embodiment of the present application, the text to be detected is an encrypted patent application text to be submitted. The patent application text can be a patent application text in various languages, and the present application does not limit it.
[0028] When receiving the encrypted patent application text to be submitted, the text comparison device can receive all the encrypted patent application texts of the user at one time; the text comparison device can also receive several encrypted segment files uploaded by the user based on the encrypted patent application text, decrypt the several encrypted segment files, obtain several decrypted segment files, perform document parsing and document merging on the several decrypted segment files, and obtain the final text to be detected.
[0029] When comparing patent application texts for abnormal patent application behavior in the embodiment of the present application, the comparison is mainly aimed at patent application documents submitted simultaneously and successively by the applicant and its related parties. Therefore, the text range in the comparison database should also be limited to the patent application texts submitted on the same day and previously by the applicant and its related parties.
[0030] Specifically, in the patent comparison scenario, the innovation subject information includes at least the information of the applicant who submitted the document to be detected, and may also include information of other applicants associated with it. Generally speaking, when making an abnormality judgment for a new application document submitted by an applicant (individual or unit), it is necessary to combine the unit information associated with the inventor, the inventor information associated with the unit (generally the technical personnel of the unit) and / or other unit information associated with the unit (for example, a business unit may have multiple branches in multiple regions) to search for relevant patent application documents to determine whether the submitted inventions are obviously the same, or are essentially formed by a simple combination of different invention features or elements.
[0031] In one embodiment, the innovation subject information includes information of a first applicant who submits the text to be detected and information of a second applicant associated with the first applicant information.
[0032] The first applicant subject information is the applicant subject information that submits the text to be detected. In a patent application, the applicant subject information includes applicant information and / or inventor information, wherein the applicant information can be personal information when submitting a patent application in the name of an individual, or it can be unit information such as company information when submitting a patent application in the name of an organization.
[0033] The second applicant entity information is information about other related applicant entities retrieved based on the first applicant entity information. In one embodiment, the inventor's organization information can be retrieved based on the inventor information, and information about other related organizations (e.g., branches) can also be retrieved based on the organization information. Furthermore, information about other inventors associated with the organization can also be retrieved based on the organization information.
[0034] In one embodiment, the step of obtaining the text to be detected and the creative subject information of the text to be detected in step S101 includes: Step S1011, obtaining information of the first applicant who submitted the text to be detected; Step S1012: obtaining second applicant information associated with the first applicant information based on the first applicant information; Step S1013: Obtain the innovation subject information according to the first applicant subject information and the second applicant subject information.
[0035] In this embodiment, the second applicant information associated with the first applicant information can be obtained from a user information database. The user information database is a database that can be used to query other applicant information associated with the first applicant information. In other words, it can query the associated unit information based on the inventor information, or query the associated inventor and other associated unit information based on the unit information. Specifically, it can include an enterprise organization information database and / or a patent database.
[0036] In step S102, the text to be detected is segmented into paragraphs to obtain a plurality of paragraphs of text.
[0037] In one embodiment, the text to be detected can be segmented according to natural paragraphs, and the segmented text can be processed to remove stop words and screen words to obtain a plurality of paragraph texts. Specifically, the paragraph separators in the patent application to be submitted can be identified, and the patent application to be submitted can be segmented into a plurality of paragraph texts according to the paragraph separators. Of course, in other embodiments, other methods can be used for paragraph segmentation processing. Specifically, the text to be detected can be segmented according to user-defined paragraph segmentation rules or a paragraph segmentation model obtained through neural network training.
[0038] In step S103, each paragraph text is segmented into text segments to obtain a plurality of text segments corresponding to each paragraph text.
[0039] The embodiment of the present application is to determine similar texts of each paragraph text of the text to be detected, and further segment each paragraph text into text fragments in this step, so that the subsequent steps can match the similar texts of each text fragment separately through each text fragment.
[0040] In this embodiment, the text segment can be a text segment that constitutes a complete sentence. Specifically, when segmenting a paragraph text, segmentation can be performed based on punctuation marks in the paragraph text. The specific segmentation rules can be set according to the situation. For example, the text segment can be segmented based on sentence ending symbols such as periods and exclamation marks. Of course, in some embodiments, other methods can be used to segment text segments, such as using a pre-trained text segmentation model.
[0041] In step S104, feature segmentation recognition is performed on each of the text segments to obtain a feature segmentation set corresponding to each of the text segments.
[0042] In the embodiments of the present application, feature segmentation is essentially the feature element of the technical solution. Those skilled in the art understand that the technical solution is composed of technical means, and the technical means are inseparable from the feature elements. For example, "obtaining user data sent by the client device" contains the feature elements of "client device" and "user data". When performing feature segmentation recognition, the recognized feature segmentation is the feature element of the technical solution.
[0043] Among them, feature word segmentation recognition refers to the recognition of feature word segmentations in text fragments, that is, entity nouns. For example, the text fragment "Beijing is the capital of China" can obtain entity nouns such as "Beijing", "China", and "Capital" through word segmentation recognition. For example, the text fragment "Beijing, Shanghai, and Guangzhou" can obtain entity nouns such as "Beijing", "Shanghai", and "Guangzhou" through word segmentation recognition. Specifically, in this step, a customized ES word segmenter can be used to perform feature word segmentation recognition on the text fragment. In order to more effectively identify the feature word segmentations that constitute the technical solution, in one embodiment, a feature word segmentation recognition model trained with a technical feature word segmentation database is used for word segmentation recognition, thereby ensuring that the feature word segmentations that constitute the technical solution are recognized. The technical feature word segmentation database can be obtained based on the analysis of patent document data, which stores the feature word segmentations used in various technical solutions. Therefore, when the text fragment is segmented by the feature word segmentation recognition model, the feature word segmentations will not be recognized for the cliché text in the patent application text or other text content that does not record the actual technical content. Therefore, the accuracy of searching for similar texts in the paragraph text can be avoided from being interfered by non-technical content.
[0044] Correspondingly, the feature segmentation set includes all feature segmentations obtained from the text segmentation recognition.
[0045] Furthermore, in one embodiment, the feature segmentation obtained by the segmentation recognition can also be expanded with equivalent feature segmentation, so as to obtain a more complete feature segmentation set. It should be understood that in the text comparison method of the embodiment of the present application, the determination of similar texts and similar text segments based on the feature segmentation set of the text fragment is based on the original feature segmentation obtained by the segmentation recognition, that is, the number of feature segmentations in the feature segmentation set is determined by the number of original feature segmentations obtained by the text segmentation recognition, and the feature segmentations selected from the feature segmentation set when performing the search refer to the selection from the original feature segmentations. If, in one embodiment, the feature segmentation set is also expanded with equivalent feature segmentations, without changing the technical concept of the present application, it can be understood that the equivalent feature segmentations are mainly used to expand the search of the original feature segmentations when searching similar texts (see subsequent description), and when performing word recognition on similar text segments, if equivalent feature segmentations are identified instead of directly identifying the original feature segmentations, in this case, the identified equivalent feature segmentations can be equivalently identified as feature segmentations, and the specific technical implementation can be designed based on this concept, and is not further limited here.
[0046] Specifically in the field of patents, the same entity object in a technical means can generally have multiple different names, corresponding to different entity nouns (that is, different feature segmentations). Therefore, in one embodiment, the original feature segmentations obtained by the text segmentation recognition can be expanded based on the pre-trained feature segmentation expansion model or based on the feature segmentation library to obtain other equivalent feature segmentations. For example, for "equipment", the equivalent feature segmentation that can be replaced may be identified as "device". For example, for "fixed component", the equivalent feature segmentation that can be replaced may be identified as "fixed component", "fixed element", etc., and so on. Based on the feature segmentations obtained by the original recognition and the other expanded equivalent feature segmentations, the feature segmentation set corresponding to the text segment is obtained, so that the comparison of the patent application text is more accurate. Based on this principle, specifically, a feature segmentation expansion model for expanding equivalent feature segmentations can be obtained through neural network training, or a feature segmentation library storing various equivalent feature segmentations can be established after analyzing existing patent documents.
[0047] In step S105, a search formula is constructed based on a preset search formula construction template, several feature segmentations in each feature segmentation set, and the innovation subject information, and several similar texts corresponding to each text segment are obtained based on the search formula.
[0048] In the embodiment of the present application, after each paragraph text is divided into several text segments, similar texts corresponding to each text segment are retrieved respectively based on the feature segmentation set of each text segment.
[0049] Please refer to Figure 2 In one embodiment, the step of calculating the similarity between each of the text segments and the corresponding similar text segments in step S1071 includes: Step S10711: Select a text segment from each of the text segments as a first text segment, select a similar text segment from several similar text segments corresponding to the first text segment as a first similar text segment, use the first character string of the first text segment as a current character string, compare the current character string of the first text segment with each character string of the first similar text segment, and if there is a character string in the first similar text segment that is consistent with the current character string of the first text segment, merge the current character string of the first text segment with the character string next to the current character string of the first text segment to obtain a merged character string; The character string includes but is not limited to a single character and a single word. The first text segment and the first similar text segment both include a plurality of character strings.
[0050] In this embodiment, starting with the first string of the first text segment, the first similar text segment is searched to see if it contains the first string of the first text segment. If the first similar text segment contains the first string of the first text segment, the first string of the first text segment is merged with the second string of the first text segment to obtain a merged string. For example, if the first text segment is "ASDFHGJ" and the first similar text segment is "AQDFGKJ", then the first similar text segment contains the first string "A". The first string "A" of the first text segment is merged with the second string "S" of the first text segment to obtain a merged string "AS".
[0051] If the first similar text segment does not contain the first character string of the first text segment, the duplicate check value of the first character of the first text segment in the first text segment is 0. Starting from the second character string of the first text segment, the first similar text segment is searched to see whether it contains the second character string of the first text segment.
[0052] Step S10712, comparing the merged character string with each character string of the first similar text segment; In an embodiment of the present application, the merged string is compared one by one with each string of the first similar text segment to determine whether the first similar text segment contains the merged string, that is, to determine whether the string in the first text segment coincides with the string of the first similar text segment.
[0053] Step S10713: If there is no string consistent with the merged string in the first similar text segment, obtain the string length of the current string of the first text segment and the string length of the first text segment; determine the duplicate check value of the current string of the first text segment based on the string length of the current string of the first text segment and the string length of the first text segment; update the next string of the current string of the first text segment to the current string, and compare the current string of the first text segment with each string of the first similar text segment until the duplicate check value of each string of the first text segment relative to the first text segment is determined; In an embodiment of the present application, if the first similar text segment does not contain the merged string, the ratio of the string length of the current string (i.e., the string before merging) to the string length of the first text segment is determined as the duplicate check value of the current string of the first text segment. The string length of the first text segment is the sum of the string lengths of all strings in the first text segment. For example, if the first similar text segment does not contain the merged string "AS", the string length 1 of the current string (i.e., the first string "A") is divided by the string length 7 of the first text segment to obtain the duplicate check value of the current string as 1 / 7.
[0054] Update the next string of the current string of the first text segment to the current string, search whether the first similar text segment contains the current string of the first text segment, and repeat steps S10711 to S10712 until the duplicate check values of each string of the first text segment relative to the first text segment are determined.
[0055] Step S10714: summing the duplicate check values of each character string in the first text segment to obtain the similarity between the first text segment and the first similar text segment; In this embodiment, after obtaining the duplicate check values of each string in the first text segment, the sum of the duplicate check values of each string is used as the similarity between the first text segment and the first similar text segment. For example, if the first text segment is "ASDFHGJ" and the first similar text segment is "AQDFGKJ", then the duplicate check value of the string "A" in the first text segment and the first text segment is 1 / 7, the duplicate check value of the string "S" in the first text segment and the first text segment is 0, the duplicate check value of the string "DF" in the first text segment and the first text segment is 2 / 7, the duplicate check value of the string "H" in the first text segment and the first text segment is 0, the duplicate check value of the string "G" in the first text segment and the first text segment is 0, and the duplicate check value of the string "J" in the first text segment and the first text segment is 1 / 7. Therefore, the similarity between the first text segment and the first similar text segment is 4 / 7.
[0056] Step S10715: Selecting a next similar text segment from the plurality of similar text segments corresponding to the first text segment and updating it as the first similar text segment; using the first character string of the first text segment as the current character string; and comparing the current character string of the first text segment with each character string of the current first similar text segment until the similarity between the first text segment and each first similar text segment is obtained; Step S10716: Select the next text segment from each of the text segments and update it as the first text segment; select a similar text segment from several similar text segments corresponding to the next text segment selected from each of the text segments and update it as the first similar text segment; take the first character string of the current first text segment as the current character string; compare the current character string of the first text segment with each character string of the current first similar text segment until the similarity between each of the text segments and the corresponding several similar text segments is obtained.
[0057] Among them, the search formula construction template is used to construct a search formula. Specifically, when constructing a search formula based on the search formula construction template, several feature participles in the feature participle set are used to determine the keywords retrieved by the search formula, and the innovation subject information is used to determine the scope of the innovation subject (submitted patent application text) retrieved by the search formula. In addition, the search formula construction template also presets patent text retrieval information, and the patent text retrieval information is used to determine the scope of the retrieved patent documents such as abstract documents, claims documents and / or specification documents.
[0058] In one embodiment, the step S105 of constructing a search formula based on a preset search formula construction template, a plurality of feature segmentations in each feature segmentation set, and the innovation subject information, and obtaining a plurality of similar texts corresponding to each text segment based on the search formula includes: All the feature participles in the feature participle set are selected as retrieval participle information, a retrieval formula is constructed based on a preset retrieval formula construction template, the retrieval participle information and the innovation subject information, the retrieval formula is input into a text retrieval database for retrieval, if the number of retrieved texts is less than a preset number threshold, the feature participles are reduced from the selected feature participles as retrieval participle information, a retrieval formula is reconstructed based on the retrieval formula construction template, the retrieval participle information and the innovation subject information, the reconstructed retrieval formula is input into the text retrieval database for retrieval, until the number of retrieved texts is greater than or equal to the preset number threshold.
[0059] The text retrieval database is determined based on the specific scenario. In a patent comparison scenario, it may be a patent database. In this case, the same text refers to the text of the same patent application, specifically including documents such as the abstract, claims, and specification. In other embodiments, the text retrieval database is determined based on other practical scenarios, and the scope of the same text is determined. For example, in a paper retrieval scenario, the text retrieval database refers to a paper database. In this case, the same text refers to the same paper.
[0060] Among them, in the patent text comparison of this embodiment, the scope of the retrieved documents covers the abstract, claims and instructions. When searching according to the search segmentation information, that is, whether the content of the abstract, claims and instructions and other documents contains all the characteristic segmentations in the search segmentation information, specifically, it can be searched from the abstract first, then from the claims, and finally from the instructions, or it can be searched from the abstract, claims and instructions at the same time. When searching in combination with the information of the innovation subject, it is to retrieve the patent application text submitted by the innovation subject and avoid retrieving texts that are not related to the innovation subject information. Based on this principle, when constructing the search formula, it can be constructed according to the search formula construction rules of the text retrieval database, and there is no specific limitation.
[0061] Specifically, if based on the common search rules in existing patent search databases, those skilled in the art know that each feature participle should be combined in an "and" relationship. For example, if the search participle information has n feature participles, the search formula composed of the search participle information part is "feature participle 1 and feature participle 2 and feature participle 3....and feature participle n". Furthermore, in combination with the scope of the searched documents, for example, including the abstract, claims, and description, and the scope of the innovative subject, the search formula is "(abstract=feature participle 1 and feature participle 2 and feature participle 3....and feature participle n) or (claims=feature participle 1 and feature participle 2 and feature participle 3....and feature participle n) or (description=feature participle 1 and feature participle 2 and feature participle 3....and feature participle n) and (innovative subject=innovative subject information)", wherein the "(innovative subject=innovative subject information)" part in the search formula is determined according to the specific situation. For example, if the innovative subject information includes the first enterprise information and the second enterprise information, then the corresponding search formula for this part can be "(applicant=first enterprise name or second enterprise name)". Based on this principle, those skilled in the art understand that the search formula construction method mentioned in this embodiment is mainly for auxiliary explanation, and the specific means of constructing the search formula in the embodiment of this application can be set according to the specific situation.
[0062] In one embodiment, in the step of constructing a search formula based on the search formula construction template, search word information, and the innovation subject information, inputting the search formula into a text search database for retrieval, and if the number of retrieved texts is less than a preset number threshold, reducing the feature word from the selected feature word as the retrieval word information, firstly, the scope of the documents to be retrieved by the search formula is determined to be abstract documents. If the number of retrieved texts is less than the preset number threshold, the scope of the retrieval documents is further determined to be claim documents. If the number of retrieved texts is less than the preset number threshold, the scope of the retrieval documents is further determined to be specification documents. If the number of texts ultimately retrieved is still less than the preset number threshold according to the specification document scope, the step of reducing the feature word from the selected feature word as the retrieval word information is performed again, and re-constructing the search formula based on the search formula construction template, the retrieval word information, and the innovation subject information for retrieval until the number of retrieved texts is greater than or equal to the preset number threshold.
[0063] As described above, in this embodiment, the characteristic participles in the segmentation information are searched together using an "and" relationship. If characteristic participles need to be reduced, they are also reduced from these characteristic participles. Specifically, the characteristic participles can be reduced one at a time, or multiple at a time. Specifically, in this embodiment, one characteristic participle is reduced at a time.
[0064] In one embodiment, if the feature segmentation set is a set obtained after expansion of equivalent feature segmentations, that is, equivalent feature segmentations are also expanded for at least part of the feature segmentations, that is, the feature segmentation set includes the feature segmentations originally identified and the equivalent feature segmentations expanded based on the feature segmentations originally identified, then after the feature segmentations are selected, the equivalent feature segmentations corresponding to the selected feature segmentations are also obtained, and the retrieval segmentation information is obtained based on the feature segmentations and the equivalent feature segmentations. When constructing a retrieval formula based on the retrieval segmentation information, these equivalent feature segmentations are spliced with the original feature segmentations using an "or" relationship, thereby improving the comprehensiveness of the retrieval and avoiding the inability of the method of this application to retrieve similar texts due to the applicant's use of equivalent replacement of feature segmentations in other patent application texts submitted. For example, if feature participle 1 is expanded to obtain equivalent feature participle 1 and equivalent feature participle 2, the corresponding expanded search formula is "(feature participle 1 or equivalent feature participle 1 or equivalent feature participle 2) and feature participle 2 and feature participle 3....and feature participle n". Similarly, when one feature participle is reduced, the corresponding equivalent feature participle in the search participle information is also reduced. For example, in the above example, if feature participle 1 is subtracted, the corresponding equivalent feature participle should also be subtracted. At this time, the search formula obtained based on the new search participle information is "feature participle 2 and feature participle 3....and feature participle n". Similarly, if feature participle 2 is subtracted, feature participle 1 and its equivalent feature participles are not affected. At this time, the search formula obtained based on the new search participle information is "(feature participle 1 or equivalent feature participle 1 or equivalent feature participle 2) and feature participle 3....and feature participle n".
[0065] In this embodiment, several similar texts are retrieved for each text fragment. Specifically, all the feature segmentation words in the feature segmentation set are first used as search segmentation word information. Combined with the innovation subject information and the search formula construction template, the text content range to be retrieved is determined, and a search formula is constructed. The search formula is input into the text database to retrieve the text containing the search segmentation word information in the text submitted by the innovation subject. If the number of retrieved texts is lower than the preset number threshold, the feature segmentation word is reduced by one to obtain new search segmentation word information, and a new search formula is constructed with the innovation subject information for retrieval. Similarly, if the number of retrieved texts is lower than the preset number threshold, the feature segmentation word is continued to be reduced by one to obtain new search segmentation word information, and a new search formula is constructed with the innovation subject information for retrieval until a number of texts not lower than the preset number threshold is retrieved. It is determined that the retrieved texts are similar texts corresponding to the text fragments, thereby fully retrieving a sufficient number of similar texts from the text submitted by the innovation subject to fully perform comparison and analysis.
[0066] In step S106, based on the feature segmentation set of each text segment, a number of similar text segments that match the text segment are intercepted from the corresponding similar texts.
[0067] In this step, similar text segments are extracted from similar texts based on the feature segmentation set corresponding to the text segments, and the extracted similar text segments include a preset number of feature segmentations, where the preset number is determined according to the specific scenario, for example, 6.
[0068] In one embodiment, the step of extracting a plurality of similar text segments that match the text segment from the corresponding similar text based on the feature segmentation set of each text segment in step S106 includes: According to the feature segmentation set corresponding to each text fragment, in each corresponding similar text, the text fragment containing a preset number of feature segmentations and starting with the first feature segmentation among the preset number of feature segmentations and ending with the last feature segmentation is determined as a similar text fragment, and each of the similar text fragments is intercepted to obtain several similar text fragments that match the text fragment.
[0069] In this embodiment, each similar text is subjected to word recognition. When the first feature segmentation word belonging to the feature segmentation set is recognized, the feature segmentation words are continuously recognized starting from the feature segmentation word. When the Nth feature segmentation word is recognized and just meets the preset number, the Nth feature segmentation word obtained at the end is used to determine that the text segment is a similar text segment. Furthermore, starting from the second feature segmentation word, similar text segments ending with the N+1th feature segmentation word are continuously determined until all similar text segments in the similar text are determined. This embodiment is based on the feature segmentation word set obtained by the original recognition and not expanded by equivalent feature segmentation words. If, in one embodiment, the feature segmentation word set is expanded by equivalent feature segmentation words, then if equivalent feature segmentation words are obtained in the similar text recognition in this step, it is also considered that the corresponding original feature segmentation words are recognized. Based on this principle, similar text segments that are similar in a substantial sense can be determined. The specific details will not be elaborated.
[0070] It can be understood that in the embodiment of the present application, the text to be detected may be divided into multiple paragraphs, each paragraph may be divided into multiple segments, each segment may be matched to multiple similar texts, and each similar text may contain multiple similar text segments.
[0071] In step S107 , similar texts of each paragraph text are determined according to the similarities between each text segment and the corresponding similar text segments.
[0072] In this embodiment, each paragraph text may be divided into multiple text segments, and each text segment may correspond to several similar text segments. For any paragraph text, similarity calculations are performed on its text segments and the corresponding similar text segments. Based on the similarity between each text segment and the corresponding similar text segment, the similar text to which the similar text segment that meets the similarity condition belongs is determined as the similar text of the paragraph text. The similarity condition is set according to the situation. In one embodiment, it can be set to meet a preset similarity threshold. In other embodiments, it can also be set to other conditions.
[0073] In one embodiment, the step of determining similar texts of each paragraph text according to the similarity between each text segment and the corresponding similar text segments in step S107 includes: Step S1071, calculating the similarity between each of the text segments and a corresponding number of similar text segments; Step S1072 : Determine the similar text corresponding to the similar text segment with the highest similarity as the similar text of the corresponding paragraph text.
[0074] In this embodiment, each paragraph may be divided into multiple text segments, and each text segment may correspond to several similar text segments. For any paragraph, similarity calculations are performed on its text segment and each of its corresponding similar text segments. Ultimately, the similar text to which the similar text segment with the highest similarity belongs is determined as the similar text of the paragraph.
[0075] In this embodiment, after calculating the similarity between the first text segment and a first similar text segment, steps S10711 to S10714 are used to calculate the similarity between the first text segment and the next first similar text segment, until the similarity between the first text segment and each first similar text segment is obtained. Then, the next text segment is selected as the first text segment, and similar text segments corresponding to the first text segment are similarly calculated according to steps S10711 to S10714, until the similarity between the next first text segment and each corresponding similar text segment is obtained.
[0076] By matching and finding the character strings in the first text segment that overlap with the first similar text segment, and calculating the duplicate check value of each character string in the first text segment relative to the first text segment, the similarity between the first text segment and each first similar text segment can be automatically and quickly obtained.
[0077] Please refer to Figure 3 In one embodiment, after the step of comparing the merged character string with each character string of the first similar text segment in step S10712, the method further includes: Step S10717: If there is a string in the first similar text segment that is consistent with the merged string, update the merged string to the current string of the first text segment, merge the current string of the first text segment with the next string of the current string of the first text segment to obtain a merged string, and compare the merged string with each string of the first similar text segment until the duplicate check value of each string of the first text segment and the first text segment is determined.
[0078] In an embodiment of the present application, if the first similar text segment contains the merged string, the merged string is further merged with the next string, and the merged string is compared with each string of the first similar text segment. If the first similar text segment no longer contains the merged string, step S10713 is executed until the duplicate check values of each string of the first text segment and the first text segment are determined. For example, the first text segment is "ABCDEF", the first similar text segment is "ABPHJK", and the first similar text segment includes the string "A" of the first text segment. Then the string "A" of the first text segment is merged with the string "B" of the first text segment to obtain the merged string "AB". The first similar text segment includes the merged string "AB". The string "C" of the first text segment is merged with the merged string "AB" to obtain the merged string "ABC". The first similar text segment does not include the merged string "ABC", and step S10713 is executed.
[0079] By comparing the merged string with the first similar text segment, the various strings in the first text segment that overlap with the similar text can be determined, and based on the duplicate check values of each overlapping string in the first text segment, the similarity of the first text segment to the first similar text segment can be determined.
[0080] In step S108 , a comparison result of the text to be detected is obtained based on similar texts of each paragraph text.
[0081] In summary, the embodiment of the present application ultimately determines similar texts of each paragraph of the text to be detected.
[0082] In one embodiment, a comparison result of the text to be detected is obtained based on similar texts of each paragraph text and corresponding similarity values.
[0083] In one embodiment, it is determined whether the similarity values corresponding to the similar texts of each paragraph text reach a preset similarity threshold, and the abnormality probability of the text to be detected is determined based on whether each paragraph text reaches the preset similarity threshold.
[0084] Specifically, in one embodiment, if there is no similarity value corresponding to any paragraph text that is greater than a preset similarity threshold, it can be determined that the probability of the text to be detected being abnormal is 0. In an optional embodiment, if there is a similarity value corresponding to similar text of a paragraph text that is greater than a preset similarity threshold, it is determined whether the similar texts of these paragraph texts are the same text, and the proportion of the paragraph texts in which the similar texts are the same text in the text to be detected is calculated, and the probability of the text to be detected being abnormal is determined based on the proportion value. Specifically, the probability of abnormal application can be determined based on the corresponding high and low proportion values. Based on this principle, the probability calculation rules can be set according to the scenario, and are not limited here. In another optional embodiment, if there are similar texts of paragraph texts whose corresponding similarity values are greater than a preset similarity threshold, it is determined whether the similar texts of these paragraph texts are the same text, and the number of paragraph texts of which the similar texts are the same text is calculated. If the number of paragraph texts of which the similar texts are the same texts is greater than the preset number threshold, it is determined that the probability of the text to be detected being abnormal is high, and the specific probability value is determined according to the defined rules, for example, it can be 85%; if the number of paragraph texts of which the similar texts are the same texts is not greater than the preset paragraph number threshold, it is determined that the probability of the text to be detected being abnormal is low, and similarly, the specific probability value is determined according to the defined rules, for example, it can be 60%.
[0085] With respect to the text comparison method of the embodiment of the present application, the following provides a test of its technical effect.
[0086] 1. Test Data Preparation 1. Text Data Collection Patent application text samples: The samples cover normal patent application texts and abnormal patent application texts. There are 20 normal patent application texts, among which normal patent application texts are patent application texts that are not collected / recorded in the public patent database; abnormal patent application texts are collected from the public patent database and modified to varying degrees, including multiple patent application texts with obviously identical invention contents, or which are essentially formed by simple combinations of different invention features or elements. The samples are no less than 50.
[0087] Text formatting: Collected text samples are formatted uniformly to ensure they are all identifiable and processable by text comparison methods, such as plain text (.txt) and rich text (.doc, .docx). For text in special formats, such as patent application documents in PDF format, specialized conversion tools are used to convert them into these processable formats, while preserving the original text's layout and content as much as possible.
[0088] 2. Information collection of innovation entities Applicant Information: For each patent application sample, collect the corresponding applicant information, including applicant information (e.g., individual name, organization name, etc.) and inventor information. Also, collect information on other entities associated with the applicant, such as the inventor's organization, other entities associated with the organization (e.g., branches), and other inventors associated with the organization, in order to construct a complete innovation entity information network. This related information can be collected and organized through various channels, including corporate organizational databases, patent databases, and internet searches.
[0089] Information Collation and Annotation: Collected applicant information and related information are collated and annotated to form structured innovation subject information data. For example, applicant information and related subject information can be stored in separate data tables, with associations established between them using association keys (such as applicant number, inventor number, etc.). Furthermore, each patent application document sample is annotated with its corresponding innovation subject information to ensure accurate matching and comparison between the document and its innovation subject information during testing.
[0090] 3. Preparation of similar text data Similar Text Generation: Based on collected patent application text samples, texts similar to the original samples are generated through manual editing or text generation tools to serve as comparative data for testing the similar text retrieval function. These similar texts can be generated by making simple modifications to the original samples (such as replacing some sentences or paragraphs, adjusting the order of sentences or paragraphs, adding or deleting content, etc.) to simulate anomalous patent application texts that may occur in practice. The number of similar texts generated should be no less than 50, and should cover varying degrees of similarity to comprehensively test the method's similar text retrieval capabilities and the accuracy of similarity calculations.
[0091] Similar text annotation: Generate similar text and annotate it to clarify its degree of similarity to the original sample text (e.g., high similarity, moderate similarity, low similarity, etc.) and the specific content of the similarity (e.g., which paragraphs and sentences are similar). This annotation information will serve as a reference for evaluating the accuracy of similar text retrieval and similarity calculation results during testing.
[0092] 2. Test case description Use case 1: Comparison of 20 patent application documents not collected or recorded in public patent databases was performed using the text comparison method of this embodiment to observe whether it can correctly identify them as normal documents and provide accurate comparison results. The expected result is that these documents can be accurately identified as normal patent application documents and not mistakenly identified as abnormal documents.
[0093] Use case 2: Abnormal patent application text comparison: 50 published patent documents from several applicants are collected and sample patent application documents with abnormal invention content (obviously identical content, or essentially formed by simple combinations and variations of different invention features or elements) are generated. The text comparison method of the present embodiment is then used to process these documents to see whether the identical content between these documents can be accurately identified and identified as abnormal. The expected result is that the similarities between these documents can be identified with high accuracy and a clear abnormality determination result can be given.
[0094] 3. Test results Normal patent application text comparison results: Among the 20 normal patent application text samples tested, the text comparison method of the embodiment of the present application identified 0 abnormal texts, with an accuracy rate of 100%.
[0095] Abnormal patent application text comparison results: Of 50 patent application text samples with abnormal invention content, the text comparison method of the present embodiment successfully identified 48 abnormal texts, with an accuracy rate of 96%. Among them, the two abnormal texts that were not detected were texts with low similarity. The ability to accurately identify the common content between these texts and determine them as abnormal texts demonstrates a strong ability to identify texts with abnormal content.
[0096] In summary, the text comparison method of the present embodiment can accurately identify normal patent application text as normal patent application text and does not misidentify it as abnormal text, indicating that it has high accuracy when processing normal text. At the same time, it can accurately identify abnormal patent application text as abnormal patent application text.
[0097] Please refer to Figure 4 The second aspect of the embodiment of the present application further discloses a text comparison device, comprising: The to-be-detected text acquisition module 201 is used to acquire the to-be-detected text and the innovation subject information of the to-be-detected text; The to-be-detected text paragraph segmentation module 202 is used to segment the to-be-detected text into paragraphs to obtain a plurality of paragraphs of text; The paragraph text segmentation module 203 is used to segment each paragraph text into text segments to obtain a plurality of text segments corresponding to each paragraph text; The text segment feature segmentation recognition module 204 is used to perform feature segmentation recognition on each text segment to obtain a feature segmentation set corresponding to each text segment; The text segment similar text matching module 205 is used to construct a search formula based on a preset search formula construction template, several feature segmentation words in each feature segmentation word set, and the innovation subject information, and obtain several similar texts corresponding to each text segment based on the search formula; The similar text similar segment interception module 206 is used to intercept a number of similar text segments that match the text segment from the corresponding similar text according to the feature segmentation set of each text segment; wherein the similar text segment includes a preset number of feature segmentations in the feature segmentation set; A paragraph similar text determination module 207 is configured to determine similar texts of each paragraph text based on similarities between each text segment and corresponding similar text segments; The comparison result generating module 208 is configured to obtain a comparison result of the text to be detected based on similar texts of each paragraph text.
[0098] It should be noted that the text comparison device provided in the above embodiment, when executing the text comparison method, only uses the division of the above functional modules as an example. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the text comparison device and the text comparison method provided in the above embodiment are based on the same concept. The implementation process is detailed in the above text comparison method embodiment and will not be repeated here.
[0099] A third aspect of the embodiments of the present application further discloses a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of any of the methods described in any of the embodiments of the present application. That is, those skilled in the art will appreciate that all or part of the steps in the methods of the aforementioned embodiments can be implemented by instructing the relevant hardware through a program. The program, stored in a storage medium, includes instructions for causing a device (such as a microcontroller or chip) or a processor to perform all or part of the steps of the methods described in each embodiment of the present application. The computer program may include computer program code, which may be in source code form, object code form, an executable file, or some intermediate form. The aforementioned storage medium includes any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a removable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. It should be noted that the content contained in computer-readable media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable media does not include electrical carrier signals and telecommunications signals.
[0100] Please refer to Figure 5 The fourth aspect of the embodiment of the present application also discloses an electronic device 301, including: a processor 302, a memory 303, and a computer program 304 stored in the memory 303 and executable on the processor 302. When the processor 302 executes the computer program 304, the steps of the method described in any one of the embodiments of the present application are implemented.
[0101] The processor 302 may include one or more processing cores. The processor 302 connects to various components within the electronic device 301 using various interfaces and circuits. It executes instructions, programs, code sets, or instruction sets stored in the memory 303 and accesses data from the memory 303 to perform various functions and process data within the electronic device 301. Optionally, the processor 302 may be implemented in the form of at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 302 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the touchscreen display; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 302 and may be implemented as a separate chip.
[0102] The memory 303 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 303 includes a non-transitory computer-readable storage medium. The memory 303 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 303 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch control instructions), instructions for implementing the aforementioned method embodiments, and the data storage area may store data involved in the aforementioned method embodiments. The memory 303 may also optionally be at least one storage device located remotely from the aforementioned processor 302.
[0103] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present application, and the present application is intended to encompass such modifications and variations.
Claims
1. A text comparison method, characterized in that: The following steps are involved: Obtaining the text to be detected and the innovation subject information of the text to be detected; Segmenting the text to be detected into paragraphs to obtain a plurality of paragraph texts; Segmenting each paragraph text into text segments to obtain a plurality of text segments corresponding to each paragraph text; Performing feature segmentation recognition on each of the text segments to obtain a feature segmentation set corresponding to each of the text segments; Based on a preset search formula construction template, several characteristic segmentation words in each characteristic segmentation word set and the innovation subject information, a search formula is constructed, and based on the search formula, several similar texts corresponding to each of the text fragments are obtained; According to the feature segmentation set of each text segment, a number of similar text segments matching the text segment are intercepted from the corresponding similar text; wherein the similar text segments include a preset number of feature segmentations in the feature segmentation set; Determining similar texts for each of the paragraph texts based on similarities between each of the text segments and the corresponding similar text segments; A comparison result of the text to be detected is obtained based on the similar text of each paragraph text.
2. The text comparison method according to claim 1, characterized in that: The step of constructing a search formula based on a preset search formula construction template, several feature segmentations in each feature segmentation set, and the innovation subject information, and obtaining several similar texts corresponding to each text segment based on the search formula includes: All the feature participles in the feature participle set are selected as retrieval participle information, a retrieval formula is constructed based on a preset retrieval formula construction template, the retrieval participle information and the innovation subject information, the retrieval formula is input into a text retrieval database for retrieval, if the number of retrieved texts is less than a preset number threshold, the feature participles are reduced from the selected feature participles as retrieval participle information, a retrieval formula is reconstructed based on the retrieval formula construction template, the retrieval participle information and the innovation subject information, the reconstructed retrieval formula is input into the text retrieval database for retrieval, until the number of retrieved texts is greater than or equal to the preset number threshold.
3. The text comparison method according to claim 1, characterized in that: The step of extracting a plurality of similar text segments that match the text segment from corresponding similar texts based on the feature segmentation set of each text segment comprises: According to the feature segmentation set corresponding to each text fragment, in each corresponding similar text, the text fragment containing a preset number of feature segmentations and starting with the first feature segmentation among the preset number of feature segmentations and ending with the last feature segmentation is determined as a similar text fragment, and each of the similar text fragments is intercepted to obtain several similar text fragments that match the text fragment.
4. The text comparison method according to claim 1, wherein: The step of determining similar texts of each paragraph text according to the similarity between each text segment and the corresponding similar text segments includes: Calculating the similarity between each of the text segments and a corresponding number of similar text segments; The similar text corresponding to the similar text segment with the highest similarity is determined as the similar text of the corresponding paragraph text.
5. The text comparison method according to claim 4, characterized in that: The step of calculating the similarity between each of the text segments and the corresponding similar text segments includes: Selecting a text segment from each of the text segments as a first text segment, selecting a similar text segment from several similar text segments corresponding to the first text segment as a first similar text segment, taking the first character string of the first text segment as a current character string, comparing the current character string of the first text segment with each character string of the first similar text segment, and if there is a character string in the first similar text segment that is consistent with the current character string of the first text segment, merging the current character string of the first text segment with a character string next to the current character string of the first text segment to obtain a merged character string; comparing the merged character string with each character string of the first similar text segment; If there is no string consistent with the merged string in the first similar text segment, obtain the string length of the current string of the first text segment and the string length of the first text segment; determine the duplicate check value of the current string of the first text segment based on the string length of the current string of the first text segment and the string length of the first text segment; update the next string of the current string of the first text segment to the current string, and compare the current string of the first text segment with each string of the first similar text segment until the duplicate check value of each string of the first text segment relative to the first text segment is determined; Summing the duplicate check values of each character string in the first text segment to obtain the similarity between the first text segment and the first similar text segment; Selecting a next similar text segment from the plurality of similar text segments corresponding to the first text segment and updating it as the first similar text segment, taking the first character string of the first text segment as the current character string, and comparing the current character string of the first text segment with each character string of the current first similar text segment until the similarity between the first text segment and each first similar text segment is obtained; The next text segment is selected from each text segment and updated as the first text segment. A similar text segment is selected from several similar text segments corresponding to the next text segment selected from each text segment and updated as the first similar text segment. The first character string of the current first text segment is used as the current character string. The current character string of the first text segment is compared with each character string of the current first similar text segment until the similarity between each text segment and the corresponding several similar text segments is obtained.
6. The text comparison method according to claim 5, characterized in that: After the step of comparing the merged character string with each character string of the first similar text segment, the method further includes: If there is a string in the first similar text segment that is consistent with the merged string, update the merged string to the current string of the first text segment, merge the current string of the first text segment with the next string of the current string of the first text segment to obtain a merged string, and compare the merged string with each string of the first similar text segment until the duplicate check value of each string of the first text segment and the first text segment is determined.
7. The text comparison method according to claim 1, characterized in that: The step of obtaining the text to be detected and the innovation subject information of the text to be detected includes: Obtaining information of the first applicant who submitted the text to be tested; Obtaining, based on the first applicant information, second applicant information associated with the first applicant information; The innovation subject information is obtained according to the first applicant subject information and the second applicant subject information.
8. A text comparison device, characterized in that: include: A module for acquiring text to be detected, used to acquire the text to be detected and the innovation subject information of the text to be detected; A paragraph segmentation module for the text to be detected is used to segment the text to be detected into paragraphs to obtain a plurality of paragraph texts; A paragraph text segmentation module is used to segment each paragraph text into text segments to obtain a plurality of text segments corresponding to each paragraph text; A text segmentation feature recognition module is used to perform feature segmentation recognition on each text segment to obtain a feature segmentation set corresponding to each text segment; A text segment similar text matching module is used to construct a search formula based on a preset search formula construction template, a number of characteristic segmentation words in each characteristic segmentation word set, and the innovation subject information, and obtain a number of similar texts corresponding to each text segment based on the search formula; A similar text similar segment interception module is used to intercept a number of similar text segments that match the text segment from the corresponding similar text according to the feature segmentation set of each text segment; wherein the similar text segment includes a preset number of feature segmentations in the feature segmentation set; A paragraph similar text determination module, configured to determine similar texts of each paragraph text based on similarities between each text segment and corresponding similar text segments; The comparison result generating module is used to obtain the comparison result of the text to be detected based on the similar text of each paragraph text.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory, wherein the computer program implements the text comparison method according to any one of claims 1 to 7 when executed by the processor.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the text comparison method according to any one of claims 1 to 7.