Method for generating a database query
The method addresses the challenge of incomplete text generation in patent documents by using semantic-syntactic parsing and classification models to extract and classify text segments, ensuring accurate extraction and classification of technical results.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-03-26
AI Technical Summary
Existing methods for automated text processing, particularly in the formation of patent documents, fail to accurately extract and classify specific parts such as the technical result, leading to incomplete or inaccurate generation of texts.
A method and system for automated natural language processing that involves identifying and parsing segments of text, applying semantic-syntactic analysis to extract main entities and associated entities, forming a text corpus, and using classification and clustering models to associate statements with these entities, enabling accurate classification and extraction of technical results.
Ensures accurate classification and extraction of specific texts, expanding the capabilities for automated processing of natural language texts, particularly in patent documents, by generating a query to a database that contains text parses associated with statements.
Smart Images

Figure 00000026_0000 
Figure 00000027_0000
Abstract
Description
METHOD OF FORMING A REQUEST TO THE DATABASE
[0001] AREA OF TECHNOLOGY
[0002] The proposed technical solution relates to methods of automated text processing and can be used in the formation of text corpora.
[0003] LEVEL OF TECHNOLOGY
[0004] Various methods for automated generation of texts in natural language are known, for example, those disclosed in the documents: US20070136321A1 dated 06 / 14 / 2007, US20100205125A1 dated 08 / 12 / 20120, US20190377780A1 dated 12 / 12 / 2019, US20130080883A1 dated 03 / 28 / 2013, US10713443B 1 dated 07 / 14 / 2020, US10747953B1 dated 08 / 18 / 2020, US11023662B2 dated 06 / 01 / 2021, US11341323B1 dated 05 / 24 / 2022. However, the known solutions do not allow the formation of texts in natural language with sufficient completeness, specific to a certain area, such as, for example, patent documents, in which, in addition to the actual formation of disclosures, it is also necessary to ensure the formation of special parts of the patent document, such as, for example, the level of technology (Background of the Invention), for which in most cases it is necessary to have a statement in the form of a disclosure of the technical problem solved by the invention or the technical result achieved by using the invention (technical effect (9.2.8 Case Law of the Boards of Appeal), (technical) result (MPEP 716.02(a))).
[0005] Patent No. US11593564B2 of February 28, 2023 (D1) discloses systems, methods, and storage media for extracting patent document templates from a patent corpus. Examples of implementation include: obtaining a patent corpus; obtaining one or more parameters; determining one or more subsets of the patent corpus by filtering the patent corpus based on one or more parameters; identifying one or more clusters of documents within individual ones of one or more subsets of the patent corpus; obtaining a patent document template corresponding to the first cluster of documents; and / or performing other operations. However, the solution known from D1 does not disclose any methods or techniques for extracting, in particular, statements concerning the technical result.
[0006] From the patent US5774833A of 30.06.1998 (D2), a method for processing patent text on a computer is known, including determining the boundaries of parts of the patent text, loading at least one of the parts of the patent text into working memory, analyzing at least one of the parts of the patent text and reporting the results. The user. Moreover, the alphanumeric data of the drawing can also be compared with the patent text. Method D2 can be combined with a word processor. Method D2 can recognize and report claim dependencies, specific characteristics of the patent text, and patent errors based on legal standards, standards of practice, USPTO standards, or even user preferences. However, the solution known from D2 does not disclose any methods or techniques for extracting, in particular, claims concerning the technical result.
[0007] Thus, there is a problem of automated extraction of claims from specific natural language texts, such as, but not limited to, patent documents.
[0008] The solution known from D2 can be chosen as the closest analogue.
[0009] DISCLOSURE OF THE INVENTION
[0010] The technical problem solved by the claimed invention is the creation of a method and / or a computer device and / or a system and / or a machine-readable data carrier that does not have the disadvantages of analogues and thus ensures the accurate classification of specific texts.
[0011] The technical result achieved by implementing the claimed invention, in addition to fulfilling its intended purpose, is to eliminate the shortcomings of similar technologies, thereby ensuring accurate classification of specific texts. Another technical result is the expansion of the arsenal of technical means—methods for automated processing of natural language text.
[0012] The technical result is achieved due to the fact that a method for generating a query to a database, executable by a processor of a computer device, is provided, in which a query to a database is generated, containing at least one statement obtained using a method for generating a text corpus; wherein the said database contains at least a plurality of text parses, each of which is associated with one or more of the said statements; wherein the method for generating a text corpus is a method executable by the processor of a computer device, in which at least by means of a method for automated processing of text in a natural language a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the first segment, and a text corpus is formed from the received pairs of texts; wherein the method for automated processing of text in a natural language is a method executable by a processor of a computing device, which consists of at least performing the following steps: a step of identifying a text in a natural language, including at least three segments; a step of identifying segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of a text in a natural language; a step of marking in the first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment at least one part to be parsed; a step of parsing the marked parts by means of semantic-syntactic analysis;a step of extracting from the part of the first segment subjected to semantic-syntactic parsing at least the main entity of the first segment, and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated end entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step of associating said statement with said main entity of the first segment.
[0013] BRIEF DESCRIPTION OF DRAWINGS
[0014] Illustrative embodiments of the present invention are described below in detail with reference to the accompanying drawings, which are incorporated herein by reference, and in which:
[0015] Fig. 1, by way of example and not limitation, shows an exemplary flow chart of a method 100 for automated natural language text processing.
[0016] Fig. 2, by way of example and not limitation, shows an exemplary diagram of a system 200 for automated natural language processing.
[0017] IMPLEMENTATION OF THE INVENTION
[0018] In a preferred embodiment of the present invention, there is provided a method executable by a processor of a computer device for automated processing of natural language text, which method comprises at least the following steps: a step of identifying a natural language text comprising at least three segments; a step of identifying segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step of marking up in the selected first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment at least one part to be parsed; a step of parsing the marked parts using semantic- syntactic parsing; a step of extracting from the part of the selected first segment subjected to semantic-syntactic parsing at least the main entity of the first segment, and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step of associating said statement with said main entity.
[0019] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-combined to obtain identifiable text in a natural language.
[0020] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-associated to obtain identifiable text in a natural language.
[0021] In a particular embodiment of the present invention, the said method is provided, characterized in that the part to be parsed marked in the selected first segment is the first sentence in a natural language.
[0022] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.
[0023] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested in nested associated entities, until a nested associated entity is retrieved in which no entity is nested.
[0024] In a particular embodiment of the present invention, the said method is provided, characterized in that the said part to be parsed, marked in the selected first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing.
[0025] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the selected first segment and all associated entities associated with it are extracted from the first part.
[0026] In a particular embodiment of the present invention, the said method is provided, characterized in that at least the main entity of the second part and at least an associated entity linked to the main entity of the second part are extracted from the second part.
[0027] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.
[0028] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested in nested associated entities, until a nested associated entity is retrieved in which no entity is nested.
[0029] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the second part is associated with the main entity of the selected first segment.
[0030] In a particular embodiment of the present invention, the said method is provided, characterized in that for each extracted entity, lemmatization is at least partially performed.
[0031] In a particular embodiment of the present invention, the said method is provided, characterized in that the marking is carried out using a classification model and / or a clustering model.
[0032] In a particular embodiment of the present invention, the said method is provided, characterized in that the division of the part of the first segment to be analyzed is carried out using a classification model and / or a clustering model.
[0033] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computing device, is provided, in which, at least by means of said method automated processing of natural language text produces a plurality of text pairs, each of which includes at least one statement related to the main essence of the selected first segment, and forms a text corpus from the resulting text pairs.
[0034] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computer device, is provided, in which, at least by means of said method of automated processing of text in a natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the second part of the part to be parsed marked up in the selected first segment, and a text corpus is formed from the obtained pairs of texts.
[0035] In another preferred embodiment of the present invention, a method for pre-training, or training, or further training a classification model, executable by a processor of a computer device, is provided, in which at least the classification model is trained to classify a text analysis as associated with any statement obtained using any mentioned method for forming a text corpus, when receiving at the input of the classification model a text analysis containing at least a main entity and at least an associated terminal entity associated with the main entity, wherein the training is carried out using a text corpus obtained by any mentioned method for forming a text corpus.
[0036] In another preferred embodiment of the present invention, a method for pre-training, or training, or additional training of a clustering model, executable by a processor of a computer device, is provided, in which at least the clustering model is trained to classify a text analysis as associated with any statement obtained using any mentioned method for forming a text corpus, when receiving at the input of the clustering model a text analysis containing at least a main entity and at least an associated leaf entity associated with the main entity, wherein the training is carried out using a text corpus obtained by any mentioned method for forming a text corpus.
[0037] In another preferred embodiment of the present invention, a method for classifying a text parsing, executable by a processor of a computing device, is provided, in which at least a text parsing is fed to the input of said classification model, containing at least the main an entity and an associated terminal entity associated with the main entity, and classify said text analysis as being associated with any statement obtained using any said method of forming a text corpus.
[0038] In another preferred embodiment of the present invention, a method for classifying text parsing, executable by a processor of a computing device, is provided, which at least provides at the input of said clustering model a text parsing containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text parsing as being associated with any statement obtained using any said method for forming a text corpus.
[0039] In another preferred embodiment of the present invention, a method for generating a database, executable by a processor of a computer device, is provided, in which a plurality of text parses classified by any said classification method are identified, they are associated with one or more corresponding statements, and a database is generated containing at least a plurality of said text parses, each of which is associated with one or more said statements.
[0040] In another preferred embodiment of the present invention, a method for generating a query to a database, executable by a processor of a computer device, is provided, in which a query to a database is generated, containing at least one statement obtained using any of the mentioned methods for generating a corpus of texts; wherein the mentioned database contains at least a plurality of text parses, each of which is associated with one or more of the mentioned statements.
[0041] In another preferred embodiment of the present invention, a method for selecting text parses, executable by a processor of a computer device, is provided, in which at least a query is generated and sent to a database, containing at least some statement obtained using some method for forming a corpus of texts, and at least one text parse is obtained, associated with said statement; wherein said database contains at least a plurality of text parses, each of which is associated with one or more of said statements.
[0042] In another preferred embodiment of the present invention, there is provided a method executable by a processor of a computer device for generating a text entry, in which at least one statement obtained using any of the above is received and a syntactically and semantically correct sentence is generated, including the statement or a derivative thereof.
[0043] In another preferred embodiment of the present invention, a method for generating a database of text records, executable by a processor of a computer device, is provided, which includes at least obtaining a plurality of text records by means of said method for generating text records and recording the obtained text records in a database; the approval is obtained using any of said methods for generating a corpus of texts.
[0044] In another preferred embodiment of the present invention, a method for generating natural language text, executable by a processor of a computer device, is provided, which comprises at least receiving a text entry from a database formed by said method for forming a database including an assertion of text entries, after which generating natural language text, including at least said text entry and a text that is not said entry.
[0045] In another preferred embodiment of the present invention, a method for generating natural language text, executable by a processor of a computer device, is provided, which at least identifies a natural language text entry, including at least a main entity and at least an associated end entity linked to it, and a statement linked to the main entity, after which a text entry including this statement is obtained from a database formed by said method for forming a database of text entries including the statement, and generates natural language text, including at least the identified text entry, and a text entry obtained from said database
[0046] In another preferred embodiment of the present invention, there is provided a computer device for automated natural language processing, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any of the mentioned methods of automated natural language text processing.
[0047] In another preferred embodiment of the present invention, a computer device for generating a text corpus is provided, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any of the said methods for generating the text corpus.
[0048] In another preferred embodiment of the present invention, a computer device is provided for pre-training, or training, or re-training a classification model, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of pre-training, or training, or re-training a classification model.
[0049] In another preferred embodiment of the present invention, a computer device is provided for pre-training, or training, or re-training a clustering model, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of pre-training, or training, or re-training a clustering model.
[0050] In another preferred embodiment of the present invention, there is provided a computer device for classifying text parsing, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said text parsing classification method.
[0051] In another preferred embodiment of the present invention, a computing device for generating a database is provided, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any said method of generating a database.
[0052] In another preferred embodiment of the present invention, a computing device is provided for generating a query to a database, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a query to a database.
[0053] In another preferred embodiment of the present invention, there is provided a computer device for selecting text parses, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for selecting text parses.
[0054] In another preferred embodiment of the present invention, there is provided a computer device for generating a text entry, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a text entry.
[0055] In another preferred embodiment of the present invention, a computer device is provided for generating a database of text entries including an approval, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a database of text entries including an approval.
[0056] In another preferred embodiment of the present invention, there is provided a computer device for generating natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating natural language text.
[0057] In another preferred embodiment of the present invention, there is provided a machine-readable storage medium, including a non-transitive machine-readable storage medium, containing program code that, when executed by a processor, causes the processor to perform the actions of any said method.
[0058] The following are embodiments of the present invention, revealing examples of its implementation in specific embodiments. However, the description itself is not intended to limit the scope of rights granted by this patent. Rather, it should be understood that the claimed invention may also be implemented in other ways that incorporate different elements and conditions, or combinations of elements and conditions similar to those described herein, in combination with other existing and future technologies.
[0059] In this description, a patent document refers to, without limitation, a patent or patent application, that is, as a rule, a specific text, consisting of at least three segments: a first segment, which is a patent claim, a second segment, which is a description (description, specification), and a third segment, which is an abstract (abstract), wherein, without limitation, the numbering of the segments above is not given in the order of their sequence in the text in natural language, but only for the simplicity of presentation and association in the present document, as will be shown below. Preferably, without limitation, the formulas of patent documents do not relate to new chemical compounds and other similar objects created for the first time, since with respect to such objects the statement of the technical result does not require disclosure.
[0060] In Fig. 1, by way of example and not limitation, an exemplary flow chart of a method 100 for automated processing of natural language text is shown. The method 100 comprises at least the following steps: a step 101 of identifying a natural language text that includes at least three segments; a step 102 of identifying the segments; a step 103 of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step 104 of marking up only one part to be parsed in the selected first segment and a step of marking up at least one part to be parsed in the selected second segment and / or in the selected third segment; a step 105 of parsing the marked parts by means of semantic-syntactic parsing;step 106 of extracting from the part of the selected first segment subjected to semantic-syntactic parsing at least the main entity of the first segment and at least an associated entity associated with the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement;step 107 of associating said statement with said principal entity. For example, but not limited to, step 101 of identifying natural language text involves reliably identifying natural language text, such as, but not limited to, a patent document. For example, but not limited to, segments of the patent document text may be represented by separate natural language texts, for which, if necessary, they are first associated, for example, by assigning them some relatedness features, such as unique identifiers, the decoding of which makes it possible to determine the relatedness of the segments among themselves. Such methods and techniques for associating entities in data storage systems are widely known in the prior art and, accordingly, in detail; are not further described. For example, but not limited to, a patent document may be a single text, typically with continuous page numbers; in such a case, prior association of segments with one another is not required.
[0061] Preferably, but not limited to, step 102 involves identifying segments of the natural language text; for example, but not limited to, in the case where the natural language text is a patent document, at least the patent claims, which is the first segment in the context of the present invention, the description, which is the second segment in the context of the present invention, and the abstract, which is the third segment in the context of the present invention, are identified. Without limitation, in the case where the natural language text is a non-patent document, the segments are identified in a similar manner. In this case, without limitation, it should mainly be assumed that the first segment includes a set of entities, while the second and third segments include at least one or more statements characteristic of said set of entities.An example of a suitable non-patent document may be, for example, but not limited to, a scientific article in which the role of the first segment may be assigned to the abstract, and the roles of the second and third segments, respectively, to the main text and conclusion.
[0062] Preferably, but not limited to, step 103 comprises selecting the identified segments, including, but not limited to, selecting the first segment as comprising the set of entities and one of the second segment or the third segment. In this case, but not limited to, the choice of which of the second segment or the third segment to select depends on the probability with which the claim will be present in a particular segment. More specifically, but not limited to, the presence of the claim in a particular segment will generally depend on the approach to the requirements imposed on the natural language text. For example, but not limited to, when the natural language text is a patent document, then, depending on the patent system in which it is formed, the third segment (abstract) will at least contain a claim about the technical result.For example, but not limited to, Russian patent documents are typically published with an abstract formulating the technical result. Moreover, as of the filing date of the present application, the preparation of the abstract, when deciding whether to grant a patent, is the responsibility of the examiner, who is competent in formulating the technical result, thus eliminating inaccuracies in its presentation. At the same time, in patent documents, In US patent documents, the abstract is typically prepared by the applicant, resulting in abstracts in US patent documents lacking a clear structure, and the presence of a technical result claim is not guaranteed. Thus, for example, without limitation, when choosing between the second and third segments in the US, it would be more effective to select the second segment, since at least the Background of the Invention or Brief Summary of Invention sections are likely to contain a technical result claim, as required by MPEP 608.01(c) and 608.01(d). At the same time, for example, without limitation, for each natural language text, a particular segment can be selected based on the presence of the claim itself.For this purpose, for example, but not limited to, a semantic-syntactic analysis can be performed for each natural language text, for example, using a semantic parser, the technology of which is widely represented in the prior art and, accordingly, is not described below. Based on the results of the semantic-syntactic analysis, it becomes possible to identify the specific portion of the natural language text containing the desired assertion. Preferably, but not limited to, the first segment of the natural language text is preliminarily excluded, as it is known to contain no assertions, and whose presence during the semantic-syntactic analysis could distort the results.
[0063] Preferably, but not limited to, step 104 involves annotating the parts of the selected segments. Preferably, but not limited to, only one part of the selected first segment to be parsed is annotated, namely, the part containing the set of entities. For example, without limitation, in the case where the natural language text is a patent document, the first sentence, i.e., the first claim, is selected as the annotated part, as it is most likely to contain the required set of entities, i.e., the set of essential features. Furthermore, without limitation, there are situations where the set of entities (features) in the first claim, in addition to the essential features, also includes nonessential features.In this case, such a selected first segment may be subjected to additional analysis after finding the technical result statement in the second or third segment and checking which of the features do not affect the achievement of the technical result; based on the results of such additional analysis, the marked portion is cleared of unnecessary entities. In this case, preferably, but not limited to, the marked and optionally cleared portion of the selected first segment to be analyzed may be subjected to analysis to determine the presence of a generic part (first). part) and a distinctive part (the second part); in this case, for example, but not limited to, it may be sufficient to determine the linking word separating the generic part from the distinctive part. At the same time, not every patent system requires that an independent claim be composed with the use of a distinctive part, which results in the generic part not being clearly distinguished and not cut off by a linking word; in such a case, for example, but not limited to, a preliminary semantic-syntactic analysis of the entire first segment may be carried out in order to identify, at least, the main essence of the first segment and the main essences of the entities associated with the main essence of the first segment, which will be discussed in detail below.Typically, but not limited to, when the natural language text is a patent document, the primary entity will be defined as a generic concept, and the associated entities will be individual features, such as, but not limited to, method steps or product parts. The resulting set of primary entities and associated entities can be subjected to a familiarity analysis, which can determine whether the entire set is generic or only a subset of the features constitutes a generic part. For example, but not limited to, the familiarity analysis can be performed automatically, based on a pre-prepared text corpus, for each of which a first segment has been identified and a semantic-syntactic analysis has been performed on the part to be analyzed.In this case, but not limited to, separation may be performed either manually or automatically, for example, using a pre-trained classification model and / or clustering model. In this case, but not limited to, the selected second segment or the selected third segment is labeled in such a way as to obtain the desired statement with the highest probability. Typically, the statement of the desired statement is preceded by a certain text structure, or the statement of the desired statement includes certain keywords. Without limitation, several parts of the selected second segment or the selected third segment to be analyzed may be labeled in this manner. In this case, but not limited to, labeling may be performed either manually or automatically, for example, using a pre-trained classification model and / or clustering model.Preferably, but not limited to, step 105 comprises parsing the marked parts by means of semantic-syntactic parsing using, for example, but not limited to, the mentioned syntactic parser.
[0064] Preferably, but not limited to, step 106 comprises extracting from the parsed portions of the first segment at least the main entity of the first segment and at least an associated entity linked to the main entity of the first segment, wherein at least one of the extracted associated entities is an associated terminal entity. For example, without limitation, in the case where the first segment is a patent claim, the first independent claim will be subjected to semantic-syntactic parsing; in such a case, the main entity will be a generic concept that is the root of the parsing tree, the associated entities will be the features of the technical solution that are the nodes of the parsing tree, and at least one associated entity will be an associated terminal entity, that is, it will be a feature that has only one edge, that is, it will be a terminal node (leaf) of the parsing tree.Moreover, the parse tree can be constructed using the Abstract Syntax Tree (AST) principle, meaning it can be stripped of all irrelevant features, as well as repetitions, anaphora, and other elements that don't affect the claim—in the case of patent claims, those that can't affect the technical result. Examples of redundant elements obtained after parsing include, but are not limited to, commas or other separators, which are useful directly during parsing, such as as links or boundaries, but are redundant for the text corpus when used, for example, to train a classification model and / or clustering model.Furthermore, for example, without limitation, when the first segment has been divided into a generic part and a distinctive part, the main entity of the first segment, after the aforementioned analysis, will be found in the first part, that is, in the generic part, along with the entities associated with it from the generic part; while, without limitation, the second part, which is the distinctive part, will in such a case have the main entity of the second part and the entities associated with it, which, nevertheless, are also all associated with the main entity of the first segment, since all the individual features revealed by means of analysis directly or indirectly belong to only one generic concept.In this case, for example, without limitation, each associated entity extracted in this way can be iteratively subjected to its own semantic-syntactic analysis, during which the corresponding main entities of the associated entities and the associated entities of the associated entities associated with them will be discovered, that is, each associated entity can have a plurality of nested associated entities up to one or more. associated terminal entities that can no longer be subjected to semantic-syntactic parsing because their features are too simple. In this case, for example, but not limited to, lemmatization can be performed to reduce words to their original forms. Furthermore, without limitation, for each entity, including the main entity, as well as any associated entities, a generalization operation can be performed, in which the entity can be elevated to the level of a method or product, which corresponds to a certain function determined by the entity's features. In this case, preferably, without limitation, at least one statement is extracted from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing.For example, without limitation, in the case where the natural language text is a patent document, the semantic-syntactic analysis of paragraph 0010 of the present description in its original version will yield two statements: "accurate generation of a text corpus" and "automated generation of a text corpus." For example, without limitation, in the same case, the semantic-syntactic analysis of paragraph 0011 may yield three statements: "implementation of the purpose - generation of a text corpus," "accurate generation of a text corpus," and "automated generation of a text corpus." Preferably, without limitation, at step 107, at least one extracted statement is associated with the main entity of the first segment, and, thus, the statement is associated with any associated entities related to the main entity of the first segment.In this case, for example, but not limited to, some resulting statements, such as "implementation of purpose," are of no practical use, as they cannot be subsequently used in a text corpus to train a classification and / or clustering model, for example, to predict a statement corresponding to a new set of entities. This would lead to any classifier based on the classification and / or clustering model assigning any set of entities to such a statement, as such a statement is inherent in all resulting text parses. For this reason, it is preferable, but not limited to, to clean up the resulting "statement-main entity" pairs, removing redundant statements.
[0065] Thus, preferably, without limitation, it becomes possible to obtain a corpus of texts, for which purpose by means of said method 100 in step 107 by associating said statement with said main entity of the first segment (and, as a consequence, with all associated entities related to the main entity) a plurality of text pairs are obtained, each of which includes at least one statement associated with the main entity of the first segment, and a text corpus is formed from the obtained text pairs. Alternatively, without limitation, it becomes possible to obtain a text corpus, which, by means of the said method 100 at step 107 by associating the said statement with the said main entity of the second part of the part to be parsed marked in the first segment (and, as a consequence, with all associated entities associated with the main entity of the second part of the part to be parsed marked in the first segment), a plurality of text pairs are obtained, each of which includes at least one statement associated with the main entity of the second part of the part to be parsed marked in the first segment, and a text corpus is formed from the obtained text pairs.Such resulting text corpora, for example, but not limited to, may be used, either jointly or separately, for pre-training, or training, or additional training of a classification model and / or a clustering model, in which, at least, the classification model and / or the clustering model are trained to classify a text parse as related to any said statement, upon receipt at the input of the classification model and / or the clustering model of a text parse containing, at least, a main entity and an associated terminal entity associated with the main entity, wherein the training is carried out on one of the said text corpora, jointly or separately. In this case, without limitation, the main entity in this way is understood to be, for example, the main entity of the first segment or the said main entity of the second part of the part of the first segment to be parsed.To train a classification model and / or a clustering model, for example, but not limited to, a sample from a text corpus is provided, after which the classification model and / or clustering model is trained on such a sample. In this case, for example, but not limited to, the classification model is based on one of the following: logistic regression, support vector machine, decision trees method, random forest method, naive Bayes classifier, k-nearest neighbors method, neural network, boosting, gradient boosting, bagging, group argument accounting method; in this case, the clustering model is based on one of the following: k-means method, density-based spatial clustering of applications with noise (DBSCAN), hierarchical clustering method, spectral clustering method.In this case, without limitation, any suitable known or future known classification model and / or clustering model may be used, the choice of which should be based on the fact that such a model. Provided the ability to assign equal weights to different assertions. This is primarily due to the fact that the same text parsing may, in fact, correspond to several different assertions. For example, without limitation, the same invention, characterized by the same set of essential features, may achieve different technical results, each of which may be suitable and lead to a solution to a specific technical problem. This is due to the fact that the formulation of a technical problem is primarily related to the prior art, or even the prototype, in which the technical problem is formulated. In this regard, it is preferable that the classification model and clustering model employed provide the fundamental ability to classify a text parsing as corresponding to several assertions simultaneously.Moreover, for example, without limitation, the weights of the statements do not necessarily have to be the same for a decision to assign a text analysis to these statements; in fact, without limitation, it is possible to ensure the assignment of a text analysis to several statements at once by defining an acceptable proximity of the weights of the statements.
[0066] Thus, without limitation, when the corresponding classification model or clustering model is obtained, it becomes possible to provide a method for classifying text parsing, which at least involves feeding to the input of the obtained classification model and / or clustering model a text parsing containing at least a main entity and an associated terminal entity associated with the main entity, and classifying said text parsing as associated with any said statement obtained using any said method for forming a text corpus. In this case, without limitation, the main entity in this way is understood to be, for example, the main entity of the first segment or the said main entity of the second part of the part of the first segment to be parsed.
[0067] The classified text analyses can then be stored in a database of the system 200 for automated natural language processing, which will be described in detail below with reference to Fig. 4. For this purpose, without limitation, a method for forming a database is preferably provided, in which a plurality of said classified text analyses are identified, associated with one or more corresponding statements, and a database is formed containing at least a plurality of said text analyses, each of which is associated with one or more said statements. The database thus formed makes it possible to obtain text parses that correspond to any statement selected by the user. For this purpose, preferably, but without limitation, a method for generating a database query is provided, in which a database query is generated containing at least one statement obtained using any of the aforementioned methods for generating a text corpus; wherein the aforementioned database contains at least a plurality of text parses, each of which is associated with one or more of the aforementioned statements. For example, without limitation, a query may be provided containing the statement "accurate generation of a text corpus," in response to which at least one corresponding text parse will be selected. At the same time, depending on the statement, too many text parses may be selected, which may render the query irrelevant.In order to increase the relevance of a query, for example, but not limited to, the query may be supplemented with an additional statement, which will inevitably reduce the resulting sample of text analyses. Thus, preferably, but not limited to, a method for selecting text analyses is provided, in which a query is generated and sent to a database containing at least one statement obtained using any of the aforementioned methods for generating a text corpus, and at least one text analysis associated with said statement is obtained; wherein said database contains at least a plurality of text analyses, each of which is associated with one or more of the aforementioned statements.
[0068] Moreover, without limitation, the resulting method preferably provides new capabilities for generating natural language texts, in particular, but not limited to, texts such as patent document descriptions. Preferably, without limitation, a method for generating a text entry is provided, which at least one statement obtained using any of the aforementioned methods for generating a text corpus is obtained and a syntactically and semantically correct sentence is generated, including the statement or a derivative thereof. Moreover, without limitation, a text entry is understood to be a portion of a text, but not the entire text, and a text, accordingly, is understood to be a collection, including heterogeneous text entries.In this case, without limitation, the said assertion may be present in the generated text record either unchanged or in a derived form, i.e., when the assertion is subject to modifications and changes in order to ensure, at a minimum, the consistency of the generated sentence. In this case, preferably, without limitation, the resulting text records may be placed in a database of text records including approval. For this purpose, preferably, without limitation, a method for generating a database of text records including approval is provided, which at least involves obtaining a plurality of text records by means of said text record generation method and recording the obtained text records in the database; wherein the approval is obtained using any of said text corpus generation methods. For example, without limitation, such a database can subsequently be used as a source of standardized records when generating text in natural language.For this purpose, for example, but not limited to, a method for generating natural language text is provided, which involves at least obtaining a text entry from a database of text entries formed by said generation method, including an assertion, and then generating natural language text, which includes at least said text entry and text that is not said entry. More specifically, without limitation, the text that is not said entry may be, for example, but not limited to, text entered by the user.That is, without limitation, a method for generating text in a natural language is provided, which involves at least identifying a text entry in a natural language, including at least a main entity and at least an associated leaf entity linked thereto, and a statement linked to the main entity, after which a text entry including this statement is obtained from said database of text entries including the statement, and generating text in a natural language, including at least the identified text entry and the text entry obtained from said database. In this case, without limitation, the identified entry is a text entry entered by a user. More specifically, without limitation, the text entry entered by a user is a syntactically and semantically correct sentence.More specifically, but not limited to, the user-entered input constitutes an independent patent claim. This enables, for example, but not limited to, the generation of natural language text based on user-entered text and including a portion of text not actually entered by the user, i.e., generated automatically. Furthermore, without limitation, the text generation itself is accomplished using methods and means known in the art, such as NLP processors, which are accordingly not described in detail below.
[0069] Thus, preferably, without being limited, as shown in Fig. 2, a computer device 201 can be provided, in which, in the context of the present invention, at least one of or any combination of: a computer device for automated processing of text in natural language, a computer device for forming a text corpus, a computer device for pre-training, or training, or further training a classification model, a computer device for pre-training, or training, or further training a clustering model, a computer device for classifying text parsing, a computer device for forming a database, a computer device for forming a query in a database, a computer device for selecting text parsings can be embodied.Such a computing device 201 most typically comprises at least: one or more processors 2011; a memory 2012 containing program code that, when executed by the processor 2011, causes the processor 2011 to perform actions of any of the aforementioned methods described with reference to Fig. 1 for automated processing of text in a natural language, and / or a method for forming a text corpus, and / or a method for pre-training, or training, or further training of a classification model, and / or a method for pre-training, or training, or further training of a clustering model, and / or a method for classifying text parsing, and / or a method for forming a database, and / or a method for forming a query to a database, and / or a method for selecting text parsing.At the same time, such a computing device can be implemented as a thin client, which will mean that all, or at least most, computing operations are performed on the system server 202, which thus also contains at least a processor 2021 and memory 2022, which are thus essentially similar to, respectively, the processor 2011 and memory 2012.By way of example, but not limitation, the memory 2012, 2022 (machine-readable storage medium 2012, 2022) may include non-volatile memory (NVRAM); random access memory (RAM); read-only memory (ROM); electrically erasable programmable read-only memory (EEPROM); flash memory or other memory technologies; CDROM, digital versatile disc (DVD) or other optical or holographic storage media; magnetic cassettes, magnetic tape, magnetic disk storage device or other magnetic storage devices; as well as any other storage medium that can be used to store and encode the desired information. In this case, without limitation, the memory 2012, 2022 includes a storage medium based on a computer storage device in the form of volatile or non-volatile memory, or a combination thereof. In this case, not. Exemplary hardware devices include, but are not limited to, solid-state memory, hard disk drives, optical disk drives, and so forth. However, without limitation, the machine-readable storage medium 2012, 2022 (memory 2012, 2022) is non-transient (permanent, non-transitive), so it does not include a transient (transitive) propagating signal.In this case, without being limited, an exemplary environment can be stored in the memory 2012, 2022, in which, using computer commands or codes, including those stored in the memory 2022 of the server 202, a procedure for automated processing of text in a natural language can be carried out, and / or a procedure for forming a corpus of texts, and / or a procedure for pre-training, or training, or further training of a classification model, and / or a procedure for pre-training, or training, or further training of a clustering model, and / or a procedure for classifying text parsing, and / or a procedure for forming a database, and / or a procedure for forming a query to a database, and / or a procedure for selecting text parsing. In this case, without being limited, the computer device 201, when not a thin client, contains one or more processors 2011, which are intended to execute computer commands or codes stored in the memory 2012 of the device 201 in order to ensure the execution of the said procedures.In this case, without being limited, the server 202 may be essentially similar to the computer device 201, when it is not a thin client, and, accordingly, contain one or more processors 2021, which are designed to execute computer commands or codes stored in the memory 2022 of the server 202 in order to ensure the execution of the mentioned procedures. In this case, without being limited, the system 200 may also include a database (DB) 203. DB 203 may be, but is not limited to: a hierarchical DB, a network DB, a relational DB, an object DB, an object-oriented DB, an object-relational DB, a spatial DB, a combination of two or more of the listed DBs, and the like.In this case, without limitation, the DB 203 at least stores classified text parsing associated with the corresponding statements, and can also store data for analysis, classification models, clustering models and other information in the memory 2021, 2022 or in a suitable memory of another computing device associated with the computing device 201 and / or with the server 202, which may be, but is not limited to, a memory similar to any memory 2021, 2022, as shown earlier, and which can be accessed via the server 202. In addition, without limitation, the server 202 is provided, which, in addition to the functions described earlier, stores and facilitates the manipulation of computer commands or codes previously described in this document, which, accordingly, are not further. are described. In this case, without being limited to, the server 202, in addition to the functions described earlier, can provide for the regulation of data exchange in the system 200. In this case, without being limited to, the data exchange within the system 200 is carried out thanks to one or more data transmission networks 204. In this case, without being limited to, the data transmission networks 204 may include, but are not limited to, one or more local area networks (LAN) and / or wide area networks (WAN), or may represent an information telecommunications network, the Internet, or an Intranet, or a virtual private network (VPN), or a combination thereof, and the like. In this case, without being limited to, the server 202 also has the ability to provide a virtual computing environment for ensuring interaction between the components of the system. In this case, without being limited to, the network 204 serves to ensure interaction between the computer device 201, the server 402 and, optionally, the database 203.In this case, without limitation, the non-thin client computing device 201 and / or the server 202 may be directly connected to the database 203 using wired and wireless communication methods and techniques known in the art, which, accordingly, are not described in detail below, or, without limitation, the database 203 may be implemented in the memory 2012, 2022. In this case, without limitation, a suitable non-thin client computing device 201 may perform the role of the server 202 of the system 200 for other computing devices 201 that are thin clients. In this case, most typically, without limitation, the components of the computing device 201 and the components of the server 202 are interconnected, including by means of some data bus.
[0070] The present description of the implementation of the claimed invention demonstrates only particular embodiments and does not limit other embodiments of the claimed invention, since possible other alternative embodiments of the claimed invention, not going beyond the scope of the information set out in this application, should be obvious to a specialist in the given field of technology, having the usual qualifications, for whom the claimed invention is intended.
Claims
CLAUSES OF THE INVENTION 1. A method for generating a query to a database, executable by a processor of a computer device, which comprises generating a query to a database containing at least one statement obtained using a method for generating a text corpus; wherein said database contains at least a plurality of text parses, each of which is associated with one or more of said statements; wherein the method for generating a text corpus is a method executable by the processor of the computer device, in which at least by means of a method for automated processing of text in a natural language a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of a first segment, and a text corpus is formed from the obtained pairs of texts;wherein the method for automated processing of text in a natural language is a method executable by a processor of a computer device, which consists of at least performing the steps of: a step 101 of identifying a text in a natural language, including at least three segments; a step 102 of identifying segments; a step 103 of selecting at least a first segment and at least a second segment and / or at least a third segment of a text in a natural language; a step 104 of marking in the first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment, at least one part to be parsed; a step 105 of parsing the marked parts by means of semantic-syntactic parsing;step 106 of extracting from the part of the first segment subjected to semantic-syntactic parsing at least the main entity of the first segment and at least an associated entity associated with the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; step 107 of associating said statement with said main entity of the first segment.
2. The method according to claim 1, characterized in that the first segment, the second segment, and the third segment are pre-combined to obtain an identifiable text in a natural language.
3. The method according to claim 1, characterized in that the first segment, the second segment, and the third segment are pre-associated to obtain identifiable text in natural language.
4. Method according to item 1, characterized by the fact that the part to be parsed marked in the first segment is the first sentence in a natural language.
5. The method according to claim 4, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.
6. The method according to claim 1, characterized in that the said part to be parsed, marked in the first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing; wherein the main entity of the first segment and all associated entities related to it are extracted from the first part; wherein the main entity of the second part, and at least an associated entity related to the main entity of the second part, are extracted from the second part.
7. The method according to claim 6, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and for each associated entity at least the main entity of the associated entity and at least a nested associated entity linked to the main entity of the associated entity are extracted, wherein at least one of the nested associated entities is a nested associated end entity.
8. The method according to paragraph 7, characterized in that the main essence of the second part is associated with the main essence of the first segment.
9. The method according to any of paragraphs 1-8, characterized in that before forming the text corpus, at least one extracted statement is removed.
Citation Information
Patent Citations
Method and system of semantic processing text documents
RU2630427C2
Projecting Semantic Information from a Language Independent Syntactic Model
US20090326924A1
Text analysis using phrase definitions and containers
US20100250235A1
Semantic query language
US20170220673A1
Indexing system
US9501506B1