Method for generating a database

The method employs semantic-syntactic parsing and classification models to address the challenge of accurately extracting and classifying text parts in patent documents, enhancing automated text processing by generating precise text corpora.

WO2026063826A1PCT designated stage Publication Date: 2026-03-26KRAVCHENKO ARTEM ALEKSANDROVICH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-03-26

Smart Images

  • Figure 00000030_0000
    Figure 00000030_0000
  • Figure 00000031_0000
    Figure 00000031_0000
Patent Text Reader

Abstract

The proposed technical solution relates to methods for automated text processing and can be used in the generation of text corpora. The claimed invention solves the technical problem of creating a method and / or a computing device and / or a system and / or a machine-readable data carrier which overcome the disadvantages of the prior art and, in so doing, provide for the precise automated generation of a text corpus which can subsequently be used for pre-training, or training, or post-training classification models and / or clustering models.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD OF FORMING A DATABASE

[0001] AREA OF TECHNOLOGY

[0002] The proposed technical solution relates to methods of automated text processing and can be used in the formation of text corpora.

[0003] LEVEL OF TECHNOLOGY

[0004] Various methods for automated generation of texts in natural language are known, for example, those disclosed in the documents: US20070136321 A1 dated 06 / 14 / 2007, US20100205125A1 dated 08 / 12 / 20120, US20190377780A1 dated 12 / 12 / 2019, US20130080883A1 dated 03 / 28 / 2013, US10713443B1 dated 07 / 14 / 2020, US10747953B1 dated 08 / 18 / 2020, US11023662B2 dated 06 / 01 / 2021, US11341323B1 dated 05 / 24 / 2022. However, the known solutions do not allow the formation of texts in natural language with sufficient completeness, specific to a certain area, such as, for example, patent documents, in which, in addition to the actual formation of disclosures, it is also necessary to ensure the formation of special parts of the patent document, such as, for example, the level of technology (Background of the Invention), for which in most cases it is necessary to have a statement in the form of a disclosure of the technical problem solved by the invention or the technical result achieved by using the invention (technical effect (9.2.8 Case Law of the Boards of Appeal), (technical) result (MPEP 716.02(a))).

[0005] Patent No. US11593564B2 of 28.02.2023 (D1) discloses systems, methods, and storage media for extracting patent document templates from a patent corpus. Examples of implementation include: obtaining a patent corpus; obtaining one or more parameters; determining one or more subsets of the patent corpus by filtering the patent corpus based on one or more parameters; identifying one or more clusters of documents within individual ones of one or more subsets of the patent corpus; obtaining a patent document template corresponding to the first cluster of documents; and / or performing other operations. However, the solution known from D1 does not disclose any methods or techniques for extracting, in particular, statements concerning the technical result.

[0006] Patent US5774833A of June 30, 1998 (D2) discloses a method for processing patent text on a computer. The method includes determining the boundaries of portions of the patent text, loading at least one portion of the patent text into working memory, analyzing at least one portion of the patent text, and communicating the results to the user. Furthermore, alphanumeric data from a drawing can also be compared with the patent text. The D2 method can be combined with a word processor. The D2 method can recognize and report claim dependencies, specific patent text characteristics, and patent errors based on legal standards, standards of practice, USPTO standards, or even user preferences. However, the solution disclosed in D2 does not disclose any methods or techniques for extracting, in particular, claims concerning the technical result.

[0007] Thus, there is a problem of automated extraction of claims from specific natural language texts, such as, but not limited to, patent documents.

[0008] The solution known from D2 can be chosen as the closest analogue.

[0009] DISCLOSURE OF THE INVENTION

[0010] The technical problem solved by the claimed invention is the creation of a method and / or a computer device and / or a system and / or a machine-readable data carrier that does not have the disadvantages of analogues and thus ensures the accurate classification of specific texts.

[0011] The technical result achieved by implementing the claimed invention, in addition to fulfilling its intended purpose, is to eliminate the shortcomings of similar technologies, thereby ensuring accurate classification of specific texts. Another technical result is the expansion of the arsenal of technical means—methods for automated processing of natural language text.

[0012] The technical result is achieved due to the fact that a method for generating a database, executed by a processor of a computer device, is provided, in which a plurality of classified text parses are identified, associated with one or more corresponding statements, and a database is generated containing at least a plurality of said text parses, each of which is associated with one or more said statements; wherein the classified text parses are obtained by means of a method for classifying text parsing, executable by a processor of a computer device, in which at least a text parse is fed to the input of a classification model and / or a clustering model, containing at least a main entity and an associated leaf entity associated with the main entity, and said text parse is classified as being associated with any statement obtained using the method for generating a text corpus; wherein the method for generating a text corpus is a method executable by a processor of a computer device, in which at least by means of a method for automated processing of text in a natural language a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main entity of the first segment, and a text corpus is formed from the obtained pairs of texts;wherein the method for automated processing of text in a natural language is a method executable by a processor of a computer device, which consists of at least performing the following steps: a step of identifying text in a natural language, including at least three segments; a step of identifying segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of text in a natural language; a step of marking in the first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment, at least one part to be parsed; a step of parsing the marked parts by means of semantic-syntactic parsing;a step of extracting from the part of the first segment subjected to semantic-syntactic parsing at least the main entity of the first segment, and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step of associating said statement with said main entity of the first segment; wherein the classification model is obtained by means of a method of pre-training, or training, or additional training of a classification model, executable by a processor of a computer device, in which at least the classification model is trained to classify the parsed text as related to some; an assertion obtained using said method of formation, when a text parsing classification model receives at least a main entity and at least an associated end entity associated with the main entity, wherein training is carried out using a text corpus obtained by means of said method of forming a text corpus;wherein the clustering model is obtained by means of a pre-training method, or training, or additional training of a clustering model, executed by a processor of a computer device, in which, at least, the clustering model is trained to classify the text parsing as associated with any statement obtained using the said method of forming a text corpus, when receiving at the input of the clustering model a text parsing containing at least a main entity and at least an associated terminal entity associated with the main entity, wherein the training is carried out using a text corpus obtained by means of the said method of forming a text corpus.

[0013] BRIEF DESCRIPTION OF DRAWINGS

[0014] Illustrative embodiments of the present invention are described below in detail with reference to the accompanying drawings, which are incorporated herein by reference, and in which:

[0015] Fig. 1 shows, by way of example and not limitation, an exemplary flow chart of a method 100 for automated natural language text processing.

[0016] Fig. 2, by way of example and not limitation, shows an exemplary diagram of a system 200 for automated natural language processing.

[0017] IMPLEMENTATION OF THE INVENTION

[0018] In a preferred embodiment of the present invention, a method for automated processing of natural language text, executable by a processor of a computer device, is provided, which method comprises at least the following steps: a step of identifying natural language text, including at least three segments; a step of identifying segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of natural language text; a step of marking in the selected first segment only one a part to be parsed, and a step of marking up in the selected second segment and / or in the selected third segment, at least one part to be parsed; a step of parsing the marked up parts by means of semantic-syntactic parsing; a step of extracting from the part of the selected first segment subjected to semantic-syntactic parsing, at least the main entity of the first segment, and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated leaf entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing, at least one statement; a step of associating said statement with said main entity.

[0019] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-combined to obtain identifiable text in a natural language.

[0020] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-associated to obtain identifiable text in a natural language.

[0021] In a particular embodiment of the present invention, the said method is provided, characterized in that the part to be parsed marked in the selected first segment is the first sentence in a natural language.

[0022] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

[0023] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested within nested associated entities, until a nested associated entity is retrieved that does not have any entity nested within it.

[0024] In a particular embodiment of the present invention, the said method is provided, characterized in that the said part to be parsed, marked in the selected first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing.

[0025] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the selected first segment and all associated entities associated with it are extracted from the first part.

[0026] In a particular embodiment of the present invention, the said method is provided, characterized in that at least the main entity of the second part and at least an associated entity linked to the main entity of the second part are extracted from the second part.

[0027] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

[0028] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested in nested associated entities, until a nested associated entity is retrieved in which no entity is nested.

[0029] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the second part is associated with the main entity of the selected first segment.

[0030] In a particular embodiment of the present invention, the said method is provided, characterized in that for each extracted entity, lemmatization is at least partially performed.

[0031] In a particular embodiment of the present invention, the said method is provided, characterized in that the marking is carried out using a classification model and / or a clustering model.

[0032] In a particular embodiment of the present invention, the said method is provided, characterized in that the division of the part of the first segment to be analyzed is carried out using a classification model and / or a clustering model.

[0033] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computer device, is provided, in which, at least by means of said method of automated processing of text in natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the selected first segment, and a text corpus is formed from the obtained pairs of texts.

[0034] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computer device, is provided, in which, at least by means of said method of automated processing of text in a natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the second part of the part to be parsed marked up in the selected first segment, and a text corpus is formed from the obtained pairs of texts.

[0035] In another preferred embodiment of the present invention, there is provided a method executable by a processor of a computer device for pre-training, or training, or additional training of a classification model, in which at least the classification model is trained to classify a text parse as associated with any statement obtained using any of the mentioned methods for forming a text corpus, when receiving at the input of the classification model a text parse containing at least a main entity and at least an associated terminal entity associated with the main entity, wherein the training is carried out using a corpus of texts obtained by any of the mentioned methods of forming a corpus of texts.

[0036] In another preferred embodiment of the present invention, a method for pre-training, or training, or additional training of a clustering model, executable by a processor of a computer device, is provided, in which at least the clustering model is trained to classify a text analysis as associated with any statement obtained using any mentioned method for forming a text corpus, when receiving at the input of the clustering model a text analysis containing at least a main entity and at least an associated leaf entity associated with the main entity, wherein the training is carried out using a text corpus obtained by any mentioned method for forming a text corpus.

[0037] In another preferred embodiment of the present invention, a method for classifying a text analysis, executable by a processor of a computer device, is provided, which at least provides at the input of said classification model a text analysis containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text analysis as being associated with any statement obtained using any said method for forming a text corpus.

[0038] In another preferred embodiment of the present invention, a method for classifying text parsing, executable by a processor of a computing device, is provided, which at least provides at the input of said clustering model a text parsing containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text parsing as being associated with any statement obtained using any said method for forming a text corpus.

[0039] In another preferred embodiment of the present invention, a method for generating a database, executable by a processor of a computer device, is provided, in which a plurality of text parses classified by any of the aforementioned classification methods are identified, associated with one or more corresponding statements, and a database is generated containing, at least a set of the mentioned text analyses, each of which is related to one or more of the mentioned statements.

[0040] In another preferred embodiment of the present invention, a method for generating a query to a database, executable by a processor of a computer device, is provided, in which a query to a database is generated, containing at least one statement obtained using any of the mentioned methods for generating a corpus of texts; wherein the mentioned database contains at least a plurality of text parses, each of which is associated with one or more of the mentioned statements.

[0041] In another preferred embodiment of the present invention, a method for selecting text parses, executable by a processor of a computer device, is provided, in which at least a query is generated and sent to a database, containing at least some statement obtained using some method for forming a corpus of texts, and at least one text parse is obtained, associated with said statement; wherein said database contains at least a plurality of text parses, each of which is associated with one or more of said statements.

[0042] In another preferred embodiment of the present invention, there is provided a method executable by a processor of a computer device for generating a text entry, in which at least one statement obtained using any of the above is received and a syntactically and semantically correct sentence is generated, including the statement or a derivative thereof.

[0043] In another preferred embodiment of the present invention, a method for generating a database of text records, executable by a processor of a computer device, is provided, which includes at least obtaining a plurality of text records by means of said method for generating text records and recording the obtained text records in a database; the approval is obtained using any of said methods for generating a corpus of texts.

[0044] In another preferred embodiment of the present invention, there is provided a method for generating natural language text, executable by a processor of a computing device, wherein at least At least, a text entry is obtained from a database generated by said method of generating a database including approval of text entries, after which a text in natural language is generated, including at least said text entry and a text that is not said entry.

[0045] In another preferred embodiment of the present invention, a method for generating natural language text, executable by a processor of a computer device, is provided, which at least identifies a natural language text entry, including at least a main entity and at least an associated end entity linked to it, and a statement linked to the main entity, after which a text entry including this statement is obtained from a database formed by said method for forming a database of text entries including the statement, and generates natural language text, including at least the identified text entry, and a text entry obtained from said database

[0046] In another preferred embodiment of the present invention, there is provided a computer device for automated processing of natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of automated processing of natural language text.

[0047] In another preferred embodiment of the present invention, a computer device for generating a text corpus is provided, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any of the said methods for generating the text corpus.

[0048] In another preferred embodiment of the present invention, a computer device is provided for pre-training, or training, or re-training a classification model, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of pre-training, or training, or re-training a classification model.

[0049] In another preferred embodiment of the present invention, a computer device for pre-training, or training, or retraining of a clustering model, comprising at least: a processor; memory containing program code that, when executed by the processor, causes the processor to perform actions of any of the aforementioned methods of pretraining, or training, or retraining of the clustering model.

[0050] In another preferred embodiment of the present invention, there is provided a computer device for classifying text parsing, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said text parsing classification method.

[0051] In another preferred embodiment of the present invention, a computing device for generating a database is provided, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any said method of generating a database.

[0052] In another preferred embodiment of the present invention, a computing device is provided for generating a query to a database, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a query to a database.

[0053] In another preferred embodiment of the present invention, there is provided a computer device for selecting text parses, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for selecting text parses.

[0054] In another preferred embodiment of the present invention, there is provided a computer device for generating a text entry, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a text entry.

[0055] In another preferred embodiment of the present invention, a computing device is provided for generating a database of assertion-containing text records, comprising at least: a processor; a memory containing program code that, when executed the processor causes the processor to perform actions of any mentioned method of forming a database including the approval of text records.

[0056] In another preferred embodiment of the present invention, there is provided a computer device for generating natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating natural language text.

[0057] In another preferred embodiment of the present invention, there is provided a machine-readable storage medium, including a non-transitive machine-readable storage medium, containing program code that, when executed by a processor, causes the processor to perform the actions of any said method.

[0058] The following are embodiments of the present invention, revealing examples of its implementation in specific embodiments. However, the description itself is not intended to limit the scope of rights granted by this patent. Rather, it should be understood that the claimed invention may also be implemented in other ways that incorporate different elements and conditions, or combinations of elements and conditions similar to those described herein, in combination with other existing and future technologies.

[0059] In the present description, a patent document shall mean, without limitation, a patent or a patent application, that is, as a rule, a specific text consisting of at least three segments: the first segment, which is the patent claims (patent claims), the second segment, which is the description (description, specification), and the third segment, which is the abstract (abstract), wherein, without limitation, the numbering of the segments above is not given in the order of their sequence in the text in natural language, but only for the simplicity of presentation and association in the present document, as will be shown below. Preferably, without limitation, the formulas of patent documents do not relate to new chemical compounds and other similar objects created for the first time, since with respect to such objects the statement of the technical result does not require disclosure.

[0060] Fig. 1 shows, by way of example and not limitation, an exemplary diagram of the method 100 for automated text processing on natural language. The method 100 consists of at least performing the following steps: a step 101 of identifying a natural language text comprising at least three segments; a step 102 of identifying the segments; a step 103 of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step 104 of marking up in the selected first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment at least one part to be parsed; a step 105 of parsing the marked parts by means of semantic-syntactic parsing;step 106 of extracting from the part of the selected first segment subjected to semantic-syntactic parsing at least the main entity of the first segment and at least an associated entity associated with the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement;step 107 of associating said statement with said main entity. For example, but not limited to, step 101 of identifying natural language text ensures reliable identification of natural language text, for example, but not limited to, a patent document. For example, but not limited to, the text segments of a patent document may be represented by separate natural language texts, for which, if necessary, they are first associated, for example, by assigning them some relatedness features, for example, unique identifiers, the decoding of which makes it possible to determine the relatedness of the segments with each other. Such methods and techniques for associating entities in data storage systems are widely known in the prior art and, accordingly, are not described in detail below. For example, but not limited to, a patent document may represent a single text, typically having sequential page numbers;in this case, preliminary association of segments with each other is not required.

[0061] Preferably, but not limited to, step 102 includes identifying segments of natural language text; for example, but not limited to, in the case where the natural language text is a patent document, identifying at least the patent claims, which are the first segment in the context of the present invention, the description, which is the second segment in the context of the present invention, and the abstract, which is the third segment in the context of the present invention. Without limitation, in the case where the natural language text is a non-patent document, the identification of segments is carried out in a similar manner. In this case, without limitation, it should be primarily assumed that the first segment includes a set of entities, while the second and third segments include at least one or more statements characteristic of the said set of entities. An example of a suitable non-patent document may be, for example, but not limited to, a scientific article, in which the role of the first segment may be assigned to the abstract, and the roles of the second and third segments, respectively, to the main text and conclusion.

[0062] Preferably, but not limited to, step 103 comprises selecting the identified segments, including, but not limited to, selecting the first segment as comprising the set of entities and one of the second segment or the third segment. In this case, but not limited to, the choice of which of the second segment or the third segment to select depends on the probability with which the claim will be present in a particular segment. More specifically, but not limited to, the presence of the claim in a particular segment will generally depend on the approach to the requirements imposed on the natural language text. For example, but not limited to, when the natural language text is a patent document, then, depending on the patent system in which it is formed, the third segment (abstract) will at least contain a claim about the technical result.For example, but not limited to, Russian patent documents are typically published with an abstract formulating the technical result. Moreover, as of the filing date of the present application, the preparation of the abstract when deciding whether to grant a patent rests with the examiner, who is competent in formulating the technical result, thus eliminating inaccuracies in its presentation. Meanwhile, in US patent documents, the abstract is usually prepared by the applicant themselves, resulting in abstracts in US patent documents lacking a clear structure, and the presence of a statement of the technical result in them is not guaranteed. Thus, for example, without limitation, when choosing between the second and third segments. In the case of the United States, it would be more efficient to select the second segment, since at least the Background of the Invention or Brief Summary of Invention sections are likely to contain a claim about the technical result, as required by MPEP 608.01 (c) and 608.01 (d). At the same time, for example, but not limited to, for each natural language text, a particular segment can be selected based on the presence of the claim itself. For this purpose, for example, but not limited to, for each natural language text, a semantic-syntactic analysis can be performed, for example, using a semantic parser, the technology of which is widely represented in the prior art and, accordingly, is not described below; based on the results of the semantic-syntactic analysis, it becomes possible to determine the specific part of the natural language text containing the sought-after claim.Preferably, without limitation, the first segment is preliminarily excluded from the natural language text, as it obviously does not contain assertions, and the presence of which during semantic-syntactic analysis can distort the results.

[0063] Preferably, but not limited to, step 104 involves annotating the parts of the selected segments. Preferably, but not limited to, only one part of the selected first segment to be parsed is annotated, namely, the part containing the set of entities. For example, without limitation, in the case where the natural language text is a patent document, the first sentence, i.e., the first claim, is selected as the annotated part, as it is most likely to contain the required set of entities, i.e., the set of essential features. Furthermore, without limitation, there are situations where the set of entities (features) in the first claim, in addition to the essential features, also includes nonessential features.In this case, such a selected first segment may be subjected to additional analysis after finding the technical result statement in the second or third segment and verifying which features do not affect the achievement of the technical result. Based on the results of such additional analysis, the marked portion is cleared of unnecessary entities. In this case, preferably, but not limited to, the marked and optionally cleared portion of the selected first segment to be analyzed may be subjected to analysis to determine the presence of the generic part (first part) within it. the distinctive part (second part); in this case, for example, but not limited to, it may be sufficient to determine the linking word separating the generic part from the distinctive part. At the same time, not every patent system requires that an independent claim be composed with the use of a distinctive part, which results in the generic part not being clearly distinguished and not separated by a linking word; in such a case, for example, but not limited to, a preliminary semantic-syntactic analysis of the entire first segment may be performed in order to identify, at a minimum, the main entity of the first segment and the main entities of the entities associated with the main entity of the first segment, which will be discussed in detail below.Typically, but not limited to, when the natural language text is a patent document, the primary entity will be defined as a generic concept, and the associated entities will be individual features, such as, but not limited to, method steps or product parts. The resulting set of primary entities and associated entities can be subjected to a familiarity analysis, which can determine whether the entire set is generic or only a subset of the features constitutes a generic part. For example, but not limited to, the familiarity analysis can be performed automatically, based on a pre-prepared text corpus, for each of which a first segment has been identified and a semantic-syntactic analysis has been performed on the part to be analyzed.In this case, without limitation, the separation can be carried out both manually and automatically, for example, using a pre-trained classification model and / or clustering model. In this case, without limitation, in the selected second segment or in the selected third segment, the marking is carried out in such a way as to obtain the desired statement with the highest probability. Typically, the presentation of the desired statement is preceded by a certain text structure, or the presentation of the desired statement includes certain keywords. Without limitation, several parts of the selected second segment or the selected third segment to be analyzed can be marked in this manner. In this case, without limitation, the marking can be carried out both manually and automatically, for example, using a pre-trained classification model and / or clustering model. Preferably, without limitation, within the framework of step 105, the marked parts are analyzed by semantic-syntactic analysis using, for example, but not limited to, the mentioned syntactic parser.

[0064] Preferably, but not limited to, step 106 comprises extracting from the parsed portions of the first segment at least the main entity of the first segment and at least an associated entity linked to the main entity of the first segment, wherein at least one of the extracted associated entities is an associated terminal entity. For example, without limitation, in the case where the first segment is a patent claim, the first independent claim will be subjected to semantic-syntactic parsing; in such a case, the main entity will be a generic concept that is the root of the parsing tree, the associated entities will be the features of the technical solution that are nodes of the parsing tree, and at least one associated entity will be an associated terminal entity, that is, it will be a feature that has only one edge, that is, it will be a terminal node (leaf) of the parsing tree.Moreover, the parse tree can be constructed using the Abstract Syntax Tree (AST) principle, meaning it can be stripped of all irrelevant features, as well as repetitions, anaphora, and other elements that don't affect the claim—in the case of patent claims, those that can't affect the technical result. Examples of redundant elements obtained after parsing include, but are not limited to, commas or other separators, which are useful directly during parsing, such as as links or boundaries, but are redundant for the text corpus when used, for example, to train a classification model and / or clustering model.Furthermore, for example, but not limited to, when the first segment has been divided into generic and distinctive parts, the principal entity of the first segment, after the aforementioned analysis, will be found in the first part, that is, in the generic part, along with the entities associated with it from the generic part; while, for example, but not limited to, the second part, which is the distinctive part, will then have the principal entity of the second part and the entities associated with it, which, nevertheless, are also all associated with the principal entity of the first segment, since all individual features revealed by analysis directly or indirectly belong to only one generic concept. Moreover, for example, but not limited to, each associated one thus extracted. An entity may be iteratively subjected to its own semantic-syntactic analysis, during which the corresponding principal entities of the associated entities and the associated entities of the associated entities will be discovered. That is, each associated entity may have multiple nested associated entities, up to one or more associated terminal entities that can no longer be subjected to semantic-syntactic analysis because they are too simple in their features. In this case, for example, but not limited to, lemmatization may be performed to reduce words to their original forms. Furthermore, for each entity, including the principal entity, as well as any associated entities, a generalization operation may be performed, in which the entity can be elevated to the level of a method or product that corresponds to a certain function determined by the entity's features.In this case, preferably, but not limited to, at least one statement is extracted from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing. For example, without limitation, in the case where the natural language text is a patent document, the semantic-syntactic parsing of paragraph 0010 of the present description in its original version will yield two statements: "accurate generation of a text corpus" and "automated generation of a text corpus." For example, without limitation, in the same case, the semantic-syntactic parsing of paragraph 0011 may yield three statements: "implementation of the purpose - generation of a text corpus," "accurate generation of a text corpus," and "automated generation of a text corpus."In this case, preferably, but not limited to, at step 107, at least one extracted statement is associated with the main entity of the first segment, and, thus, the statement is associated with any associated entities related to the main entity of the first segment. In this case, for example, but not limited to, some of the obtained statements, such as, for example, "implementation of purpose", do not have practical use, since they cannot be subsequently used in the text corpus for training a classification model and / or a clustering model, for example, for the purpose of predicting a statement corresponding to a new set of entities, since this will lead to the fact that any classifier based on the classification model and / or. The clustering model will assign any set of entities to such a statement, since such a statement is inherent in all resulting text parses. For this reason, it is preferable, but not limited to, to clean up the resulting "statement-main entity" pairs, removing redundant statements.

[0065] In this manner, preferably, without limitation, it becomes possible to obtain a corpus of texts, for which purpose, by means of said method 100, in step 107, by associating said statement with said main entity of the first segment (and, as a consequence, with all associated entities related to the main entity), a plurality of pairs of texts are obtained, each of which includes at least one statement related to the main entity of the first segment, and a corpus of texts is formed from the obtained pairs of texts.Alternatively, without limitation, it becomes possible to obtain a corpus of texts, which by means of said method 100 in step 107 by associating said statement with said main entity of the second part of the part to be parsed marked in the first segment (and, as a consequence, with all associated entities related to the main entity of the second part of the part to be parsed marked in the first segment) a plurality of pairs of texts are obtained, each of which includes at least one statement related to the main entity of the second part of the part to be parsed marked in the first segment, and a corpus of texts is formed from the obtained pairs of texts.Such resulting text corpora, for example, but not limited to, may be used, either jointly or separately, for pre-training, or training, or additional training of a classification model and / or a clustering model, in which, at least, the classification model and / or the clustering model are trained to classify a text parse as related to any said statement, upon receipt at the input of the classification model and / or the clustering model of a text parse containing, at least, a main entity and an associated terminal entity associated with the main entity, wherein the training is carried out on one of the said text corpora, jointly or separately. In this case, without limitation, the main entity in this way is understood to be, for example, the main entity of the first segment or the said main entity of the second part of the part of the first segment to be parsed.To train a classification model and / or a clustering model, for example, but not limited to, a sample is provided from a text corpus, after which the training of the classification model and / or is provided. clustering models on such a sample. In this case, for example, but not limited to, the classification model is based on one of the following: logistic regression, support vector machine, decision trees, random forest, naive Bayes, k-nearest neighbors, neural network, boosting, gradient boosting, bagging, group argument method; in this case, without limitation, the clustering model is based on one of the following: k-means, density-based spatial clustering of applications with noise (DBSCAN), hierarchical clustering, spectral clustering. In this case, any suitable known or future known classification model and / or clustering model may be used, provided that the choice of the model requires that it provides the ability to obtain equal weights for different statements.This is primarily due to the fact that the same text parsing may, in reality, correspond to several different claims. For example, without limitation, the same invention, characterized by the same set of essential features, may achieve different technical results, each of which may be suitable and lead to a solution to a specific technical problem. This is due to the fact that the formulation of a technical problem is primarily related to the prior art, or even the prototype, in which the technical problem is formulated. Therefore, it is preferable that the classification model and clustering model employed ensure the fundamental possibility of classifying a text parsing as corresponding to several claims simultaneously.Moreover, for example, without limitation, the weights of the statements do not necessarily have to be the same for a decision to assign a text analysis to these statements; in fact, without limitation, it is possible to ensure the assignment of a text analysis to several statements at once by defining an acceptable proximity of the weights of the statements.

[0066] Thus, without being limited, when the corresponding classification model or clustering model is obtained, it becomes possible to provide a method for classifying text parsing, in which, at least, a text parsing containing at least the main entity and is fed to the input of the obtained classification model and / or clustering model An associated terminal entity associated with the main entity, and classifying said text parsing as related to any said statement obtained using any said method of forming a text corpus. In this case, without limitation, the main entity in this manner may be understood to mean, for example, the main entity of the first segment or the main entity of the second part of the first segment to be parsed.

[0067] The classified text analyses can subsequently be stored in the database of the automated natural language processing system 200. For this purpose, without limitation, a method for creating a database is preferably provided, which involves identifying a plurality of said classified text analyses, associating them with one or more corresponding statements, and creating a database containing at least a plurality of said text analyses, each of which is associated with one or more said statements. From the database thus created, it becomes possible to obtain text analyses corresponding to any statement selected by the user.For this purpose, preferably, but without limitation, a method for generating a database query is provided, which comprises generating a database query containing at least one assertion obtained using any of the aforementioned methods for generating a text corpus; wherein the aforementioned database contains at least a plurality of text parses, each of which is associated with one or more of the aforementioned assertions. For example, without limitation, a query may be provided containing the assertion "accurate generation of a text corpus," in response to which at least one corresponding text parse will be selected. At the same time, depending on the assertion, too many text parses may be selected, which may render the query irrelevant. In order to increase the relevance of the query, for example, without limitation, the query may be supplemented with an additional assertion, which will inevitably reduce the resulting selection of text parses.Thus, preferably, without being limited, a method for selecting text parses is provided, in which a query is generated and sent to a database containing at least some statement obtained using any of the mentioned methods for forming a corpus of texts, and at least one text parse is obtained. associated with said statement; wherein said database contains at least a plurality of text analyses, each of which is associated with one or more of said statements.

[0068] Moreover, without limitation, the resulting method preferably provides new capabilities for generating natural language texts, in particular, but not limited to, texts such as patent document descriptions. Preferably, without limitation, a method for generating a text entry is provided, which at least one statement obtained using any of the aforementioned methods for generating a text corpus is obtained and a syntactically and semantically correct sentence is generated, including the statement or a derivative thereof. Moreover, without limitation, a text entry is understood to be a portion of a text, but not the entire text, and a text, accordingly, is understood to be a collection, including heterogeneous text entries.In this case, without limitation, the said assertion may be present in the generated text record both in unchanged form and in a derived form, i.e., when the assertion is subject to modifications and changes in order to ensure, at least, the consistency of the generated sentence. In this case, preferably, without limitation, the obtained text records can be placed in a database of text records including the assertion. For this purpose, preferably, without limitation, a method for creating a database of text records including the assertion is provided, in which, at least, a plurality of text records are obtained by means of the said method of generating a text record and the obtained text records are written into the database; wherein the assertion is obtained using any of the said methods for forming a text corpus.For example, but not limited to, such a database can subsequently be used as a source of standardized records when generating natural language text. For this purpose, for example, but not limited to, a method for generating natural language text is provided, which comprises at least obtaining a text record from a database of text records comprising an assertion, generated by said generation method, and then generating natural language text, which includes at least said text record and text that is not said record. More specifically, but not limited to, the text that is not said record may be, for example, not. without limitation, text entered by the user. That is, without limitation, a method for generating text in a natural language is provided, which involves at least identifying a text entry in a natural language, including at least a main entity and at least an associated leaf entity linked thereto, and a statement linked to the main entity, after which a text entry including this statement is obtained from said database of text entries including the statement, and generating text in a natural language, including at least the identified text entry and the text entry obtained from said database. In this case, without limitation, the identified entry is a text entry entered by the user. More specifically, without limitation, the text entry entered by the user is a syntactically and semantically correct sentence.More specifically, but not limited to, the user-entered input constitutes an independent patent claim. This enables, for example, but not limited to, the generation of natural language text based on user-entered text and including a portion of text not actually entered by the user, i.e., generated automatically. Furthermore, without limitation, the text generation itself is accomplished using methods and means known in the art, such as NLP processors, which are accordingly not described in detail below.

[0069] Thus, preferably, without being limited, as shown in Fig. 2, a computing device 201 may be provided, in which, in the context of the present invention, at least one of or any combination of: a computing device for automated natural language text processing, a computing device for generating a text corpus, a computing device for pre-training, or training, or re-training a classification model, a computing device for pre-training, or training, or re-training a clustering model, a computing device for classifying text parsing, a computing device for generating a database, a computing device for generating a query to the database, a computing device for selecting text parsings may be embodied. Such computing device 201 most typically comprises at least: one or more processors 2011; memory 2012 containing code a program which, when executed by the processor 2011, causes the processor 2011 to perform actions of any of the aforementioned methods described with reference to Fig. 1 for automated processing of natural language text, and / or a method for generating a text corpus, and / or a method for pre-training, or training, or further training of a classification model, and / or a method for pre-training, or training, or further training of a clustering model, and / or a method for classifying text parsing, and / or a method for generating a database, and / or a method for generating a query to a database, and / or a method for selecting text parsing.At the same time, such a computing device can be implemented as a thin client, which will mean that all, or at least most, computing operations are performed on the system server 202, which thus also contains at least a processor 2021 and memory 2022, which are thus essentially similar to, respectively, the processor 2011 and memory 2012.By way of example, but not limitation, the memory 2012, 2022 (the computer-readable storage medium 2012, 2022) may include non-volatile memory (NVRAM); random access memory (RAM); read-only memory (ROM); electrically erasable programmable read-only memory (EEPROM); flash memory or other memory technologies; CDROM, digital versatile disc (DVD) or other optical or holographic storage media; magnetic cassettes, magnetic tape, magnetic disk storage device or other magnetic storage devices; as well as any other storage medium that can be used to store and encode the desired information. In this case, without being limited, the memory 2012, 2022 includes a storage medium based on a computer storage device in the form of volatile or non-volatile memory, or a combination thereof.In this case, without limitation, exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, and so on. In this case, without limitation, the machine-readable storage medium 2012, 2022 (memory 2012, 2022) is not temporary (permanent, non-transitive), so that it does not include a temporary (transitive) propagating signal. In this case, without limitation, the memory 2012, 2022 can store an exemplary environment in which, using computer commands or codes, including those stored in the memory 2022 of the server 202, the procedure of automated processing of text in natural language and / or the procedure of forming a text corpus and / or the procedure can be carried out. pre-training, or training, or further training of a classification model, and / or a procedure for pre-training, or training, or further training of a clustering model, and / or a procedure for classifying text parsing, and / or a procedure for forming a database, and / or a procedure for forming a query to a database, and / or a procedure for selecting text parsing. In this case, without limitation, the computer device 201, when not a thin client, contains one or more processors 2011, which are intended to execute computer commands or codes stored in the memory 2012 of the device 201 for the purpose of ensuring the execution of the said procedures. In this case, without limitation, the server 202 can be essentially similar to the computer device 201, when not a thin client, and, accordingly, contain one or more processors 2021, which are intended to execute computer commands or codes stored in the memory 2022 of the server 202 for the purpose of ensuring the execution of the said procedures.In this case, without being limited, the system 200 may also include a database (DB) 203. The DB 203 may be, but is not limited to: a hierarchical DB, a network DB, a relational DB, an object DB, an object-oriented DB, an object-relational DB, a spatial DB, a combination of two or more of the above DBs, and the like. In this case, without being limited to, the DB 203 at least stores classified text parsing associated with the corresponding statements, and can also store data for analysis, classification models, clustering models and other information in the memory 2021, 2022 or in a suitable memory of another computing device associated with the computing device 201 and / or with the server 202, which may be, but is not limited to, a memory similar to any memory 2021, 2022, as shown earlier, and which can be accessed via the server 202.In addition, but not limited to, a server 202 is provided which, in addition to the previously described functions, stores and facilitates the manipulation of computer commands or codes previously described in this document, which, accordingly, are not further described. Moreover, without limitation, the server 202, in addition to the previously described functions, can provide for the regulation of data exchange in the system 200. Moreover, without limitation, the data exchange within the system 200 is carried out thanks to one or more data networks 204. Moreover, without limitation, the data networks 204 may include, but are not limited to, one or more local area networks (LAN) and / or global networks. (WAN), or may be an Internet information and telecommunications network, or an Intranet, or a virtual private network (VPN), or a combination thereof, and the like. In this case, without limitation, the server 202 also has the ability to provide a virtual computing environment to ensure interaction between the system components. In this case, without limitation, the network 204 serves to ensure interaction between the computing device 201, the server 402 and, optionally, the database 203. In this case, without limitation, the non-thin client computing device 201 and / or the server 202 can be connected to the database 203 directly, using wired and wireless communication methods and techniques known in the art, which, accordingly, are not described in detail below, or, without limitation, the database 203 can be implemented in the memory 2012, 2022.In this case, without limitation, a suitable non-thin client computing device 201 can perform the role of a server 202 of the system 200 for other computing devices 201 that are thin clients. In this case, most typically, without limitation, the components of the computing device 201 and the components of the server 202 are interconnected, including via some kind of data bus.

[0070] The present description of the implementation of the claimed invention demonstrates only particular embodiments and does not limit other embodiments of the claimed invention, since possible other alternative embodiments of the claimed invention, not going beyond the scope of the information set out in this application, should be obvious to a specialist in the given field of technology, having the usual qualifications, for whom the claimed invention is intended.

Claims

CLAUSES OF THE INVENTION 1. A method for generating a database, executable by a processor of a computer device, which involves identifying a plurality of classified text parses, associating them with one or more corresponding statements, and generating a database containing at least a plurality of said text parses, each of which is associated with one or more said statements; wherein the classified text parses are obtained by means of a method for classifying text parses, executable by a processor of a computer device, which involves at least feeding to the input of a classification model and / or a clustering model a text parse containing at least a main entity and an associated leaf entity associated with the main entity, and classifying said text parse as associated with any statement obtained using the method for generating a corpus of texts;wherein the method for forming a text corpus is a method executable by a processor of a computer device, in which at least by means of a method for automated processing of text in a natural language a plurality of pairs of texts are obtained, each of which includes at least one statement related to the main essence of the first segment, and a text corpus is formed from the obtained pairs of texts; wherein the method for automated processing of text in a natural language is a method executable by a processor of a computer device, which consists of at least performing the steps of: a step 101 of identifying a text in a natural language, including at least three segments; a step 102 of identifying segments; a step 103 of selecting at least a first segment and at least a second segment and / or at least a third segment of a text in a natural language;step 104 of marking in the first segment only one part to be disassembled, and the step of marking in the selected second segment and / or in the selected third segment, at least one part to be disassembled; a step 105 of parsing the marked parts by means of semantic-syntactic parsing; a step 106 of extracting from the part of the first segment subjected to semantic-syntactic parsing at least the main entity of the first segment and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step 107 of associating said statement with said main entity of the first segment;wherein the classification model is obtained by means of a pre-training method, or training, or additional training of a classification model, executed by a processor of a computer device, in which, at least, the classification model is trained to classify a text analysis as associated with any statement obtained using the said method of formation, when a text analysis containing at least a main entity and at least an associated terminal entity associated with the main entity is received at the input of the classification model, wherein the training is carried out using a text corpus obtained by means of the said method of forming a text corpus;wherein the clustering model is obtained by means of a pre-training method, or training, or additional training of a clustering model, executed by a processor of a computer device, in which, at least, the clustering model is trained to classify the text parsing as associated with any statement obtained using the said method of forming a text corpus, when receiving at the input of the clustering model a text parsing containing at least a main entity and at least an associated terminal entity associated with the main entity, wherein the training is carried out using a text corpus obtained by means of the said method of forming a text corpus.

2. The method according to paragraph 1, characterized in that the first segment, the second segment and the third segment are pre-combined to obtain identifiable text in a natural language.

3. The method according to claim 1, characterized in that the first segment, the second segment, and the third segment are pre-associated to obtain identifiable text in a natural language.

4. The method according to paragraph 1, characterized in that the part to be parsed marked in the first segment is the first sentence in a natural language.

5. The method according to claim 4, characterized in that each extracted associated entity is subjected to semantic-syntactic analysis and for each associated entity at least the main entity of the associated entity and at least a nested associated entity linked to the main entity of the associated entity are extracted, wherein at least one of the nested associated entities is a nested associated end entity.

6. The method according to claim 1, characterized in that the said part to be parsed, marked in the first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing; wherein the main entity of the first segment and all associated entities related to it are extracted from the first part; wherein the main entity of the second part, and at least an associated entity related to the main entity of the second part, are extracted from the second part.

7. The method according to claim 6, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and for each associated entity at least the main entity of the associated entity and at least a nested associated entity linked to the main entity of the associated entity are extracted, wherein at least one of the nested associated entities is a nested associated end entity.

8. The method according to paragraph 7, characterized in that the main essence of the second part is associated with the main essence of the first segment.

9. The method according to any of paragraphs 1-8, characterized in that before forming the text corpus, at least one extracted statement is removed.

Citation Information

Patent Citations

  • Method of forming and structuring an electronic database

    RU2696295C1

  • System and method for synthetic form image generation

    US10546054B1

  • Database generation from natural language text documents

    US20230037077A1

  • Database query generation using natural language text

    US20230096857A1

  • Systems, methods, and media for retrieving an entity from a data table using semantic search

    US20230259507A1