Computing device for classifying a text parse

The computer device and method address the challenge of accurately extracting and classifying statements in patent documents by employing semantic-syntactic parsing and model training, enabling effective generation of structured text corpora and natural language outputs.

WO2026063830A1PCT designated stage Publication Date: 2026-03-26KRAVCHENKO ARTEM ALEKSANDROVICH
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing methods for automated text processing, particularly in patent documents, fail to accurately extract and classify specific statements such as technical results due to insufficient completeness and lack of methods for extracting claims.

Method used

A computer device and method for text classification and parsing, utilizing semantic-syntactic parsing to identify and associate main entities and associated entities within segments of natural language texts, forming a text corpus through iterative parsing and training of classification or clustering models to accurately classify and generate relevant statements.

Benefits of technology

Ensures accurate classification and extraction of specific statements from natural language texts, expanding the capabilities of automated text processing by providing a robust method for generating structured text corpora and natural language outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000028_0000
    Figure 00000028_0000
  • Figure 00000029_0000
    Figure 00000029_0000
Patent Text Reader

Abstract

The invention relates to methods for automated text processing and can be used in the generation of text corpora. A computing device for classifying a text parse comprises a processor and a memory containing a program code that causes the processor to perform the actions of a method for classifying a text parse. A text parse containing a main entity and, associated with said main entity, an associated end entity, is input into a classification model. The text parse is classified as being related to an assertion obtained using a method for generating a text corpus. Using a natural language text processing method, a plurality of text pairs is obtained, each of which contains an assertion related to a main entity of a first segment, and a text corpus is generated from the obtained text pairs. The invention provides for more precise automated generation of a text corpus for use in pre-training, or training, or post-training classification models and / or clustering models.
Need to check novelty before this filing date? Find Prior Art

Description

COMPUTER DEVICE FOR CLASSIFICATION AND ANALYZING OF TEXT

[0001] AREA OF TECHNOLOGY

[0002] The proposed technical solution relates to methods of automated text processing and can be used in the formation of text corpora.

[0003] LEVEL OF TECHNOLOGY

[0004] Various methods for automated generation of texts in natural language are known, for example, those disclosed in the documents: US20070136321A1 dated 14.06.2007, US20100205125A1 dated 12.08.20120, US20190377780A1 dated 12.12.2019, US20130080883A1 dated 28.03.2013, US10713443B1 dated 14.07.2020, US10747953B1 dated 18.08.2020, US11023662B2 dated 01.06.2021, US11341323B1 dated 24.05.2022. However, the known solutions do not allow the formation of texts in natural language with sufficient completeness, specific to a certain area, such as, for example, patent documents, in which, in addition to the actual formation of disclosures, it is also necessary to ensure the formation of special parts of the patent document, such as, for example, the level of technology (Background of the Invention), for which in most cases it is necessary to have a statement in the form of a disclosure of the technical problem solved by the invention or the technical result achieved by using the invention (technical effect (9.2.8 Case Law of the Boards of Appeal), (technical) result (MPEP 716.02(a))).

[0005] Patent No. US11593564B2 of February 28, 2023 (D1) discloses systems, methods, and storage media for extracting patent document templates from a patent corpus. Examples of implementation include: obtaining a patent corpus; obtaining one or more parameters; determining one or more subsets of the patent corpus by filtering the patent corpus based on one or more parameters; identifying one or more clusters of documents within individual ones of one or more subsets of the patent corpus; obtaining a patent document template corresponding to the first cluster of documents; and / or performing other operations. However, the solution known from D1 does not disclose any methods or techniques for extracting, in particular, statements concerning the technical result.

[0006] From the patent US5774833A of 30.06.1998 (D2) a method for processing patent text on a computer is known, which includes determining the boundaries of parts of the patent text, loading at least one of the parts of the patent text into the working Memory, analysis of at least one portion of the patent text, and reporting of the results to the user. The alphanumeric data of the drawing can also be compared with the patent text. Method D2 can be combined with a word processor. Method D2 makes it possible to recognize and report claim dependencies, specific characteristics of the patent text, and patent errors based on legal standards, standards of practice, USPTO standards, or even user preferences. However, the prior art solution D2 does not disclose any methods or techniques for extracting, in particular, claims concerning the technical result.

[0007] Thus, there is a problem of automated extraction of claims from specific natural language texts, such as, but not limited to, patent documents.

[0008] The solution known from D2 can be chosen as the closest analogue.

[0009] DISCLOSURE OF THE INVENTION

[0010] The technical problem solved by the claimed invention is the creation of a method and / or a computer device and / or a system and / or a machine-readable data carrier that does not have the disadvantages of analogues and thus ensures the accurate classification of specific texts.

[0011] The technical result achieved by implementing the claimed invention, in addition to fulfilling its intended purpose, is to eliminate the shortcomings of similar technologies, thereby ensuring accurate classification of specific texts. Another technical result is the expansion of the arsenal of technical means—methods for automated processing of natural language text.

[0012] The technical result is achieved due to the fact that a computer device for classifying text parsing is provided, which device contains, at least: a processor; a memory containing program code, which, when executed by the processor, causes the processor to perform the actions of a method for classifying text parsing, in which, at least, a text parsing is fed to the input of the classification model, containing, at least, a main entity and an associated terminal entity associated with the main entity, and the said text parsing is classified as associated with any statement obtained using the method for generating a corpus of texts; wherein the method for generating a corpus of texts is executable by the processor of the computer device in a manner in which, at least by means of a method for automated processing of text in a natural language a plurality of text pairs are obtained, each of which includes at least one statement related to the main essence of the first segment, and a text corpus is formed from the obtained text pairs; wherein the method for automated processing of text in a natural language is a method executable by a computer processor and consists of at least performing the following steps: a step of identifying a text in a natural language, including at least three segments; a step of identifying segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of a text in a natural language; a step of marking in the first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment at least one part to be parsed; a step of parsing the marked parts by means of semantic syntactic parsing;a step of extracting from the part of the first segment subjected to semantic-syntactic parsing at least the main entity of the first segment, and at least an associated entity associated with the main entity of the first segment, wherein at least one of the associated entities is an associated end entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step of associating said statement with said main entity of the first segment;wherein the classification model is obtained by means of a pre-training method, or training, or additional training of a classification model, executed by a processor of a computer device, in which, at least, the classification model is trained to classify a text analysis as associated with any statement obtained using the said method of formation, when a text analysis containing at least a main entity and at least an associated terminal entity associated with the main entity is received at the input of the classification model, wherein the training is carried out using a text corpus obtained by means of the said method of forming a text corpus.

[0013] BRIEF DESCRIPTION OF DRAWINGS

[0014] Illustrative embodiments of the present invention are now described in detail with reference to the accompanying drawings, which are incorporated herein by reference, and in which:

[0015] Fig. 1, by way of example and not limitation, shows an exemplary flow chart of a method 100 for automated natural language text processing.

[0016] Fig. 2, by way of example and not limitation, shows an exemplary diagram of a system 200 for automated natural language processing.

[0017] IMPLEMENTATION OF THE INVENTION

[0018] In a preferred embodiment of the present invention, there is provided a method executable by a processor of a computer device for automated processing of a natural language text, which method comprises at least the following steps: a step of identifying a natural language text comprising at least three segments; a step of identifying the segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step of marking up in the selected first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment at least one part to be parsed; a step of parsing the marked parts by means of semantic syntactic parsing;a step of extracting from the part of the selected first segment subjected to semantic-syntactic parsing at least the main entity of the first segment, and at least an associated entity associated with the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step of associating said statement with said main entity.

[0019] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-combined to obtain identifiable text in a natural language.

[0020] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-associated to obtain identifiable text in a natural language.

[0021] In a particular embodiment of the present invention, the said method is provided, characterized in that the part to be parsed marked in the selected first segment is the first sentence in a natural language.

[0022] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and extracted for each associated entity, at least a main entity of the associated entity, and at least a nested associated entity related to the main entity of the associated entity, wherein at least one of the nested associated entities is a nested associated leaf entity.

[0023] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested in nested associated entities, until a nested associated entity is retrieved in which no entity is nested.

[0024] In a particular embodiment of the present invention, the said method is provided, characterized in that the said part to be parsed, marked in the selected first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing.

[0025] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the selected first segment and all associated entities associated with it are extracted from the first part.

[0026] In a particular embodiment of the present invention, the said method is provided, characterized in that at least the main entity of the second part and at least an associated entity linked to the main entity of the second part are extracted from the second part.

[0027] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

[0028] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested in nested associated entities, until a nested associated entity is retrieved in which no entity is nested.

[0029] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the second part is associated with the main entity of the selected first segment.

[0030] In a particular embodiment of the present invention, the said method is provided, characterized in that for each extracted entity, lemmatization is at least partially performed.

[0031] In a particular embodiment of the present invention, the said method is provided, characterized in that the marking is carried out using a classification model and / or a clustering model.

[0032] In a particular embodiment of the present invention, the said method is provided, characterized in that the division of the part of the first segment to be analyzed is carried out using a classification model and / or a clustering model.

[0033] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computer device, is provided, in which, at least by means of said method of automated processing of text in natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the selected first segment, and a text corpus is formed from the obtained pairs of texts.

[0034] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computer device, is provided, in which, at least by means of said method of automated processing of text in a natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the second part of the part to be parsed marked up in the selected first segment, and a text corpus is formed from the obtained pairs of texts.

[0035] In another preferred embodiment of the present invention, a method for pre-training, or training, or additional training of a classification model, executable by a processor of a computer device, is provided, in which at least the classification model is trained to classify a text analysis as associated with any statement obtained using any mentioned method for forming a text corpus, when receiving at the input of the classification model a text analysis containing at least a main entity and at least an associated terminal entity associated with the main entity, wherein the training is carried out using a text corpus obtained by any mentioned method for forming a text corpus.

[0036] In another preferred embodiment of the present invention, a method for pre-training, or training, or additional training of a clustering model, executable by a processor of a computer device, is provided, in which at least the clustering model is trained to classify a text analysis as associated with any statement obtained using any mentioned method for forming a text corpus, when receiving at the input of the clustering model a text analysis containing at least a main entity and at least an associated leaf entity associated with the main entity, wherein the training is carried out using a text corpus obtained by any mentioned method for forming a text corpus.

[0037] In another preferred embodiment of the present invention, a method for classifying a text analysis, executable by a processor of a computer device, is provided, which at least provides at the input of said classification model a text analysis containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text analysis as being associated with any statement obtained using any said method for forming a text corpus.

[0038] In another preferred embodiment of the present invention, a method for classifying text parsing, executable by a processor of a computing device, is provided, which at least provides at the input of said clustering model a text parsing containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text parsing as being associated with any statement obtained using any said method for forming a text corpus.

[0039] In another preferred embodiment of the present invention, a method for generating a database, executable by a processor of a computer device, is provided, in which a plurality of text parses classified by any said classification method are identified, they are associated with one or more corresponding statements, and a database is generated containing at least a plurality of said text parses, each of which is associated with one or more said statements.

[0040] In another preferred embodiment of the present invention, a method for generating a computer device, executable by a processor, is provided. a database query, in which a database query is generated that contains at least one statement obtained using any of the aforementioned methods for generating a text corpus; wherein the aforementioned database contains at least a plurality of text analyses, each of which is associated with one or more of the aforementioned statements.

[0041] In another preferred embodiment of the present invention, a method for selecting text parses, executable by a processor of a computer device, is provided, in which at least a query is generated and sent to a database, containing at least some statement obtained using some method for forming a corpus of texts, and at least one text parse is obtained, associated with said statement; wherein said database contains at least a plurality of text parses, each of which is associated with one or more of said statements.

[0042] In another preferred embodiment of the present invention, there is provided a method executable by a processor of a computer device for generating a text entry, in which at least one statement obtained using any of the above is received and a syntactically and semantically correct sentence is generated, including the statement or a derivative thereof.

[0043] In another preferred embodiment of the present invention, a method for generating a database of text records, executable by a processor of a computer device, is provided, which includes at least obtaining a plurality of text records by means of said method for generating text records and recording the obtained text records in a database; the approval is obtained using any of said methods for generating a corpus of texts.

[0044] In another preferred embodiment of the present invention, a method for generating natural language text, executable by a processor of a computer device, is provided, which comprises at least receiving a text entry from a database formed by said method for forming a database including an assertion of text entries, after which generating natural language text, including at least said text entry and a text that is not said entry.

[0045] In another preferred embodiment of the present invention, there is provided a method for generating, executable by a processor of a computing device, a natural language text, in which at least a natural language text entry is identified, including at least a main entity and at least an associated end entity linked to it, and a statement linked to the main entity, after which a text entry including this statement is obtained from a database formed by said method of forming a database of text entries including the statement, and a natural language text is generated, including at least the identified text entry and the text entry obtained from said database

[0046] In another preferred embodiment of the present invention, there is provided a computer device for automated processing of natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of automated processing of natural language text.

[0047] In another preferred embodiment of the present invention, a computer device for generating a text corpus is provided, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any of the said methods for generating the text corpus.

[0048] In another preferred embodiment of the present invention, a computer device is provided for pre-training, or training, or re-training a classification model, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of pre-training, or training, or re-training a classification model.

[0049] In another preferred embodiment of the present invention, a computer device is provided for pre-training, or training, or re-training a clustering model, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of pre-training, or training, or re-training a clustering model.

[0050] In another preferred embodiment of the present invention, there is provided a computer device for classifying text parsing, comprising at least: a processor; a memory containing program code that when executed by the processor, causes the processor to perform the actions of any of the mentioned text parsing classification methods.

[0051] In another preferred embodiment of the present invention, a computing device for generating a database is provided, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any said method of generating a database.

[0052] In another preferred embodiment of the present invention, a computing device is provided for generating a query to a database, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a query to a database.

[0053] In another preferred embodiment of the present invention, there is provided a computer device for selecting text parses, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for selecting text parses.

[0054] In another preferred embodiment of the present invention, there is provided a computer device for generating a text entry, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a text entry.

[0055] In another preferred embodiment of the present invention, a computer device is provided for generating a database of text entries including an approval, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a database of text entries including an approval.

[0056] In another preferred embodiment of the present invention, there is provided a computer device for generating natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating natural language text.

[0057] In another preferred embodiment of the present invention, a machine-readable storage medium is provided, including a non-transitive a machine-readable storage medium containing program code that, when executed by a processor, causes the processor to perform actions in some specified manner.

[0058] The following are embodiments of the present invention, revealing examples of its implementation in specific embodiments. However, the description itself is not intended to limit the scope of rights granted by this patent. Rather, it should be understood that the claimed invention may also be implemented in other ways that incorporate different elements and conditions, or combinations of elements and conditions similar to those described herein, in combination with other existing and future technologies.

[0059] In the present description, a patent document shall mean, without limitation, a patent or a patent application, that is, as a rule, a specific text consisting of at least three segments: the first segment, which is the patent claims (patent claims), the second segment, which is the description (description, specification), and the third segment, which is the abstract (abstract), wherein, without limitation, the numbering of the segments above is not given in the order of their sequence in the text in natural language, but only for the simplicity of presentation and association in the present document, as will be shown below. Preferably, without limitation, the formulas of patent documents do not relate to new chemical compounds and other similar objects created for the first time, since with respect to such objects the statement of the technical result does not require disclosure.

[0060] In Fig. 1, by way of example and not limitation, an exemplary flow chart of a method 100 for automated processing of natural language text is shown. The method 100 comprises at least the following steps: a step 101 of identifying a natural language text that includes at least three segments; a step 102 of identifying the segments; a step 103 of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step 104 of marking up only one part to be parsed in the selected first segment and a step of marking up at least one part to be parsed in the selected second segment and / or in the selected third segment; a step 105 of parsing the marked parts by means of semantic-syntactic parsing;step 106 of extracting from the part of the selected first segment subjected to semantic-syntactic parsing at least the main entity of the first segment and at least an associated entity associated with the main entity of the first segment, wherein at least one of the associated entities is an associated end; entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; step 107 of associating said statement with said main entity. For example, without limitation, within the framework of step 101 of identifying text in a natural language, reliable identification of text in a natural language is ensured, for example, without limitation, a patent document. For example, without limitation, segments of the text of a patent document can be represented by separate texts in a natural language, for which, if necessary, they are first associated, for example, by assigning them some features of relatedness, for example, unique identifiers, the decoding of which makes it possible to determine the relatedness of the segments among themselves.Such methods and techniques for associating entities in data storage systems are widely known in the art and, accordingly, are not described in detail below. For example, but not limited to, a patent document may be a single text, typically with continuous page numbers; in such a case, prior association of segments is not required.

[0061] Preferably, but not limited to, step 102 involves identifying segments of the natural language text; for example, but not limited to, in the case where the natural language text is a patent document, at least the patent claims, which is the first segment in the context of the present invention, the description, which is the second segment in the context of the present invention, and the abstract, which is the third segment in the context of the present invention, are identified. Without limitation, in the case where the natural language text is a non-patent document, the segments are identified in a similar manner. In this case, without limitation, it should be primarily assumed that the first segment includes a set of entities, while the second and third segments include at least one or more statements characteristic of said set of entities.An example of a suitable non-patent document may be, for example, but not limited to, a scientific article in which the role of the first segment may be assigned to the abstract, and the roles of the second and third segments, respectively, to the main text and conclusion.

[0062] Preferably, but not limited to, step 103 comprises selecting the identified segments, including, but not limited to, selecting the first segment as comprising the set of entities and one of the second segment or the third segment. In this case, but not limited to, the choice is made as to which Whether a statement should be selected from the second or third segment depends on the likelihood of its presence in a given segment. More specifically, but not limited to, the presence of a statement in a given segment will generally depend on the approach to the requirements imposed on natural language text. For example, if the natural language text is a patent document, then, depending on the patent system in which it is formed, the third segment (abstract) will contain at least a statement of the technical result. For example, but not limited to, Russian patent documents are typically published with an abstract formulating the technical result. Moreover, as of the filing date of this application, the preparation of the abstract for the purpose of deciding whether to grant a patent rests with an examiner who is competent in matters of formulating the technical result, thus eliminating inaccuracies in its presentation.Meanwhile, in US patent documents, the applicant typically prepares the abstract themselves, resulting in abstracts in US patent documents lacking a clear structure, and the presence of a technical result claim is not guaranteed. Thus, for example, without limitation, when choosing between the second and third segments in the US, it would be more effective to select the second segment, since at least the Background of the Invention or Brief Summary of Invention sections are likely to contain a technical result claim, as required by MPEP 608.01(c) and 608.01(d). At the same time, for example, without limitation, for each natural language text, a particular segment can be selected based on the presence of the claim itself.For this purpose, for example, but not limited to, a semantic-syntactic analysis can be performed for each natural language text, for example, using a semantic parser, the technology of which is widely represented in the prior art and, accordingly, is not described below. Based on the results of the semantic-syntactic analysis, it becomes possible to identify the specific portion of the natural language text containing the desired assertion. Preferably, but not limited to, the first segment of the natural language text is preliminarily excluded, as it is known to contain no assertions, and whose presence during the semantic-syntactic analysis could distort the results.

[0063] Preferably, but not limited to, step 104 involves marking up the parts in the selected segments. Preferably, but not limited to, only one part to be parsed is marked up in the selected first segment, namely, the part that contains the set of entities. For example, but not limited to, in the case where the natural language text is a patent document, the marked up part is The first sentence, i.e., the first claim, is selected as the one most likely to possess the required set of entities, i.e., the set of essential features. However, without limitation, there are situations where the set of entities (features) in the first claim, in addition to the essential features, also includes nonessential features. In this case, the selected first segment may be subject to additional analysis after finding the technical result statement in the second or third segment and verifying which features do not affect the achievement of the technical result. Based on the results of this additional analysis, the marked portion is cleared of unnecessary entities.In this case, preferably, but not limited to, the marked and optionally purified part of the selected first segment to be parsed may be subjected to analysis in order to determine the presence of a generic part (the first part) and a distinctive part (the second part); in this case, for example, but not limited to, it may be sufficient to determine a linking word separating the generic part from the distinctive part. At the same time, not every patent system requires that an independent claim be composed with the use of a distinctive part, which results in the generic part not being clearly distinguished and not cut off by a linking word; in such a case, for example, but not limited to, a preliminary semantic-syntactic analysis of the entire first segment may be carried out in order to identify at least the main entity of the first segment and the main entities of the entities associated with the main entity of the first segment, which will be discussed in detail below.Typically, but not limited to, when the natural language text is a patent document, the primary entity will be defined as a generic concept, and the associated entities will be individual features, such as, but not limited to, method steps or product parts. The resulting set of primary entities and associated entities can be subjected to a familiarity analysis, which can determine whether the entire set is generic or only a subset of the features constitutes a generic part. For example, but not limited to, the familiarity analysis can be performed automatically, based on a pre-prepared text corpus, for each of which a first segment has been identified and a semantic-syntactic analysis has been performed on the part to be analyzed.Moreover, without limitation, separation can be performed either manually or automatically, for example, using a pre-trained classification model and / or clustering model. This can be done, without limitation, in the selected second segment or in the selected third segment. The tagging is performed in such a way as to obtain the desired statement with the highest probability. Typically, the statement of the desired statement is preceded by a certain text structure, or the statement of the desired statement includes certain keywords. Without limitation, several parts of the selected second segment or the selected third segment to be parsed may be tagged in this manner. Moreover, without limitation, the tagging may be performed both manually and automatically, for example, using a pre-trained classification model and / or a clustering model. Preferably, but not limited to, step 105 involves parsing the tagged parts through semantic-syntactic analysis using, for example, but not limited to, the aforementioned syntactic parser.

[0064] Preferably, but not limited to, step 106 comprises extracting from the parsed portions of the first segment at least the main entity of the first segment and at least an associated entity linked to the main entity of the first segment, wherein at least one of the extracted associated entities is an associated terminal entity. For example, without limitation, in the case where the first segment is a patent claim, the first independent claim will be subjected to semantic-syntactic parsing; in such a case, the main entity will be a generic concept that is the root of the parsing tree, the associated entities will be the features of the technical solution that are the nodes of the parsing tree, and at least one associated entity will be an associated terminal entity, that is, it will be a feature that has only one edge, that is, it will be a terminal node (leaf) of the parsing tree.In this case, without limitation, the parse tree can be constructed based on the principle of an abstract syntax tree (AST), meaning it can be cleared of all irrelevant features, as well as repetitions, anaphora, and other elements that do not affect the claim, i.e., in the case of a patent claim, which cannot affect the technical result. An example of redundant elements obtained after parsing is, for example, but not limited to, a comma or other separator, which are useful directly during parsing, for example, as links or boundaries, but are redundant for the text corpus when it is used, for example, to train a classification model and / or a clustering model. Furthermore, for example, but not limited to, when the first segment was divided into generic and distinctive parts, the main essence of the first segment after said parsing will be found in the first part, i.e., in the generic part, along with...associated with it entities from the generic part; in this case, without limitation, the second part, which is the distinctive part, will in this case have the main entity of the second part and entities associated with it, which, nevertheless, are also all associated with the main entity of the first segment, since all individual features revealed through analysis directly or indirectly belong to only one generic concept.In this case, for example, but not limited to, each associated entity extracted in this manner may be iteratively subjected to its own semantic-syntactic analysis, during which the corresponding principal entities of the associated entities and the associated associated entities of the associated entities will be discovered. That is, each associated entity may have multiple nested associated entities, up to one or more associated terminal entities that can no longer be subjected to semantic-syntactic analysis because they are too simple features. In this case, for example, but not limited to, lemmatization may be performed to reduce words to their original forms.Furthermore, without limitation, for each entity, including the main entity, as well as any associated entities, a generalization operation may be performed, in which the entity may be elevated to the level of a method or product, which corresponds to a certain function determined by the features of the entity. In this case, preferably, without limitation, at least one statement is extracted from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing. For example, without limitation, in the case where the natural language text is a patent document, upon semantic-syntactic parsing of paragraph 0010 of the present description in its original version, two statements will be obtained: "accurate generation of a text corpus" and "automated generation of a text corpus."For example, but not limited to, in the same case, three assertions may be obtained during the semantic-syntactic analysis of paragraph 0011: "implementation of the purpose - generation of a text corpus," "precise generation of a text corpus," and "automated generation of a text corpus." Preferably, but not limited to, at step 107, at least one extracted assertion is associated with the primary entity of the first segment, and thus the assertion is associated with any associated entities related to the primary entity of the first segment. For example, but not limited to, some of the resulting assertions, such as "implementation of the purpose," are of no practical use, as they cannot be subsequently used as part of the text corpus for model training. Classification and / or clustering models, for example, to predict a statement corresponding to a new set of entities, as this would result in any classifier based on the classification and / or clustering model assigning any set of entities to such a statement, since such a statement is inherent in all resulting text parses. For this reason, it is preferable, but not limited to, to clean up the resulting "statement-main entity" pairs, removing redundant statements.

[0065] In this manner, preferably, without limitation, it becomes possible to obtain a corpus of texts, for which purpose, by means of said method 100, in step 107, by associating said statement with said main entity of the first segment (and, as a consequence, with all associated entities related to the main entity), a plurality of pairs of texts are obtained, each of which includes at least one statement related to the main entity of the first segment, and a corpus of texts is formed from the obtained pairs of texts.Alternatively, without limitation, it becomes possible to obtain a corpus of texts, which by means of said method 100 in step 107 by associating said statement with said main entity of the second part of the part to be parsed marked in the first segment (and, as a consequence, with all associated entities related to the main entity of the second part of the part to be parsed marked in the first segment) a plurality of pairs of texts are obtained, each of which includes at least one statement related to the main entity of the second part of the part to be parsed marked in the first segment, and a corpus of texts is formed from the obtained pairs of texts.Such resulting text corpora, for example, but not limited to, may be used, either jointly or separately, for pre-training, or training, or additional training of a classification model and / or a clustering model, in which, at least, the classification model and / or the clustering model are trained to classify a text parse as related to any said statement, upon receipt at the input of the classification model and / or the clustering model of a text parse containing, at least, a main entity and an associated terminal entity associated with the main entity, wherein the training is carried out on one of the said text corpora, jointly or separately. In this case, without limitation, the main entity in this way is understood to be, for example, the main entity of the first segment or the said main entity of the second part of the part of the first segment to be parsed.To train a classification model and / or a clustering model, for example, but not limited to, a sample is provided from a text corpus, after which the classification model is trained. and / or clustering models on such a sample. In this case, for example, but not limited to, the classification model is based on one of the following: logistic regression, support vector machine, decision tree method, random forest method, naive Bayes classifier, k-nearest neighbors method, neural network, boosting, gradient boosting, bagging, group argument accounting method; in this case, without limitation, the clustering model is based on one of the following: k-means method, density-based spatial clustering of applications with noise (DBSCAN), hierarchical clustering method, spectral clustering method. In this case, without limitation, any suitable known or future known classification model and / or clustering model may be used, when choosing which it is necessary to proceed from the fact that such a model provides the possibility of obtaining equal weights for different statements.This is primarily due to the fact that the same text parsing may, in fact, correspond to several different claims. For example, without limitation, the same invention, characterized by the same set of essential features, may achieve different technical results, each of which may be suitable and lead to a solution to a specific technical problem. This is due to the fact that the formulation of a technical problem is primarily related to the prior art, or even the prototype, in which the technical problem is formulated. Therefore, it is preferable that the classification model and clustering model employed ensure the fundamental possibility of classifying a text parsing as corresponding to several claims simultaneously.Moreover, for example, without limitation, the weights of the statements do not necessarily have to be the same for a decision to assign a text analysis to these statements; in fact, without limitation, it is possible to ensure the assignment of a text analysis to several statements at once by defining an acceptable proximity of the weights of the statements.

[0066] Thus, without limitation, when the corresponding classification model or clustering model is obtained, it becomes possible to provide a method for classifying text parsing, which at least feeds to the input of the obtained classification model and / or clustering model a text parsing containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text parsing as associated with any said statement obtained using any said method of forming a text corpus. In this case, without limitation, under the main essence is thus understood to be, for example, the main essence of the first segment or the said main essence of the second part of the part of the first segment to be analyzed.

[0067] The classified text analyses can subsequently be stored in the database of the automated natural language processing system 200. For this purpose, without limitation, a method for creating a database is preferably provided, which involves identifying a plurality of said classified text analyses, associating them with one or more corresponding statements, and creating a database containing at least a plurality of said text analyses, each of which is associated with one or more said statements. From the database thus created, it becomes possible to obtain text analyses corresponding to any statement selected by the user.For this purpose, preferably, but without limitation, a method for generating a query to a database is provided, in which a query to the database is generated, containing at least one statement obtained using any of the aforementioned methods for generating a text corpus; wherein the aforementioned database contains at least a plurality of text parses, each of which is associated with one or more of the aforementioned statements. For example, without limitation, a query may be provided containing the statement "accurate generation of a text corpus," in response to which at least one corresponding text parse will be selected. At the same time, depending on the statement, too many text parses may be selected, which may render the query irrelevant. In order to increase the relevance of the query, for example, without limitation, the query may be supplemented with an additional statement, which will inevitably reduce the resulting selection of text parses.Thus, preferably, without being limited, a method for selecting text analyses is provided, in which a query is generated and sent to a database, containing at least some statement obtained using any mentioned method for forming a text corpus, and at least one text analysis associated with said statement is obtained; wherein said database contains at least a plurality of text analyses, each of which is associated with one or more mentioned statements.

[0068] Moreover, without being limited to, the resulting preferably provides new possibilities for generating texts in natural language, in particular, but not limited to, such texts as descriptions of patent documents. Preferably, without being limited to, a method for generating a text entry is provided, in which, at least, At least one assertion obtained using any of the aforementioned methods for generating a text corpus is received, and a syntactically and semantically correct sentence is generated that includes the assertion or a derivative thereof. In this case, without limitation, a text entry is understood to mean a portion of the text, but not the entire text, and a text, accordingly, is understood to mean a collection of, including heterogeneous, text entries. In this case, without limitation, the aforementioned assertion may be present in the generated text entry both unchanged and in derivative form, i.e., when the assertion is subject to modifications and changes in order to ensure, at least, the consistency of the generated sentence. In this case, preferably, without limitation, the obtained text entries can be placed in a database of text entries including the assertion.For this purpose, preferably, but without limitation, a method for generating a database of text records is provided, which at least involves obtaining a plurality of text records by means of said text record generation method and recording the obtained text records in the database; wherein the approval is obtained using any of said text corpus generation methods. For example, without limitation, such a database can subsequently be used as a source of standardized records when generating natural language text.For this purpose, for example, but not limited to, a method for generating natural language text is provided, which involves at least obtaining a text entry from a database of text entries formed by said generation method, including an assertion, and then generating natural language text, which includes at least said text entry and text that is not said entry. More specifically, without limitation, the text that is not said entry may be, for example, but not limited to, text entered by the user.That is, without limitation, a method for generating natural language text is provided, which involves at least identifying a natural language text entry comprising at least a main entity and at least an associated leaf entity linked thereto, and a statement linked to the main entity, after which a text entry comprising this statement is obtained from said database of text entries comprising the statement, and generating natural language text comprising at least the identified text entry and the text entry obtained from said database. In this case, without limitation, the identified entry is a user-entered text entry. More specifically, not. Without limitation, the user-entered text entry represents a syntactically and semantically correct sentence. More specifically, without limitation, the user-entered entry represents an independent patent claim. Thus, for example, without limitation, the generation of natural language text is achieved, based on the user-entered text and including a portion of the text not actually entered by the user, i.e., generated automatically. Moreover, without limitation, the text generation itself is accomplished using methods and means known in the art, such as, for example, NLP processors, which, accordingly, are not described in detail below.

[0069] Thus, preferably, without being limited, as shown in Fig. 2, a computer device 201 can be provided, in which, in the context of the present invention, at least one of or any combination of: a computer device for automated processing of text in natural language, a computer device for forming a text corpus, a computer device for pre-training, or training, or further training a classification model, a computer device for pre-training, or training, or further training a clustering model, a computer device for classifying text parsing, a computer device for forming a database, a computer device for forming a query in a database, a computer device for selecting text parsings can be embodied.Such a computing device 201 most typically comprises at least: one or more processors 2011; a memory 2012 containing program code that, when executed by the processor 2011, causes the processor 2011 to perform actions of any of the aforementioned methods described with reference to Fig. 1 for automated processing of text in a natural language, and / or a method for forming a text corpus, and / or a method for pre-training, or training, or further training of a classification model, and / or a method for pre-training, or training, or further training of a clustering model, and / or a method for classifying text parsing, and / or a method for forming a database, and / or a method for forming a query to a database, and / or a method for selecting text parsing.At the same time, such a computing device can be implemented as a thin client, which will mean that all, or at least most, computing operations are performed on the server 202 of the system, which thus also contains at least a processor 2021 and a memory 2022, which are thus essentially similar, respectively, to the processor 2011 and the memory 2012. By way of example, but not limitation, the memory 2012, 2022 (the machine-readable storage medium 2012, 2022) can include non-volatile memory (NVRAM); random access memory (RAM);. Read-only memory (ROM); Electrically erasable programmable read-only memory (EEPROM); Flash memory or other memory technologies; CDROM, Digital Versatile Disc (DVD) or other optical or holographic storage media; Magnetic cassettes, magnetic tape, magnetic disk storage device or other magnetic storage devices; as well as any other storage medium that can be used to store and encode the required information. In this case, without limitation, memory 2012, 2022 includes a storage medium based on a computer storage device in the form of volatile or non-volatile memory, or a combination thereof. In this case, without limitation, exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, and so on.In this case, without being limited to, the machine-readable data carrier 2012, 2022 (memory 2012, 2022) is not temporary (permanent, non-transitive), so that it does not include a temporary (transitive) propagating signal. In this case, without being limited to, an exemplary environment can be stored in the memory 2012, 2022, in which, using computer commands or codes, including those stored in the memory 2022 of the server 202, a procedure for automated processing of text in a natural language can be carried out, and / or a procedure for forming a text corpus, and / or a procedure for pre-training, or training, or additional training of a classification model, and / or a procedure for pre-training, or training, or additional training of a clustering model, and / or a procedure for classifying text parsing, and / or a procedure for forming a database, and / or a procedure for forming a query to a database, and / or a procedure for selecting text parsing.In this case, without limitation, the computer device 201, when not a thin client, comprises one or more processors 2011, which are designed to execute computer instructions or codes stored in the memory 2012 of the device 201 in order to ensure the execution of the said procedures. In this case, without limitation, the server 202 can be essentially similar to the computer device 201, when not a thin client, and, accordingly, comprise one or more processors 2021, which are designed to execute computer instructions or codes stored in the memory 2022 of the server 202 in order to ensure the execution of the said procedures. In this case, without being limited, the system 200 may also include a database (DB) 203. The DB 203 may be, but is not limited to: a hierarchical DB, a network DB, a relational DB, an object DB, an object-oriented DB, an object-relational DB, a spatial DB, a combination of two or more of the above DBs, and the like.At the same time, without limitation, DB 203 at least stores classified text analyses associated with. with the relevant statements, and can also store data for analysis, classification models, clustering models and other information in memory 2021, 2022 or in a suitable memory of another computing device associated with computing device 201 and / or server 202, which may be, but is not limited to, a memory similar to any memory 2021, 2022, as shown earlier, and which can be accessed via server 202. In addition, without limitation, server 202 is provided, which, in addition to the previously described functions, stores and facilitates the manipulation of computer commands or codes previously described in this document, which, accordingly, are not further described. Moreover, without limitation, server 202, in addition to the previously described functions, can ensure the regulation of data exchange in system 200.In this case, without limitation, the exchange of data within the system 200 is carried out thanks to one or more data networks 204. In this case, without limitation, the data networks 204 may include, but are not limited to, one or more local area networks (LAN) and / or wide area networks (WAN), or may represent an information telecommunications network, the Internet, or an Intranet, or a virtual private network (VPN), or a combination thereof, and the like. In this case, without limitation, the server 202 also has the ability to provide a virtual computing environment for ensuring interaction between the components of the system. In this case, without limitation, the network 204 serves to ensure interaction between the computer device 201, the server 402 and, optionally, the database 203.In this case, without limitation, the non-thin client computing device 201 and / or the server 202 may be directly connected to the database 203 using wired and wireless communication methods and techniques known in the art, which, accordingly, are not described in detail below, or, without limitation, the database 203 may be implemented in the memory 2012, 2022. In this case, without limitation, a suitable non-thin client computing device 201 may perform the role of the server 202 of the system 200 for other computing devices 201 that are thin clients. In this case, most typically, without limitation, the components of the computing device 201 and the components of the server 202 are interconnected, including by means of some data bus.

[0070] The present description of the implementation of the claimed invention demonstrates only particular embodiments and does not limit other embodiments of the claimed invention, since possible other alternative embodiments of the claimed invention, not going beyond the scope of the information set out in this application, should be obvious to a person skilled in the art. field of technology, having ordinary qualifications, for whom the claimed invention is intended.

Claims

CLAUSES OF THE INVENTION 1. A computer device for classifying a text parse, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of a method for classifying a text parse, which at least includes feeding to the input of a classification model a text parse containing at least a main entity and an associated leaf entity associated with the main entity, and classifying said text parse as being associated with any statement obtained using the method for generating a corpus of texts;wherein the method for forming a text corpus is a method executable by a processor of a computer device in which, at least by means of a method for automated processing of text in a natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the first segment, and a corpus of texts is formed from the obtained pairs of texts; wherein the method for automated processing of text in a natural language is a method executable by a processor of a computer and consists of at least performing the steps of: a step of identifying text in a natural language, including at least three segments; a step of identifying segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of text in a natural language;a step of marking in the first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment at least one part to be parsed; a step of parsing the marked parts by means of semantic-syntactic parsing; a step of extracting from the part of the first segment subjected to semantic-syntactic parsing at least the main entity of the first segment, and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated end entity,; and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step of associating said statement with said main entity of the first segment; wherein the classification model is obtained by means of a method of pre-training, or training, or additional training of a classification model, executed by a processor of a computer device, in which at least the classification model is trained to classify the parsing of a text as associated with any statement obtained using said method of formation, when a parsing of a text containing at least the main entity and at least an associated terminal entity associated with the main entity is received at the input of the classification model, wherein the training is carried out using a corpus of texts obtained by means of said method of forming a corpus of texts.

2. The method according to claim 1, characterized in that the first segment, the second segment, and the third segment are pre-combined to obtain an identifiable text in a natural language.

3. The method according to claim 1, characterized in that the first segment, the second segment, and the third segment are pre-associated to obtain identifiable text in natural language.

4. Method according to item 1, characterized by the fact that the part to be parsed, marked in the first segment, is the first sentence in a natural language.

5. The method according to claim 4, characterized in that each extracted associated entity is subjected to semantic-syntactic analysis and for each associated entity at least the main entity of the associated entity and at least a nested associated entity linked to the main entity of the associated entity are extracted, wherein at least one of the nested associated entities is a nested associated end entity.

6. The method according to claim 1, characterized in that the said part to be parsed, marked in the first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing; wherein the main entity of the first segment and all associated entities connected with it are extracted from the first part; wherein the main entity of the second part, and at least the associated entity linked to the main entity of the second part is extracted from the second part.

7. The method according to claim 6, characterized in that each extracted associated entity is subjected to semantic-syntactic analysis and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

8. The method according to paragraph 7, characterized in that the main essence of the second part is associated with the main essence of the first segment.

9. The method according to any of paragraphs 1-8, characterized in that before forming the text corpus, at least one extracted statement is removed.

Citation Information

Patent Citations

  • Text segmentation

    RU2666277C1

  • Systems and methods for providing adaptive surface texture in auto-drafted patent documents

    US11023662B2

  • Patent application preparation system and template creator

    US11341323B1

  • Systems and methods for extracting patent document templates from a patent corpus

    US11593564B2