Computing device for automated processing of natural language text

The computer device and method address the challenge of incomplete automated text processing in patent documents by segmenting, parsing, and extracting entities, creating a text corpus for model training and statement extraction.

WO2026063823A1PCT designated stage Publication Date: 2026-03-26KRAVCHENKO ARTEM ALEKSANDROVICH
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing methods for automated natural language processing, particularly in the context of patent documents, fail to accurately extract and form texts with specific completeness, such as patent documents, lacking techniques for extracting statements concerning the technical result.

Method used

A computer device and method for automated text processing that identifies and parses natural language texts into segments, performs semantic-syntactic parsing, and extracts main entities and associated entities, forming a text corpus for pre-training or training classification and clustering models.

Benefits of technology

Ensures accurate automated generation of a text corpus for training models, expanding the arsenal of methods for natural language processing and enabling the extraction of relevant statements from patent documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000026_0000
    Figure 00000026_0000
  • Figure 00000027_0000
    Figure 00000027_0000
Patent Text Reader

Abstract

The proposed technical solution relates to methods for automated text processing and can be used in the generation of text corpora. What is proposed is a computing device for the automated processing of natural language text. The claimed invention solves the technical problem of creating a method and / or a computing device and / or a system and / or a machine-readable data carrier which overcome the disadvantages of the prior art and, in so doing, provide for the precise automated generation of a text corpus which can subsequently be used for pre-training, or training, or post-training classification models and / or clustering models.
Need to check novelty before this filing date? Find Prior Art

Description

COMPUTER DEVICE FOR AUTOMATED PROCESSING OF NATURAL LANGUAGE TEXT

[0001] AREA OF TECHNOLOGY

[0002] The proposed technical solution relates to methods of automated text processing and can be used in the formation of text corpora.

[0003] LEVEL OF TECHNOLOGY

[0004] Various methods for automated generation of texts in natural language are known, for example, those disclosed in the documents: US20070136321A1 dated 14.06.2007, US20100205125A1 dated 12.08.20120, US20190377780A1 dated 12.12.2019, US20130080883A1 dated 28.03.2013, US10713443B1 dated 14.07.2020, US10747953B1 dated 18.08.2020, US11023662B2 dated 01.06.2021, US11341323B1 dated 24.05.2022. However, the known solutions do not allow the formation of texts in natural language with sufficient completeness, specific to a certain area, such as, for example, patent documents, in which, in addition to the actual formation of disclosures, it is also necessary to ensure the formation of special parts of the patent document, such as, for example, the level of technology (Background of the Invention), for which in most cases it is necessary to have a statement in the form of a disclosure of the technical problem solved by the invention or the technical result achieved by using the invention (technical effect (9.2.8 Case Law of the Boards of Appeal), (technical) result (MPEP 716.02(a))).

[0005] Patent No. US11593564B2 of February 28, 2023 (D1) discloses systems, methods, and storage media for extracting patent document templates from a patent corpus. Examples of implementation include: obtaining a patent corpus; obtaining one or more parameters; determining one or more subsets of the patent corpus by filtering the patent corpus based on one or more parameters; identifying one or more clusters of documents within individual ones of one or more subsets of the patent corpus; obtaining a patent document template corresponding to the first cluster of documents; and / or performing other operations. However, the solution known from D1 does not disclose any methods or techniques for extracting, in particular, statements concerning the technical result.

[0006] From the patent US5774833A of 30.06.1998 (D2) a method for processing patent text on a computer is known, which includes determining the boundaries of parts of the patent text, loading at least one of the parts of the patent text into the working Memory, analysis of at least one portion of the patent text, and reporting of the results to the user. The alphanumeric data of the drawing can also be compared with the patent text. Method D2 can be combined with a word processor. Method D2 makes it possible to recognize and report claim dependencies, specific characteristics of the patent text, and patent errors based on legal standards, standards of practice, USPTO standards, or even user preferences. However, the prior art solution D2 does not disclose any methods or techniques for extracting, in particular, claims concerning the technical result.

[0007] Thus, there is a problem of automated extraction of claims from specific natural language texts, such as, but not limited to, patent documents.

[0008] The solution known from D2 can be chosen as the closest analogue.

[0009] DISCLOSURE OF THE INVENTION

[0010] The technical problem solved by the claimed invention is the creation of a method and / or a computer device and / or a system and / or a machine-readable data carrier that does not have the disadvantages of analogs and thus ensures the accurate automated formation of a text corpus, which can subsequently be used for pre-training, or training, or additional training of classification models and / or clustering models.

[0011] The technical result achieved by implementing the claimed invention, in addition to fulfilling its intended purpose, is to eliminate the shortcomings of similar technologies, thereby ensuring the accurate automated generation of a text corpus, which can subsequently be used for pre-training, training, or further training of classification and / or clustering models. Another technical result is the expansion of the arsenal of technical means—methods for automated processing of natural language text.

[0012] The technical result is achieved due to the fact that a computer device is provided for automated processing of text in a natural language, containing at least: a processor; memory containing program code, which, when executed by the processor, causes the processor to perform actions of a method for automated processing of text in a natural language, consisting of at least performing the steps of: a step of identifying text in a natural language, including at least three segments; a step of identifying segments; a step of selecting, at least a first segment and at least a second segment and / or at least a third segment of a natural language text; a step of marking up in the first segment only one part to be parsed, and a step of marking up in a selected second segment and / or in a selected third segment at least one part to be parsed; a step of parsing the marked parts by means of semantic-syntactic parsing; a step of extracting from the part of the first segment subjected to semantic-syntactic parsing at least a main entity of the first segment and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated leaf entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement;the stage of associating said statement with said main entity of the first segment.;

[0013] BRIEF DESCRIPTION OF DRAWINGS

[0014] Illustrative embodiments of the present invention are described below in detail with reference to the accompanying drawings, which are incorporated herein by reference, and in which:

[0015] Fig. 1, by way of example and not limitation, shows an exemplary flow chart of a method 100 for automated natural language text processing.

[0016] Fig. 2, by way of example and not limitation, shows an exemplary diagram of a system 200 for automated natural language processing.

[0017] IMPLEMENTATION OF THE INVENTION

[0018] In a preferred embodiment of the present invention, there is provided a method executable by a processor of a computer device for automated processing of a natural language text, which method comprises at least the following steps: a step of identifying a natural language text comprising at least three segments; a step of identifying the segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step of marking up in the selected first segment only one part to be parsed, and a step of marking in the selected second segment and / or in the selected third segment at least one part to be parsed; a step of parsing the marked parts by means of semantic syntactic parsing;a stage of extracting from the part of the selected first segment subjected to semantic-syntactic parsing at least the main entity of the first segment and at least an associated entity linked to the main entity; a first segment, wherein at least one of the associated entities is an associated end entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic syntactic parsing at least one statement; a step of associating said statement with said main entity.

[0019] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-combined to obtain identifiable text in a natural language.

[0020] In a particular embodiment of the present invention, the said method is provided, characterized in that the first segment, the second segment and the third segment are pre-associated to obtain identifiable text in a natural language.

[0021] In a particular embodiment of the present invention, the said method is provided, characterized in that the part to be parsed marked up in the selected first segment is the first sentence in a natural language.

[0022] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

[0023] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested in nested associated entities, until a nested associated entity is retrieved in which no entity is nested.

[0024] In a particular embodiment of the present invention, the said method is provided, characterized in that the said part to be parsed, marked in the selected first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing.

[0025] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the selected first segment and all associated entities associated with it are extracted from the first part.

[0026] In a particular embodiment of the present invention, the said method is provided, characterized in that at least the main entity of the second part and at least an associated entity linked to the main entity of the second part are extracted from the second part.

[0027] In a particular embodiment of the present invention, the said method is provided, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

[0028] In a particular embodiment of the present invention, said method is provided, characterized in that the actions of said method are performed iteratively for all nested associated entities, including all associated entities nested in nested associated entities, until a nested associated entity is retrieved in which no entity is nested.

[0029] In a particular embodiment of the present invention, the said method is provided, characterized in that the main entity of the second part is associated with the main entity of the selected first segment.

[0030] In a particular embodiment of the present invention, the said method is provided, characterized in that for each extracted entity, lemmatization is at least partially performed.

[0031] In a particular embodiment of the present invention, the said method is provided, characterized in that the marking is carried out using a classification model and / or a clustering model.

[0032] In a particular embodiment of the present invention, the said method is provided, characterized in that the division of the part of the first segment to be analyzed is carried out using a classification model and / or a clustering model.

[0033] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computer device, is provided, in which, at least by means of said method of automated processing of text in natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the selected first segment, and form a corpus of texts from the resulting pairs of texts.

[0034] In another preferred embodiment of the present invention, a method for generating a text corpus, executable by a processor of a computer device, is provided, in which, at least by means of said method of automated processing of text in a natural language, a plurality of pairs of texts are obtained, each of which includes at least one statement associated with the main essence of the second part of the part to be parsed marked up in the selected first segment, and a text corpus is formed from the obtained pairs of texts.

[0035] In another preferred embodiment of the present invention, a method for pre-training, or training, or further training a classification model, executable by a processor of a computer device, is provided, in which at least the classification model is trained to classify a text analysis as associated with any statement obtained using any mentioned method for forming a text corpus, when receiving at the input of the classification model a text analysis containing at least a main entity and at least an associated terminal entity associated with the main entity, wherein the training is carried out using a text corpus obtained by any mentioned method for forming a text corpus.

[0036] In another preferred embodiment of the present invention, a method for pre-training, or training, or additional training of a clustering model, executable by a processor of a computer device, is provided, in which at least the clustering model is trained to classify a text analysis as associated with any statement obtained using any mentioned method for forming a text corpus, when receiving at the input of the clustering model a text analysis containing at least a main entity and at least an associated leaf entity associated with the main entity, wherein the training is carried out using a text corpus obtained by any mentioned method for forming a text corpus.

[0037] In another preferred embodiment of the present invention, a method for classifying a text parse, executable by a processor of a computing device, is provided, which at least provides to the input of said classification model a text parse, containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text parse as being associated with any statement, obtained using any of the mentioned methods of forming a corpus of texts.

[0038] In another preferred embodiment of the present invention, a method for classifying text parsing, executable by a processor of a computing device, is provided, which at least provides to the input of said clustering model a text parsing containing at least a main entity and an associated leaf entity associated with the main entity, and classifies said text parsing as being associated with any statement obtained using any said method for forming a text corpus.

[0039] In another preferred embodiment of the present invention, a method for generating a database, executable by a processor of a computer device, is provided, in which a plurality of text parses classified by any said classification method are identified, they are associated with one or more corresponding statements, and a database is generated containing at least a plurality of said text parses, each of which is associated with one or more said statements.

[0040] In another preferred embodiment of the present invention, a method for generating a query to a database, executable by a processor of a computer device, is provided, in which a query to a database is generated, containing at least one statement obtained using any of the mentioned methods for generating a corpus of texts; wherein the mentioned database contains at least a plurality of text parses, each of which is associated with one or more of the mentioned statements.

[0041] In another preferred embodiment of the present invention, a method for selecting text parses, executable by a processor of a computer device, is provided, in which at least a query is generated and sent to a database, containing at least some statement obtained using some method for forming a corpus of texts, and at least one text parse is obtained, associated with said statement; wherein said database contains at least a plurality of text parses, each of which is associated with one or more of said statements.

[0042] In another preferred embodiment of the present invention, there is provided a method executable by a processor of a computing device for generating a text entry, in which at least one a statement obtained using any of the above, and generate a syntactically and semantically correct sentence including the statement or its derivative

[0043] In another preferred embodiment of the present invention, a method for generating a database of text records, executable by a processor of a computer device, is provided, which includes at least obtaining a plurality of text records by means of said method for generating text records and recording the obtained text records in a database; the approval is obtained using any of said methods for generating a corpus of texts.

[0044] In another preferred embodiment of the present invention, a method for generating natural language text, executable by a processor of a computer device, is provided, which comprises at least receiving a text entry from a database formed by said method for forming a database including an assertion of text entries, after which generating natural language text, including at least said text entry and a text that is not said entry.

[0045] In another preferred embodiment of the present invention, a method for generating natural language text, executable by a processor of a computer device, is provided, which at least identifies a natural language text entry, including at least a main entity and at least an associated end entity linked to it, and a statement linked to the main entity, and then obtains from a database formed by said method for forming a database of text entries including the statement, a text entry including this statement, and generates natural language text, including at least the identified text entry, and a text entry obtained from said database

[0046] In another preferred embodiment of the present invention, there is provided a computer device for automated processing of natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of automated processing of natural language text.

[0047] In another preferred embodiment of the present invention, there is provided a computer device for generating a text corpus, comprising: at least: a processor; memory containing program code that, when executed by the processor, causes the processor to perform actions of any of the aforementioned methods for generating a text corpus.

[0048] In another preferred embodiment of the present invention, a computer device is provided for pre-training, or training, or re-training a classification model, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of pre-training, or training, or re-training a classification model.

[0049] In another preferred embodiment of the present invention, a computer device is provided for pre-training, or training, or re-training a clustering model, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method of pre-training, or training, or re-training a clustering model.

[0050] In another preferred embodiment of the present invention, there is provided a computer device for classifying text parsing, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said text parsing classification method.

[0051] In another preferred embodiment of the present invention, a computing device for generating a database is provided, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of any said method of generating a database.

[0052] In another preferred embodiment of the present invention, a computing device is provided for generating a query to a database, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a query to a database.

[0053] In another preferred embodiment of the present invention, there is provided a computer device for selecting text parses, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for selecting text parses.

[0054] In another preferred embodiment of the present invention, there is provided a computer device for generating a text entry, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a text entry.

[0055] In another preferred embodiment of the present invention, a computer device is provided for generating a database of text entries including an approval, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating a database of text entries including an approval.

[0056] In another preferred embodiment of the present invention, there is provided a computer device for generating natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform the actions of any said method for generating natural language text.

[0057] In another preferred embodiment of the present invention, there is provided a machine-readable storage medium, including a non-transitive machine-readable storage medium, containing program code that, when executed by a processor, causes the processor to perform the actions of any said method.

[0058] The following are embodiments of the present invention, revealing examples of its implementation in specific embodiments. However, the description itself is not intended to limit the scope of rights granted by this patent. Rather, it should be understood that the claimed invention may also be implemented in other ways that incorporate different elements and conditions, or combinations of elements and conditions similar to those described herein, in combination with other existing and future technologies.

[0059] In this description, a patent document, without limitation, means a patent or a patent application, that is, as a rule, a specific text consisting of at least three segments: the first segment, which is the patent claims (patent claims), the second segment, which is the description (description, specification) and the third segment, which is the abstract (abstract), while, without limitation, the numbering of the segments above is not given in the order of their sequence in the text in natural language, but only for ease of presentation and association in the present document, as will be shown below. Preferably, but without limitation, the claims of patent documents do not refer to new chemical compounds and other similar objects created for the first time, since with respect to such objects, the claim of a technical result does not require identification.

[0060] In Fig. 1, by way of example and not limitation, an exemplary flow chart of a method 100 for automated processing of natural language text is shown. The method 100 comprises at least the following steps: a step 101 of identifying a natural language text that includes at least three segments; a step 102 of identifying the segments; a step 103 of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step 104 of marking up only one part to be parsed in the selected first segment and a step of marking up at least one part to be parsed in the selected second segment and / or in the selected third segment; a step 105 of parsing the marked parts by means of semantic-syntactic parsing;step 106 of extracting from the part of the selected first segment subjected to semantic-syntactic parsing at least the main entity of the first segment and at least an associated entity associated with the main entity of the first segment, wherein at least one of the associated entities is an associated terminal entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement;step 107 of associating said statement with said main entity. For example, but not limited to, step 101 of identifying natural language text ensures reliable identification of natural language text, for example, but not limited to, a patent document. For example, but not limited to, the text segments of a patent document may be represented by separate natural language texts, for which, if necessary, they are first associated, for example, by assigning them some relatedness features, for example, unique identifiers, the decoding of which makes it possible to determine the relatedness of the segments with each other. Such methods and techniques for associating entities in data storage systems are widely known in the prior art and, accordingly, are not described in detail below. For example, but not limited to, a patent document may represent a single text, typically having sequential page numbers;in this case, preliminary association of segments with each other is not required.

[0061] Preferably, but not limited to, step 102 involves identifying segments of the natural language text; for example, but not limited to, in the case where the natural language text is a patent document, at least the patent claims, which is the first segment in the context of the present invention, the description, which is the second segment in the context of the present invention, and the abstract, which is the third segment in the context of the present invention, are identified. Without limitation, in the case where the natural language text is a non-patent document, the segments are identified in a similar manner. In this case, without limitation, it should be primarily assumed that the first segment includes a set of entities, while the second and third segments include at least one or more statements characteristic of said set of entities.An example of a suitable non-patent document may be, for example, but not limited to, a scientific article in which the role of the first segment may be assigned to the abstract, and the roles of the second and third segments, respectively, to the main text and conclusion.

[0062] Preferably, but not limited to, step 103 comprises selecting the identified segments, including, but not limited to, selecting the first segment as comprising the set of entities and one of the second segment or the third segment. In this case, but not limited to, the choice of which of the second segment or the third segment to select depends on the probability with which the claim will be present in a particular segment. More specifically, but not limited to, the presence of the claim in a particular segment will generally depend on the approach to the requirements imposed on the natural language text. For example, but not limited to, when the natural language text is a patent document, then, depending on the patent system in which it is formed, the third segment (abstract) will at least contain a claim about the technical result.For example, but not limited to, Russian patent documents are typically published with an abstract formulating the technical result. Moreover, as of the filing date of the application, the preparation of the abstract for the patent application is the responsibility of the examiner, who is competent in formulating the technical result, thus eliminating any inaccuracies in its presentation. Meanwhile, in US patent documents, the abstract is typically prepared by the applicant themselves, resulting in abstracts in US patent documents lacking a clear structure, and the presence of a statement of the technical result is not guaranteed. Thus, for example, without limitation, when choosing between the second and third segments in the US case. It would be more efficient to select the second segment, since at least the Background of the Invention or Brief Summary of Invention sections are likely to contain a claim about the technical result, as required by MPEP 608.01(c) and 608.01(d). At the same time, for example, but not limited to, for each natural language text, a particular segment can be selected based on the presence of the claim itself. For this purpose, for example, but not limited to, for each natural language text, a semantic-syntactic analysis can be performed, for example, using a semantic parser, the technology of which is widely represented in the prior art and, accordingly, is not described below; based on the results of the semantic-syntactic analysis, it becomes possible to determine the specific part of the natural language text containing the sought-after claim.Preferably, without limitation, the first segment is preliminarily excluded from the natural language text, as it obviously does not contain assertions, and the presence of which during semantic-syntactic analysis can distort the results.

[0063] Preferably, but not limited to, step 104 involves annotating the parts of the selected segments. Preferably, but not limited to, only one part of the selected first segment to be parsed is annotated, namely, the part containing the set of entities. For example, without limitation, in the case where the natural language text is a patent document, the first sentence, i.e., the first claim, is selected as the annotated part, as it is most likely to contain the required set of entities, i.e., the set of essential features. Furthermore, without limitation, there are situations where the set of entities (features) in the first claim, in addition to the essential features, also includes nonessential features.In this case, such a selected first segment may be subject to additional analysis after finding the technical result statement in the second or third segment and verifying which features do not affect the achievement of the technical result. Based on the results of such additional analysis, the marked portion is purified of unnecessary entities. In this case, preferably, but not limited to, the marked and optionally purified portion of the selected first segment to be parsed may be subjected to analysis to determine the presence of a generic part (the first part) and a distinctive part (the second part). For such an analysis, for example, but not limited to, it may be sufficient to determine the linking word separating the generic part from the distinctive part. At the same time, not every patent system requires that an independent claim be drafted with the distinctive part. parts, which leads to the fact that the generic part is not clearly highlighted and not cut off by a linking word; in such a case, for example, without limitation, a preliminary semantic-syntactic analysis of the entire first segment can be carried out in order to identify, at least, the main entity of the first segment and the main entities of the entities associated with the main entity of the first segment, which will be discussed in detail below.Typically, but not limited to, when the natural language text is a patent document, the primary entity will be defined as a generic concept, and the associated entities will be individual features, such as, but not limited to, method steps or product parts. The resulting set of primary entities and associated entities can be subjected to a familiarity analysis, which can determine whether the entire set is generic or only a subset of the features constitutes a generic part. For example, but not limited to, the familiarity analysis can be performed automatically, based on a pre-prepared text corpus, for each of which a first segment has been identified and a semantic-syntactic analysis has been performed on the part to be analyzed.In this case, but not limited to, separation may be performed either manually or automatically, for example, using a pre-trained classification model and / or clustering model. In this case, but not limited to, the selected second segment or the selected third segment is labeled in such a way as to obtain the desired statement with the highest probability. Typically, the statement of the desired statement is preceded by a certain text structure, or the statement of the desired statement includes certain keywords. Without limitation, several parts of the selected second segment or the selected third segment to be analyzed may be labeled in this manner. In this case, but not limited to, labeling may be performed either manually or automatically, for example, using a pre-trained classification model and / or clustering model.Preferably, but not limited to, step 105 comprises parsing the marked parts by means of semantic-syntactic parsing using, for example, but not limited to, the mentioned syntactic parser.

[0064] Preferably, but not limited to, step 106 comprises extracting from the parsed portions of the first segment at least a main entity of the first segment and at least an associated entity associated with the main entity of the first segment, wherein at least one of the extracted associated entities is an associated leaf entity. For example, not Without limitation, in the case where the first segment is a patent claim, the first independent claim will be subject to semantic-syntactic parsing; in such a case, the main entity will be the generic concept, which is the root of the parse tree, the associated entities will be the features of the technical solution, which are the nodes of the parse tree, and at least one associated entity will be an associated leaf entity, that is, it will be a feature with only one edge, that is, it will be a leaf node (leaf) of the parse tree. In this case, without limitation, the parse tree can be constructed according to the principle of an abstract syntax tree (AST, Abstract Syntax Tree), that is, it can be cleared of all nonessential features, as well as repetitions, anaphors and other elements that do not affect the assertion, that is, in the case of a patent claim, which cannot affect the technical result.An example of redundant elements obtained after parsing is, for example, but not limited to, a comma or other separator, which are useful directly during parsing, for example, as links or boundaries, but are redundant for the text corpus when it is used, for example, for training a classification model and / or a clustering model.Furthermore, for example, but not limited to, when the first segment has been divided into a generic part and a distinctive part, the principal entity of the first segment, after said analysis, will be found in the first part, that is, in the generic part, along with the entities associated with it from the generic part; while, without limitation, the second part, which is the distinctive part, will then have the principal entity of the second part and the entities associated with it, which, nevertheless, are also all associated with the principal entity of the first segment, since all the individual features revealed by means of analysis directly or indirectly belong to only one generic concept.In this case, for example, but not limited to, each associated entity extracted in this manner may be iteratively subjected to its own semantic-syntactic analysis, during which the corresponding main entities of the associated entities and the associated associated entities of the associated entities will be discovered. That is, each associated entity may have multiple nested associated entities, up to one or more associated terminal entities that can no longer be subjected to semantic-syntactic analysis, since they are too simple features. In this case, for example, but not limited to, lemmatization may be performed to reduce words to their original forms. In addition, without limitation, for each entity, including the main entity, as well as including any associated entities, A generalization operation can be performed, in which the entity can be elevated to the level of a method or product, which corresponds to a certain function determined by the attributes of the entity. In this case, preferably, but not limited to, at least one statement is extracted from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic analysis. For example, without limitation, in the case where the natural language text is a patent document, upon semantic-syntactic analysis of paragraph 0010 of this description in its original version, two statements will be obtained: "precise generation of a text corpus" and "automated generation of a text corpus."For example, but not limited to, in the same case, during the semantic-syntactic analysis of paragraph 0011, three assertions may be obtained: "implementation of the purpose—generation of a text corpus," "precise generation of a text corpus," and "automated generation of a text corpus." Preferably, but not limited to, at step 107, at least one extracted assertion is associated with the primary entity of the first segment, and thus the assertion is associated with any associated entities related to the primary entity of the first segment.In this case, for example, but not limited to, some resulting statements, such as "implementation of purpose," are of no practical use, as they cannot be subsequently used in a text corpus to train a classification and / or clustering model, for example, to predict a statement corresponding to a new set of entities. This would lead to any classifier based on the classification and / or clustering model attributing any set of entities to such a statement, as such a statement is inherent in all resulting text parses. For this reason, it is preferable, but not limited to, to clean up the resulting "statement-main entity" pairs, removing redundant statements.

[0065] Thus, preferably, without being limited to, it becomes possible to obtain a corpus of texts, for which, by means of said method 100, at step 107, by associating said statement with said main entity of the first segment (and, as a consequence, with all associated entities related to the main entity), a plurality of pairs of texts are obtained, each of which includes at least one statement related to the main entity of the first segment, and a corpus of texts is formed from the obtained pairs of texts. Alternatively, without being limited to, it becomes possible to obtain a corpus of texts, which, by means of said method 100, at step 107, by associating said statement with said main entity the second part of the part to be parsed marked in the first segment (and, as a consequence, with all associated entities related to the main entity of the second part of the part to be parsed marked in the first segment) receive a plurality of pairs of texts, each of which includes at least one statement related to the main entity of the second part of the part to be parsed marked in the first segment, and form a corpus of texts from the obtained pairs of texts.Such resulting text corpora, for example, but not limited to, may be used, either jointly or separately, for pre-training, or training, or additional training of a classification model and / or a clustering model, in which, at least, the classification model and / or the clustering model are trained to classify a text parse as related to any said statement, upon receipt at the input of the classification model and / or the clustering model of a text parse containing, at least, a main entity and an associated terminal entity associated with the main entity, wherein the training is carried out on one of the said text corpora, jointly or separately. In this case, without limitation, the main entity in this way is understood to be, for example, the main entity of the first segment or the said main entity of the second part of the part of the first segment to be parsed.To train a classification model and / or a clustering model, for example, but not limited to, a sample from a text corpus is provided, after which the classification model and / or clustering model is trained on such a sample. In this case, for example, but not limited to, the classification model is based on one of the following: logistic regression, support vector machine, decision trees method, random forest method, naive Bayes classifier, k-nearest neighbors method, neural network, boosting, gradient boosting, bagging, group argument accounting method; in this case, the clustering model is based on one of the following: k-means method, density-based spatial clustering of applications with noise (DBSCAN), hierarchical clustering method, spectral clustering method.In this case, without limitation, any suitable known or future known classification model and / or clustering model may be used, provided that when choosing one, it is necessary to proceed from the fact that such a model provides the possibility of obtaining the same weights for different statements. This is primarily due to the fact that the same text analysis may in fact correspond to several different statements. For example, without limitation, the same invention, characterized by the same set of essential features, may achieve different Technical results, each of which may be suitable and lead to a solution to its own technical problem, which is due to the fact that the formulation of a technical problem is primarily related to the level of technology, or even the prototype, in which the technical problem is formulated. In this regard, it is preferable that the classification model and clustering model employed ensure the fundamental possibility of classifying a text parse as corresponding to several assertions simultaneously. Moreover, for example, but not limited to, the assertion weights do not necessarily have to be identical for a decision to classify a text parse as corresponding to these assertions; in fact, without limitation, classifying a text parse as corresponding to several assertions simultaneously can be ensured by defining an acceptable proximity of assertion weights.

[0066] Thus, without limitation, when the corresponding classification model or clustering model is obtained, it becomes possible to provide a method for classifying text parsing, which at least involves feeding to the input of the obtained classification model and / or clustering model a text parsing containing at least a main entity and an associated terminal entity associated with the main entity, and classifying said text parsing as associated with any said statement obtained using any said method for forming a text corpus. In this case, without limitation, the main entity in this way is understood to be, for example, the main entity of the first segment or the said main entity of the second part of the part of the first segment to be parsed.

[0067] The classified text analyses can subsequently be stored in the database of the automated natural language processing system 200. For this purpose, without limitation, a method for creating a database is preferably provided, which involves identifying a plurality of said classified text analyses, associating them with one or more corresponding statements, and creating a database containing at least a plurality of said text analyses, each of which is associated with one or more said statements. From the database thus created, it becomes possible to obtain text analyses corresponding to any statement selected by the user.For this purpose, preferably, without limitation, a method for generating a query to a database is provided, in which a query to a database is generated, containing at least one statement obtained using any of the mentioned methods for generating a corpus of texts; wherein the mentioned database. Contains at least a plurality of text parses, each associated with one or more of the aforementioned assertions. For example, but not limited to, a query containing the assertion "accurate text corpus generation" may be provided, in response to which at least one corresponding text parse will be selected. However, depending on the assertion, too many text parses may be selected, which may render the query irrelevant. To improve query relevance, for example, but not limited to, the query may be supplemented with an additional assertion, which will inevitably reduce the resulting selection of text parses.Thus, preferably, without being limited, a method for selecting text analyses is provided, in which a query is generated and sent to a database, containing at least some statement obtained using any mentioned method for forming a text corpus, and at least one text analysis associated with said statement is obtained; wherein said database contains at least a plurality of text analyses, each of which is associated with one or more mentioned statements.

[0068] Moreover, without limitation, the resulting method preferably provides new capabilities for generating natural language texts, in particular, but not limited to, texts such as patent document descriptions. Preferably, without limitation, a method for generating a text entry is provided, which at least one statement obtained using any of the aforementioned methods for generating a text corpus is obtained and a syntactically and semantically correct sentence is generated, including the statement or a derivative thereof. Moreover, without limitation, a text entry is understood to be a portion of a text, but not the entire text, and a text, accordingly, is understood to be a collection, including heterogeneous text entries.In this case, without limitation, the said assertion may be present in the generated text record both in unchanged form and in a derived form, i.e., when the assertion is subject to modifications and changes in order to ensure, at least, the consistency of the generated sentence. In this case, preferably, without limitation, the obtained text records can be placed in a database of text records including the assertion. For this purpose, preferably, without limitation, a method for creating a database of text records including the assertion is provided, in which, at least, a plurality of text records are obtained by means of the said method of generating a text record and the obtained text records are written into the database; wherein the assertion is obtained using any of the said methods of generating a corpus. texts. For example, but not limited to, such a database can subsequently be used as a source of standardized records when generating natural language text. For this purpose, for example, but not limited to, a method for generating natural language text is provided, which involves at least obtaining a text record from a database of text records formed by said generation method, including an assertion, and then generating natural language text, which includes at least said text record and a text that is not said record. More specifically, but not limited to, the text that is not said record may be, for example, but not limited to, text entered by the user.That is, without limitation, a method for generating text in a natural language is provided, which involves at least identifying a text entry in a natural language, including at least a main entity and at least an associated leaf entity linked thereto, and a statement linked to the main entity, after which a text entry including this statement is obtained from said database of text entries including the statement, and generating text in a natural language, including at least the identified text entry and the text entry obtained from said database. In this case, without limitation, the identified entry is a text entry entered by a user. More specifically, without limitation, the text entry entered by a user is a syntactically and semantically correct sentence.More specifically, but not limited to, the user-entered input constitutes an independent patent claim. This enables, for example, but not limited to, the generation of natural language text based on user-entered text and including a portion of text not actually entered by the user, i.e., generated automatically. Furthermore, without limitation, the text generation itself is accomplished using methods and means known in the art, such as NLP processors, which are accordingly not described in detail below.

[0069] Thus, preferably, without being limited, as shown in Fig. 2, a computer device 201 may be provided, in which, in the context of the present invention, at least one of or any combination of: a computer device for automated processing of natural language text, a computer device for generating a text corpus, a computer device for pre-training, or training, or re-training a classification model, a computer device for pre-training, or training, or re-training a model clustering, a computer device for classifying text parsing, a computer device for creating a database, a computer device for generating a query to a database, a computer device for selecting text parsings. Such a computer device 201 most typically comprises at least: one or more processors 2011; a memory 2012 containing program code that, when executed by the processor 2011, causes the processor 2011 to perform actions of any of the aforementioned methods described with reference to Fig. 1 for automated processing of text in natural language, and / or a method for creating a text corpus, and / or a method for pre-training, or training, or further training of a classification model, and / or a method for pre-training, or training, or further training of a clustering model, and / or a method for classifying text parsing, and / or a method for creating a database, and / or a method for generating a query to a database, and / or a method for selecting text parsings.At the same time, such a computing device can be implemented as a thin client, which will mean that all, or at least most, computing operations are performed on the system server 202, which thus also contains at least a processor 2021 and memory 2022, which are thus essentially similar to, respectively, the processor 2011 and memory 2012.By way of example, but not limitation, the memory 2012, 2022 (the computer-readable storage medium 2012, 2022) may include non-volatile memory (NVRAM); random access memory (RAM); read-only memory (ROM); electrically erasable programmable read-only memory (EEPROM); flash memory or other memory technologies; CDROM, digital versatile disc (DVD) or other optical or holographic storage media; magnetic cassettes, magnetic tape, magnetic disk storage device or other magnetic storage devices; as well as any other storage medium that can be used to store and encode the desired information. In this case, without being limited, the memory 2012, 2022 includes a storage medium based on a computer storage device in the form of volatile or non-volatile memory, or a combination thereof.In this case, without limitation, exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, and so on. In this case, without limitation, the machine-readable storage medium 2012, 2022 (memory 2012, 2022) is non-transient (permanent, non-transitive), so it does not include a temporary (transitive) propagating signal. In this case, without limitation, the memory 2012, 2022 may store an exemplary environment in which, using computer commands or codes, including those stored in the memory 2022 of the server 202, the following can be implemented: A procedure for automated processing of text in natural language, and / or a procedure for forming a text corpus, and / or a procedure for pre-training, or training, or further training of a classification model, and / or a procedure for pre-training, or training, or further training of a clustering model, and / or a procedure for classifying text parsing, and / or a procedure for forming a database, and / or a procedure for generating a query to a database, and / or a procedure for selecting text parsing. In this case, without limitation, the computer device 201, when not a thin client, comprises one or more processors 2011, which are intended to execute computer commands or codes stored in the memory 2012 of the device 201 for the purpose of ensuring the execution of the said procedures.In this case, without being limited, the server 202 may be essentially similar to the computer device 201, when it is not a thin client, and, accordingly, contain one or more processors 2021, which are designed to execute computer commands or codes stored in the memory 2022 of the server 202 in order to ensure the execution of the mentioned procedures. In this case, without being limited, the system 200 may also include a database (DB) 203. DB 203 may be, but is not limited to: a hierarchical DB, a network DB, a relational DB, an object DB, an object-oriented DB, an object-relational DB, a spatial DB, a combination of two or more of the listed DBs, and the like.In this case, without limitation, the DB 203 at least stores classified text parsing associated with the corresponding statements, and can also store data for analysis, classification models, clustering models and other information in the memory 2021, 2022 or in a suitable memory of another computing device associated with the computing device 201 and / or with the server 202, which may be, but is not limited to, a memory similar to any memory 2021, 2022, as shown earlier, and which can be accessed via the server 202. In addition, without limitation, the server 202 is provided, which, in addition to the previously described functions, stores and facilitates the manipulation of computer commands or codes previously described in this document, which, accordingly, are not further described. In this case, without limitation, the server 202, in addition to the previously described functions, can ensure the regulation of data exchange in the system 200.In this case, without limitation, the exchange of data within the system 200 is carried out thanks to one or more data transmission networks 204. In this case, without limitation, the data transmission networks 204 may include, but are not limited to, one or more local area networks (LAN) and / or wide area networks (WAN), or may represent an information telecommunications network, the Internet, or an Intranet, or a virtual private network. (VPN), or a combination thereof, and the like. In this case, without limitation, the server 202 also has the ability to provide a virtual computing environment to ensure interaction between the system components. In this case, without limitation, the network 204 serves to ensure interaction between the computer device 201, the server 402 and, optionally, the database 203. In this case, without limitation, the non-thin client computer device 201 and / or the server 202 can be connected to the database 203 directly, using wired and wireless communication methods and techniques known in the art, which, accordingly, are not described in detail further, or, without limitation, the database 203 can be implemented in the memory 2012, 2022. In this case, without limitation, a suitable non-thin client computer device 201 can perform the role of the server 202 of the system 200 for other computer devices 201, which are thin clients.In this case, most typically, without limitation, the components of the computer device 201, the components of the server 202 are connected to each other, including by means of some data bus.

[0070] The present description of the implementation of the claimed invention demonstrates only particular embodiments and does not limit other embodiments of the claimed invention, since possible other alternative embodiments of the claimed invention, not going beyond the scope of the information set out in this application, should be obvious to a specialist in the given field of technology, having the usual qualifications, for whom the claimed invention is intended.

Claims

CLAUSES OF THE INVENTION 1. A computer device for automated processing of natural language text, comprising at least: a processor; a memory containing program code that, when executed by the processor, causes the processor to perform actions of a method for automated processing of natural language text, which method consists of at least performing the steps of: a step of identifying a natural language text that includes at least three segments; a step of identifying the segments; a step of selecting at least a first segment and at least a second segment and / or at least a third segment of the natural language text; a step of marking up in the first segment only one part to be parsed, and a step of marking up in the selected second segment and / or in the selected third segment at least one part to be parsed; a step of parsing the marked parts by means of semantic-syntactic parsing;a step of extracting from the part of the first segment subjected to semantic-syntactic parsing at least the main entity of the first segment, and at least an associated entity related to the main entity of the first segment, wherein at least one of the associated entities is an associated end entity, and extracting from each part of the selected second segment and / or the selected third segment subjected to semantic-syntactic parsing at least one statement; a step of associating said statement with said main entity of the first segment.

2. The device according to claim 1, characterized in that the first segment, the second segment, and the third segment are pre-combined to obtain identifiable text in a natural language.

3. The device according to claim 1, characterized in that the first segment, the second segment, and the third segment are pre-associated to obtain identifiable text in a natural language.

4. The device according to paragraph 1, characterized in that the part to be parsed marked in the first segment is the first sentence in a natural language.

5. The device according to claim 4, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

6. The device according to claim 1, characterized in that the said part to be parsed, marked in the first segment, is divided into a first part and a second part, each of which is subjected to semantic-syntactic parsing; wherein the main entity of the first segment and all associated entities connected with it are extracted from the first part; wherein the main entity of the second part, and at least an associated entity connected with the main entity of the second part.

7. The device according to claim 6, characterized in that each extracted associated entity is subjected to semantic-syntactic parsing and at least the main entity of the associated entity and at least a nested associated entity associated with the main entity of the associated entity are extracted for each associated entity, wherein at least one of the nested associated entities is a nested associated end entity.

8. The device according to paragraph 7, characterized in that the main essence of the second part is associated with the main essence of the first segment.

Citation Information

Patent Citations

  • Retrieval device, retrieval method, and program

    JP2014056457A

  • Use of depth semantic analysis of texts on natural language for creation of training samples in methods of machine training

    RU2636098C1

  • Template-based structured document classification and extraction

    US20180144042A1

  • Text recognition method, electronic device, and storage medium

    US20210383064A1

  • Generating semantic vector representation of natural language data

    US20230306203A1