Method for transforming a structured data array
The method transforms structured data arrays into domain-specific constructs through multiple data structure formations, addressing the limitations of existing technologies to enhance indexing and searching efficiency in specialized domains.
Patent Information
- Application Number
- PCT/RU2025/050193
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-29
- Filing Date
- 2025-06-29
- Publication Date
- 2026-01-02
AI Technical Summary
Existing methods for transforming structured data arrays, such as those described in Russian patent 2544739, are inadequate for creating domain-specific constructs necessary for high-precision searches in specialized domains like jurisprudence, and are limited to natural language texts, requiring pre-selection of texts from documents.
A method for converting structured data arrays that includes forming multiple data structures to identify and group semantic parts, integrate them with system features, and create linguistic and logical constructs, ultimately forming target domain-specific constructs suitable for high-precision searches.
Enhances the efficiency of processing digitized documents by enabling accurate indexing and searching, particularly in specialized domains, by forming domain-specific constructs that facilitate reliable identification of subject roles.
Smart Images

Figure RU2025050193_02012026_PF_FP_ABST
Abstract
Description
METHOD FOR TRANSFORMING A STRUCTURED DATA ARRAY
[0001] AREA OF TECHNOLOGY
[0002] The group of inventions relates to solutions in the field of processing data arrays, in particular, to solutions in the field of processing digitalized documents containing information objects such as text and / or images, and can be used to transform a digitalized document for efficient indexing of its elements and accurate searching.
[0003] LEVEL OF TECHNOLOGY
[0004] A method for transforming a structured data array is known from Russian patent 2544739 (Igor Petrovich Rogachev), published on 20.03.2015 (D1). In the method for transforming a structured data array containing text in a natural language, known from D1, a first data structure of the structured data array is formed (101) from the final data structure of the structured data array. A database of logical connections of logical sections of elements of the first data structure is formed (102). A second data structure of the structured data array is formed (103). A database of semantic parts of the logical sections of elements of the second data structure is formed (104). Grammatically and orthographically correct semantic parts of the logical sections of elements of the second data structure are formed (105) by means of linguistic transformations over the said semantic parts. The final data structure of the structured data array is formed (106).
[0005] The method known from D1 transforms a structured data array to produce logical constructs with grammatically and orthographically correct semantic parts, which can be useful for improving the accuracy of information retrieval in non-specialized data arrays, such as fiction or journalism. However, the transformations known from D1 do not provide the formation of either basic domain-specific constructs, much less target domain-specific constructs that ensure reliable identification of subject roles, which is required for high-precision searches in specialized domains, such as jurisprudence. Moreover, the method known from D1 is only suitable for working with natural language texts, meaning it requires pre-selection of texts from the document.
[0006] The solution known from D1 can be accepted as the closest analogue.
[0007] DISCLOSURE OF THE INVENTION
[0008] The technical problem solved by the claimed invention is the creation of inventions that eliminate the shortcomings of their closest analogues and thus offer increased efficiency in processing digitized documents for subsequent indexing of their elements, their processing, and searches using them. Another technical problem solved by the claimed invention is the expansion of the arsenal of technical means—methods for converting structured data arrays containing information objects from digitized documents.
[0009] The technical result achieved by implementing the claimed invention, in addition to fulfilling its intended purpose, is the elimination of the shortcomings of the closest analogue and thus an increase in the efficiency of processing digitalized documents for the subsequent indexing of its elements, their processing and conducting searches using them.
[0010] The technical result is achieved due to the fact that a method is provided, executable by a processor or processors of a computer device, for converting a structured array of data, containing at least information objects of a digitized document, which are separate blocks of the information content of the digitized document, representing text information objects, and / or representing visual information objects, and / or representing text-visual information objects; the method is characterized by: performing step 1001 of forming a first data structure, characterized by: performing step 10011 of identifying the original data structure, at which the elements of the original data structure, representing the information objects of the digitized document, are identified;performing a step 10012 of identifying elements of the first data structure, in which the elements of the first data structure are identified, representing the semantic parts of the information objects of the digitized document, as well as the identification data of the elements of the first data structure, representing for each said semantic part the values of said semantic parts and the serial numbers of said semantic parts in the digitized document, and the first data structure is formed; and the method is characterized by: performing a step 1002 of forming a database of system features of the semantic parts, in which the system features of the semantic parts contained in the first data structure are identified, representing the format system characteristics of such semantic parts and functional ones; system characteristics of such semantic parts, as well as the values of the corresponding mentioned system characteristics, for identifying semantic parts with structural systemic features, and / or for identifying semantic parts with logical systemic features, and / or for identifying semantic parts with informational systemic features, and / or for identifying semantic parts with requisite systemic features; and forming from such identified systemic features a database of systemic features of the semantic parts; performing step 1003 of forming a second data structure, in which a second data structure is formed containing elements of the second data structure representing integrated semantic parts of the information objects of the digitized document,representing grouped semantic parts with matching system features contained in the first data structure or grouped semantic parts with unique system features contained in the first data structure, as well as containing identification data of integrated semantic parts, which are non-repeating varieties of said semantic parts with matching system features or said semantic parts with unique system features, the values of said semantic parts with matching system features or said semantic parts with unique system features, and serial numbers in the digitized document of said semantic parts with matching system features or said semantic parts with unique system features,wherein such mentioned semantic parts with coinciding systemic features or such mentioned semantic parts with unique systemic features constitute the mentioned integrated semantic parts; by performing step 1004 of forming a third data structure, in which a third data structure is formed, containing as elements of the third data structure linguistic constructions, which are the integrated semantic parts of the information objects of the digitized document contained in the second data structure, wherein such integrated semantic parts have the systemic features of text logical semantic parts, and also containing identification data of linguistic constructions, which are the values of linguistic constructions and the ordinal numbers of linguistic constructions in the digitized document,wherein the linguistic constructions in the digitized document are one of the following linguistic constructions or a combination of the following linguistic constructions: ordinary linguistic constructions of the third data structure, which are linguistic sentences, special linguistic constructions, constructs of the third data structure, which are lists or enumeration headings, reconstructed linguistic constructs of the third data structure, which are tables containing at least two rows and two columns, wherein at least one row contains column headings and / or, respectively, at least one column contains row headings;performing step 1005 of forming a fourth data structure, in which a fourth data structure is formed that contains, as elements of the fourth data structure, linguistic sentences of the fourth data structure formed from elements of the third data structure and representing either linguistic sentences that are ordinary linguistic constructions of the third data structure, or linguistic sentences obtained by transforming special linguistic constructions of the third data structure, or linguistic sentences obtained by recreating from reconstructed linguistic constructions of the third data structure;wherein the fourth data structure also includes identification data of the linguistic sentences of the fourth data structure, which are the values of the linguistic sentences of the fourth data structure and the ordinal numbers of the linguistic sentences of the fourth data structure in the fourth data structure; by performing step 1006 of forming the fifth data structure, in which a fifth data structure is formed, containing as elements of the fifth data structure the text elements of the linguistic sentences of the fourth data structure, as well as identification data of the text elements of the linguistic sentences of the fourth data structure, which are the values of the corresponding text elements of the corresponding linguistic sentences of the fourth data structure and the ordinal numbers of the corresponding text elements in the corresponding linguistic sentences of the fourth data structure;performing step 1007 of forming a database of linguistic-logical-subject features, in which linguistic, logical and subject features of the text elements of linguistic sentences of the fourth data structure are identified and a database of linguistic-logical-subject features is formed from the identified features; performing step 1008 of forming a sixth data structure, in which a sixth data structure is formed containing elements of the sixth data structure that are components of a simple judgment, which are components of the corresponding simple judgments, wherein the simple judgments are simple judgments of the corresponding linguistic sentences of the fourth data structure, and also containing identification data of the said; components of a simple proposition, representing for each corresponding component of a simple proposition from the sixth data structure the type of such corresponding component of a simple proposition, the value of such corresponding component of a simple proposition and the ordinal number of such corresponding component of a simple proposition in the corresponding linguistic sentence; by performing step 1009 of forming a seventh data structure, in which a seventh data structure is formed, containing elements of the seventh data structure, representing simple propositions of the corresponding linguistic sentences of the fourth data structure, and also containing identification data of such simple propositions, representing the values of the corresponding simple propositions and the ordinal numbers of the corresponding simple propositions in the corresponding linguistic sentences of the fourth data structure;performing step 1010 of forming an eighth data structure, in which an eighth data structure is formed, containing elements of the eighth data structure, representing final judgments of the corresponding linguistic sentences of the fourth data structure, formed from the mentioned simple judgments of the corresponding linguistic sentences of the fourth data structure, and also containing identification data of the final judgments, representing the values of the final judgments and the ordinal numbers of the final judgments in the eighth data structure;performing step 1011 of forming a ninth data structure, in which a ninth data structure is formed, containing elements of the ninth data structure, representing basic constructs of the subject area, formed from data, including data of the said sixth data structure formed as a result of performing step 1008, wherein the formation of basic constructs of the subject area is carried out on the basis of data on a formalized model of the basic construct of the subject area and data on a formalized model of the logical construct of the judgment, and also containing identification data of the basic constructs of the subject area, representing the values of the basic constructs of the subject area and the ordinal numbers of the basic constructs of the subject area in the ninth data structure;performing step 1012 of forming the final data structure, in which the final data structure is formed, containing elements of the final data structure, which are target constructions of the subject area, formed from the basic constructions of the subject area, contained in the ninth data structure, wherein the formation of the target constructions of the subject area is carried out on the basis of data on the formalized model of the target construction of the subject area, and also containing identification data of the target constructions; subject area, representing the values of the target constructs of the subject area and the ordinal numbers of the target constructs of the subject area in the final data structure.
[0011] BRIEF DESCRIPTION OF DRAWINGS
[0012] Illustrative embodiments of the present invention are described below in detail with reference to the accompanying drawings, which are incorporated herein by reference, and in which:
[0013] Fig. 1 shows, by way of example and not limitation, a general flow chart of the steps of method 1000.
[0014] Fig. 2, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1001.
[0015] Fig. 3 shows, by way of example and not limitation, the general structure of the original data structure.
[0016] Fig. 4 shows, by way of example and not limitation, the general structure of the generated first data structure.
[0017] Fig. 5, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1002.
[0018] Fig. 6 shows, by way of example and not limitation, the general structure of the generated database of system features.
[0019] Fig. 7, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1003.
[0020] Fig. 8 shows, by way of example and not limitation, the general structure of the generated second data structure.
[0021] Fig. 9, by way of example and not limitation, shows a general flow chart of the steps of step 1004.
[0022] Fig. 10 shows, by way of example and not limitation, the general structure of the generated third data structure.
[0023] Fig. 11, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1005.
[0024] Fig. 12 shows, by way of example and not limitation, the general structure of the generated fourth data structure.
[0025] Fig. 13, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1006.
[0026] Fig. 14 shows, by way of example and not limitation, the general structure of the generated fifth data structure.
[0027] Fig. 15, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1007.
[0028] Fig. 16, as an example, but not a limitation, shows the general structure of the generated database of linguistic-logical-subject features.
[0029] Fig. 17, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1008.
[0030] Fig. 18 shows, by way of example and not limitation, the general structure of the generated sixth data structure.
[0031] Fig. 19, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1009.
[0032] Fig. 20 shows, by way of example and not limitation, the general structure of the generated seventh data structure.
[0033] Fig. 21, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1010.
[0034] Fig. 22 shows, by way of example and not limitation, the general structure of the generated eighth data structure.
[0035] Fig. 23, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1011.
[0036] Fig. 24 shows, by way of example and not limitation, the general structure of the generated ninth data structure.
[0037] Fig. 25, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1012.
[0038] Fig. 26 shows, by way of example and not limitation, the general structure of the generated final data structure.
[0039] Fig. 27 illustrates, by way of example and not limitation, an exemplary diagram of a system 2000.
[0040] IMPLEMENTATION OF THE INVENTION
[0041] The possible embodiments of the present invention described in this section are presented as non-limiting examples of specific embodiments of the present invention, which are intended to be illustrative and non-limiting in all respects. Alternative embodiments of the present invention, within the scope of its protection, are obvious to those of ordinary skill in the art for whom this invention is intended. In Fig. 1, by way of example, but not limitation, a general diagram of the execution of the steps of method 1000 is shown, which is a method executed by the processor or processors of a computer device for converting a structured array of data containing at least information objects of a digitized document, which are separate blocks of the information content of the digitized document, which are textual information objects, and / or which are visual information objects, and / or which are textual and visual information objects;the method is characterized by: performing step 1001 of forming a first data structure, in which a first data structure is formed containing elements of the first data structure, representing the semantic parts of the information objects of the digitized document, as well as identification data of such semantic parts, representing the values of such semantic parts and the serial numbers of such semantic parts in the digitized document;by performing step 1002 of forming a database of system features of semantic parts, in which system features of the semantic parts contained in the first data structure are identified, representing the format system characteristics of such semantic parts and the functional system characteristics of such semantic parts, as well as the values of the corresponding mentioned system characteristics, in order to identify semantic parts with structural system features, and / or to identify semantic parts with logical system features, and / or to identify semantic parts with information system features, and / or to identify semantic parts with requisite system features; and a database of system features of semantic parts is formed from such identified system features;performing step 1003 of forming a second data structure, in which a second data structure is formed, containing elements of the second data structure, which are integrated semantic parts of information objects of the digitized document, which are grouped semantic parts contained in the first data structure with matching system features or grouped semantic parts contained in the first data structure with unique system features, and also containing identification data of the integrated semantic parts, which are non-repeating varieties of the said semantic parts with matching system features or the said semantic parts with unique system features, the values of the said semantic parts with matching system features or the said semantic parts with unique system features, and serial numbers in the digitized document of the said semantic parts; with matching systemic features or said semantic parts with unique systemic features, wherein such said semantic parts with matching systemic features or such said semantic parts with unique systemic features constitute said integrated semantic parts; by performing step 1004 of forming a third data structure, in which a third data structure is formed, containing as elements of the third data structure linguistic constructions, which are the integrated semantic parts of the information objects of the digitized document contained in the second data structure, wherein such integrated semantic parts have the systemic features of textual logical semantic parts, and also containing identification data of linguistic constructions, which are the values of linguistic constructions and the ordinal numbers of linguistic constructions in the digitized document,wherein the linguistic constructions in the digitalized document are one of the following linguistic constructions or a combination of the following linguistic constructions: ordinary linguistic constructions of the third data structure, which are linguistic sentences, special linguistic constructions of the third data structure, which are lists or enumeration headings, reconstructed linguistic constructions of the third data structure, which are tables containing at least two rows and two columns, wherein at least one row contains column headings and / or, respectively, at least one column contains row headings; performing step 1005 of forming a fourth data structure, in which a fourth data structure is formed, containing as elements of the fourth data structure linguistic sentences of the fourth data structure,formed from elements of the third data structure and representing either linguistic sentences that are ordinary linguistic constructions of the third data structure, or linguistic sentences obtained by transforming special linguistic constructions of the third data structure, or linguistic sentences obtained by recreating from reconstructed linguistic constructions of the third data structure; wherein the fourth data structure also includes identification data of the linguistic sentences of the fourth data structure, which are the values of the linguistic sentences of the fourth data structure and the ordinal numbers of the linguistic sentences of the fourth data structure in the fourth data structure; performing step 1006 of forming the fifth data structure, in which the fifth data structure is formed, containing text elements as elements of the fifth data structure, linguistic sentences of the fourth data structure, as well as identification data of the text elements of the linguistic sentences of the fourth data structure, which are the values of the corresponding text elements of the corresponding linguistic sentences of the fourth data structure and the serial numbers of the corresponding text elements in the corresponding linguistic sentences of the fourth data structure; by performing step 1007 of forming a database of linguistic-subject features, in which linguistic, logical and subject features of the text elements of the linguistic sentences of the fourth data structure are identified and a database of linguistic-logical-subject features is formed from the identified features;performing step 1008 of forming a sixth data structure, in which a sixth data structure is formed, containing elements of the sixth data structure, which are components of a simple judgment, which are components of corresponding simple judgments, wherein the simple judgments are simple judgments of corresponding linguistic sentences of the fourth data structure, and also containing identification data of said components of a simple judgment, which are, for each corresponding component of a simple judgment from the sixth data structure, the type of such corresponding component of a simple judgment, the value of such corresponding component of a simple judgment and the ordinal number of such corresponding component of a simple judgment in the corresponding linguistic sentence;performing step 1009 of forming a seventh data structure, in which a seventh data structure is formed, containing elements of the seventh data structure, representing simple judgments of the corresponding linguistic sentences of the fourth data structure, and also containing identification data of such simple judgments, representing the values of the corresponding simple judgments and the ordinal numbers of the corresponding simple judgments in the corresponding linguistic sentences of the fourth data structure;performing step 1010 of forming an eighth data structure, in which an eighth data structure is formed, containing elements of the eighth data structure, which are final judgments of the corresponding linguistic sentences of the fourth data structure, formed from the mentioned simple judgments of the corresponding linguistic sentences of the fourth data structure, and also containing identification data of the final judgments, which are the values of the final judgments and the ordinal numbers of the final judgments in the eighth data structure; performing step 1011 of forming a ninth data structure, in which a ninth data structure is formed, containing elements of the ninth data structure,; representing the basic constructs of the subject area, formed from data including data of the said sixth data structure formed as a result of the execution of step 1008, wherein the formation of the basic constructs of the subject area is carried out on the basis of data on the formalized model of the basic construct of the subject area and data on the formalized model of the logical construct of the judgment, and also containing identification data of the basic constructs of the subject area, representing the values of the basic constructs of the subject area and the ordinal numbers of the basic constructs of the subject area in the ninth data structure;performing step 1012 of forming the final data structure, in which the final data structure is formed, containing elements of the final data structure, which are target constructions of the subject area, formed from the basic constructions of the subject area, contained in the ninth data structure, wherein the formation of the target constructions of the subject area is carried out on the basis of data on the formalized model of the target construction of the subject area, and also containing identification data of the target constructions of the subject area, which are the values of the target constructions of the subject area and the ordinal numbers of the target constructions of the subject area in the final data structure;
[0042] In Fig. 2, by way of example, but not limitation, a general diagram of the execution of the steps of step 1001 of forming the first data structure 2 is shown. Step 1001 is characterized by: performing step 10011 of identifying the original data structure 1, in which the elements 11 of the original data structure 1 are identified, which represent the information objects 11 of the digitized document 1; performing step 10012 of identifying the elements 21 of the first data structure 2, in which the elements 21 of the first data structure 2 are identified, which represent the semantic parts 21 of the information objects 11 of the digitized document 1, as well as the identification data of the elements 21 of the first data structure 2, which are, for each mentioned semantic part 21, the values 211 of the mentioned semantic parts 21 and the serial numbers 212 of the mentioned semantic parts 21 in the digitized document 1, and the first data structure 2 is formed.
[0043] Fig. 3, by way of example, but not limitation, shows the general structure of the initial data structure 1 (initial structured data array 1; digitized document 1), from which the elements of the first data structure SMD are formed. Preferably, without limitation, the data source for forming the initial structured data array 1 (initial SMD 1) is, by way of example, but not limitation, a digitized document (electronic document), that is, a document, Transformed from its traditional, inherent form into digital form in the form of an electronic data file suitable for recording on electronic media. Preferably, but not limited to, a digitized (electronic) document is a systematically organized combination of individual blocks of information content (information objects). Preferably, but not limited to, each individual information object possesses semantic and logical completeness. Said blocks of information content are intended for various purposes, such as, but not limited to, providing information through the perception of textual images, and / or providing information through the perception of visual images, and / or providing information through the perception of textual and visual images.
[0044] Preferably, without limitation, the original data structure characterizing the original SMD 1 contains elements 11 representing at least information objects 11 of the original structured data array I (digitized document 1). Preferably, without limitation, the information objects 11 consist of components of information objects and may contain any number of said components, by way of example, but not limitation, individual line objects, and / or list objects, and / or tabular text objects, and / or visual objects.In this case, for example, without being limited to, the said components are preliminarily formed in information objects by 11 different technical methods and means, by way of example, but not limitation, with the help of linguistic tools (for example, without being limited to, several sentences are combined into a paragraph using known from the prior art or any other suitable technical means, which are not further described in detail), or, without being limited to, with the help of known from the prior art or any other suitable technical means of various text editors of electronic documents, which are not further described in detail (for example, without being limited to, one piece of information in the document is separated from another using the formatting tools “tabulation” or “line feed”, and similar actions), while, without being limited to, such formation of the said components in information objects. II is carried out, including but not limited to, using automation tools known from the prior art, as well as, but not limited to, using machine learning and neural network technologies. Preferably, but not limited to, the components of information objects serve as a technical tool for providing information of a certain type. By way of example, but not limitation, the components of information objects may differ in the type of information provided. To provide information through the perception of text images, For example, but not limited to, components of text-type information objects are used, such as, but not limited to, inline objects and / or list objects and / or tabular text objects, and similar text objects. To provide information through the perception of visual images, components of visual information objects are used, such as, but not limited to, logos, images, drawings, handwritten text, photographs, and similar visual objects. For example, but not limited to, to provide information through the perception of text-visual images, components of text-visual information objects are used, such as, but not limited to, consisting of combinations of the aforementioned components of text and visual information objects.Preferably, without being limited, in the original data structure 1, the elements 11, by way of example but not limitation, may be named as "IO1", "IO2", "IOZ", "IOp", where n>1 is the ordinal number of the element I in the digitized document. Preferably, without being limited, all of the above-mentioned information objects 11 of the digitized document 1 in the original data structure represent individual information objects 11, prepared in advance and placed in the original data structure 1 in the form of a structured array of individual information objects 11 of the digitized document 1. In this case, without being limited, such preparatory actions may be carried out by any method known from the prior art and, accordingly, are not described further. Preferably, without being limited, the identification of the elements 11 of the original data structure during step 10011 is performed by identifying the features of the information object of the digitized document.Such a feature may be, by way of example, but not limitation, a feature for grouping successive components of information objects in digitized document 1, which may be, but is not limited to, a control character (tag, control command) "line separator" ("line feed", "line break"), and / or "tab". As a rule, with the help of such control characters, all successive information objects 11 are separated from both sides. In this case, for example, but not limited to, the first information object 11 in digitized document 1 does not have such a feature before the first component of the information object, and the last information object 11 in digitized document 1 does not have such a feature after the last component of the information object. Elements 11 identified by the specified methods form elements 11 of the original data structure of the SMD 1.Preferably, but not limited to, such preparatory steps can be carried out by any method known in the art and, accordingly, are not further described. Preferably, but not limited to, the initial method is thus used. structured data array 1 is an array of information objects of a digitalized (electronic) document, consisting of elements 11 of the digitalized document 1, identified at step 1011.
[0045] In Fig. 4, by way of example, but not limitation, the general structure of the generated first data structure 2 of the SMD is shown. Preferably, without limitation, the first data structure contains elements 21 of the first data structure, which are semantic parts (SP) 21 of the information objects 11 of the digitized document 1 and identification data of such semantic parts, which are values 211 of the semantic parts 21 and serial numbers 212 of the semantic parts 21 in the digitized document 1. Each semantic part 21 of the information object 11 of the digitized document 1 is, by way of example, but not limitation, a separate information object 11 or a part of the information object 11 with a feature of homogeneity of the components of the information object (CHI).As an example, but not limitation, a sign of such homogeneity of the KIO may be that the KIO of a separate information object 11 belongs to only one type of data, for example, without limitation, to string text data, or to list text data, or to tabular text data, or to visual data. The meaning 211 of the semantic part 21 is, by way of example, but not limitation, a set of letters, and / or a set of words, and / or a set of digits, and / or a set of numbers, and / or a set of punctuation marks, and / or a set of other signs and symbols, and / or a table, as well as, for example, without limitation, a logo, and / or a picture, and / or a drawing, and / or handwritten text, and / or a photo and the like, of which the semantic part 21 consists. The serial number 212 of the semantic part 21, by way of example, but not limitation, is the serial number of the semantic part 21 of the information object 11 in the digitalized document 1.In the first data structure, the elements 21, by way of example but not limitation, may be referred to as "SC1", "SC2", "SC3", "SCp", where n>1 is the ordinal number of the element in the digitized document. Preferably, without limitation, the identification of the elements 21 of the first data structure during step 10012 is performed by analyzing the components of information objects (CIO) for each information object 11 of the digitized document 1. Preferably, without limitation, the essence of the analysis is to check the presence of signs of homogeneity of the components of information objects (CIO) in each individual information object 11. If all the CIOs of an individual information object 11 belong to the same type of data, then one semantic part 21 is formed from such information object 11. If the CIOs of an individual information object 11 belong to different types of data, then such information objects 11 are divided into fragments. In this case, without limitation, from. Each fragment of a separated information object 11 forms its own separated semantic part 21, except in the case where consecutive textual information objects of any kind contain a visual semantic part 21. In this case, such visual semantic parts form a nested visual semantic part 21, and fragments from textual information objects separated by a visual semantic part 21 are combined together to form their own composite semantic part 21.In this case, without limitation, in such a composite semantic part 21, instead of the embedded visual semantic part 21 removed from it, a replacement text is formed (for example, without limitation, if the embedded visual semantic part 21 was a picture, then the word “PICTURE” is used as a replacement text), placing it in the composite semantic part 21 in the place from which the mentioned embedded visual semantic part 21 was excluded, thus restoring the integrity of the sequence of the textual KIO of the individual information object 11, from which the composite semantic part 21 was formed.Preferably, without being limited to, the identification of the value 211 of the semantic part 21 during step 10012 is carried out by registering the content of the components of the information objects (a set of letters, and / or a set of words, and / or a set of digits, and / or a set of numbers, and / or a set of punctuation marks, and / or other signs and symbols, and / or a table, and / or a logo, and / or a picture, and / or a drawing, and / or handwritten text, and / or a photo, and the like), of which the semantic part 21 consists. Preferably, without being limited to, the identification of the ordinal number 212 of the semantic part 21 of the first data structure during step 10012 is carried out by calculating the location of the components of the information object of which the semantic part 21 consists in the digitized document 1. In this case, without being limited to, since the number of elements 21 can significantly exceed the number of elements 11, then the use of the ordinal numbers of elements 11 implies, for example, without being limited to, the following calculation procedure, consisting of of two stages.At the first stage, for each element 21, a preliminary number of the type "X.1" is obtained, in which the index "X" indicates the number of the element 11 from which the element 21 is formed. If the element 21 is a divided semantic part 21, a prefabricated semantic part 21 or a nested visual semantic part 21, then for such elements 21, a preliminary number of the type "X.Y" is obtained, in which the index "X" indicates the number of the element 11 from which the element 21 is formed, and the index "Y" indicates the serial number of the part of the divided, prefabricated or nested visual semantic part 21 in the sequence of components of the information object, established on the basis of the calculation of the position of the first KIO of the divided, prefabricated or nested visual semantic part 21 in the element 11. At the second stage, the nested numbering obtained in this way (for example, without limitation: 1.1, 2.1, 2.2, 3.1, 4.1, 4.2, 4.3, 5.1) allows to form the ordinal numbers of the elements of the 21st century. In digitized document 1, starting with the number "1" for the element with the embedded number "1.1." Further, without limitation, the next sequential number is assigned to element 21, either with the embedded number "1.2," or, if none exists, with the embedded number "2.1." And so on, until all embedded numbers are converted into sequential numbers for element 21 in digitized document 1. Preferably, but without limitation, such analysis for identifying and generating elements 21 can be performed by any method known in the prior art and, accordingly, is not described in detail below. For example, without limitation, such analysis can be performed traditionally by a linguist, or using a software algorithm of a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of immutable rules (an IT solution "by rules").Moreover, if there are a sufficient number of examples, it is possible to perform such an analysis using a statistical processor (neural network, AI systems) by applying neural network (AI systems) training technology.
[0046] In Fig. 5, by way of example, but not limitation, a general diagram of the execution of the steps of step 1002 of forming a database of system features 20 of semantic parts 21 is shown, in which, preferably, without limitation, system features of the semantic parts 21 contained in the first data structure 2 are identified, which represent format system characteristics 213 of such semantic parts 21 and functional system characteristics 213 of such semantic parts 21, as well as values 2131 of the corresponding mentioned system characteristics 213, for identifying semantic parts 21 with structural system features, and / or for identifying semantic parts 21 with logical system features, and / or for identifying semantic parts 21 with information system features, and / or for identifying semantic parts with requisite system features; and a database of system features 20 of semantic parts 21 is formed from such identified system features;Preferably, without limitation, step 1002 is characterized by: performing step 10021 of forming systemic features of semantic parts 21, in which, for the systemic analysis, for each semantic part 21 of information objects 11 of the digitized document 1, identification data of said semantic part 21 are provided and systemic characteristics 213 are obtained for all said semantic parts 21, as well as values 2131 of said systemic characteristics; performing step 10022 of forming a database of systemic features 20 of semantic parts 21, in which a database of systemic features of semantic parts 21 of information objects 11 of the digitized document 1 is formed, wherein the systemic feature of said semantic part 21 is all. the mentioned system characteristics 213 obtained for the mentioned semantic part 21, having the values 2131 of the system characteristics 213.
[0047] In Fig. 6, by way of example, but not limitation, the general structure of the generated database of system features 20 (DBSF 20) is shown, which is DBSF 20 of the semantic parts 21 of the information objects 11 of the digitized (electronic) document 1. Preferably, without being limited, the system characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1 contain format characteristics and functional characteristics. In this case, preferably, without being limited, the set of values 2131 of all system characteristics 213 of each semantic part 21 of the information objects 11 in the digitized document 1 is a distinctive system feature of each semantic part 21 of the information objects 11 in the digitized document 1. The format characteristics, preferably, without being limited, indicate the format features of the semantic parts 21 of the information objects 11 of the digitized document 1, which can be classified, as an example,but not limited by the nesting level, for example, as "genus-type-subtype". In this case, without being limited, the format genus of the said semantic parts 21 preferably has the following meanings: text information object of the document, visual information object of the document; without being limited, the format type of the said semantic parts 21 preferably has the following meanings: line (machine-readable text (words, numbers)), list (machine-readable text (words, numbers)), tabular (machine-readable text (words, numbers)), pictorial (photo, drawing, logo, picture), handwritten (non-machine-readable text (words, numbers)); without being limited, the format subtype of the said semantic parts 21 preferably has the following meanings: ordinary linguistic construction (linguistic sentence), special linguistic construction (linguistic construction that combines elements of a linguistic construction, as well as a method of organizing data and (or) visual information objects),reconstructed linguistic construction (a method of organizing data that has a logical basis on the basis of which it is possible to recreate a linguistic construction equivalent to information in a system of organized data), a non-linguistic construction. The functional characteristics, preferably without limitation, indicate a set of functional attributes of the semantic parts 21 of the information objects 11 of the digitized document 1, among which the following can be distinguished, by way of example, but not limitation: attributes of the structural hierarchy of the document (structural system attributes); attributes of the main semantic information (logical system attributes); attributes of accompanying or, technical information (information system features); document metadata features (requisite system features). Without being limited to, the functional characteristics indicate the functional role of the said semantic parts 21, which represent, by way of example, but not limitation, the following functional roles of the semantic parts 21: structural, i.e., the semantic part 21 with structural system features; logical, i.e., the semantic part 21 with logical system features; informational, i.e., the semantic part 21 with information system features; requisite, i.e., the semantic part 21 with document metadata features (with requisite system features).
[0048] The formation of system characteristics 213 and their values 2131 for the semantic parts 21 of the information objects 11 of the digitized document 1, preferably without limitation, is performed at step 10021 by means of a comprehensive structural-linguistic analysis of each semantic part 21 of the information objects 11 of the digitized document 1, representing, as an example, but not limitation, an analysis of the components of information objects (CIO) of the information object 11 from which the semantic part 21 is formed, from the point of view of the said system features. Based on the results of the system analysis of the said CIO of the semantic part 21, preferably without limitation, the formation of system characteristics 213 is performed and their entry at step 10022 into the B DSP 20 in the form of a list of system characteristics 213 with the values of these characteristics 2131. For example, but not limited to, for the semantic part 21 with the value 211 "Chapter 1.In the "General Provisions" section, the following system characteristics 213 with the values of system characteristics 2131 may be system features: format type "text KIO"; format type "linear machine-readable text"; format subtype "usual linguistic construction"; functional features "features of the structural hierarchy of the document"; functional role "structural". Such an analysis may be performed by any method known from the prior art and, accordingly, is not described in detail below. For example, without limitation, such an analysis may be performed traditionally by a linguist, or using a software algorithm of a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of unchangeable rules (IT solution "by rules").Moreover, given a sufficient number of examples, such an analysis can be performed using a statistical processor (neural network, AI system) by employing neural network (AI system) training technology. Preferably, but not limited to, this can be based on the identified systemic characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1 and their values 2131 ultimately form a database of system features 20 (DB DSP 20), which is DB DSP 20 of semantic parts 21 of information objects 11 of digitized document 1. In this case, system characteristics 213 of semantic parts 21 of information objects 11 of digitized document 1 and their values 2131 form system features of the mentioned semantic parts 21
[0049] In Fig. 7, by way of example, but not limitation, a general diagram of the execution of the steps of step 1003 of the formation of the second data structure 3 of the SMD is shown. Preferably, without limitation, step 103 is characterized by: the execution of step 10031 of the identification and formation of elements 31 of the second data structure 3, in which the elements 31 of the second data structure 3 are identified and formed, representing the integrated semantic parts 31 of the information objects 11 of the digitized document 1, and representing the grouped semantic parts 21 contained in the first data structure 2 with matching system features or the grouped semantic parts 21 contained in the first data structure 2 with unique system features, as well as the identification data of the integrated semantic parts 31, representing non-repeating varieties of the said semantic parts 21 with matching system features or the said semantic parts 21 with unique system features,the values of the said semantic parts 21 with matching system features or the said semantic parts 21 with unique system features, and the serial numbers in the digitalized document 1 of the said semantic parts 21 with matching system features or the said semantic parts 21 with unique system features, wherein such said semantic parts 21 with matching system features or such said semantic parts 21 with unique system features constitute the said integrated semantic parts 31; by performing step 10032 of forming the second data structure 3, in which the second data structure 3 is formed from the identified and generated elements 31 of the second data structure 3 and their identification data.
[0050] In Fig. 8, as an example, but not limitation, the general structure of the generated second data structure 3 of the SMD is shown. Preferably, without limitation, the second data structure 3 contains elements of the second data structure 31, which are integrated semantic parts 31 of the digitized (electronic) document 1 and identification data of the integrated semantic parts 31 of the digitized (electronic) document 1. Elements 31 of the second data structure 3 of the SMD are integrated semantic parts 31 of information objects 11 of the digitalized document 1 and identification data of such integrated semantic parts 31, representing the values 311 of the integrated semantic parts, serial numbers 312 of the integrated semantic parts and non-repeating varieties 313 of the integrated semantic parts in the digitalized document 1. The integrated semantic parts 31 of the information objects 11 of the digitalized document 1, preferably, without being limited to, are the grouped semantic parts 21 contained in the first data structure 2 with matching system features or semantic parts 21 with unique system features.In this case, the systemic features of semantic parts 21 are understood to be the systemic characteristics 213 and their values 2131, unique systemic features are understood to be those systemic features that are found in the digitized document 1 only in one semantic part 21, and matching systemic features are understood to be those systemic features that are found in the digitized document 1 in at least two semantic parts 21. As an example, but not limitation, the systemic features of element 31 may be as follows: "Text KIO. Tabular machine-readable text. Reconstructible linguistic construction. Features of the main semantic information. Logical functional role."Preferably, without being limited to, the values 311 of the integrated semantic parts 31 are the values 211 of the semantic parts 21 with the same or with unique system features, wherein such semantic parts 21 with the same or with unique system features constitute the said integrated semantic parts 31. Preferably, without being limited to, the ordinal number 312 of the integrated semantic part 31 are the ordinal numbers 212 of the semantic parts 21 with the same or with unique system features, wherein such semantic parts 21 with the same or with unique system features constitute the said integrated semantic parts 31. In the second data structure 3, the elements 31 are named without unique names and, by way of example, but not limitation, can be named as "ISCh1", "ISCh2", "ISChZ", "ISChn", where n>1 is the ordinal index of the element 31 in the digitized document 1, starting with "1" for each element 31 in digitized document 1.Preferably, without being limited to, the non-repeating varieties 313 of the integrated semantic parts 31 in the digitalized document 1 are non-repeating varieties of the semantic parts 21 from which the elements 31 are formed. In this case, non-repeating varieties of the semantic parts 21 are understood to mean all the unique systemic features of the semantic parts 21 contained in the digitalized document 1. As an example, but not limitation, the non-repeating variety 313 of the integrated semantic parts. 31 are unique (if element 31 contains only one element 21) or matching (if element 31 contains two or more elements 21) system features of elements 21 from which elements 31 are formed, namely: “Text KIO. Tabular machine-readable text. Reconstructed linguistic construction. Features of the main semantic information. Logical functional role.” Preferably, without limitation, the identification of elements 31 of the second data structure 3 of the SMD within the framework of step 1031 is carried out by means of a comparative analysis of the values 2131 of the system characteristics 213 of the semantic parts 21 of information objects 11 of the digitized document 1. In this case, during the identification of element 31, the serial number 312 of element 31 and the non-repeating variety 313 of element 31 are simultaneously identified.By way of example, but not limitation, the order of identification of elements 31, as well as the serial number 312 of element 31 and the non-repeating variation 313 of element 31, may be as follows. At the first stage, from the list of elements 21 of the first data structure 2, all unique elements 21 are identified, that is, such elements 21 that have unique (non-repeating) values of system characteristics 2131. At the second stage, all identified unique elements 21 are recognized as elements 31 and numbered with ordinal numbers, starting with “1”, where the number “1” is received by such element 21 that has the minimum ordinal number 212, the number “2” is received by such element 21 that has the ordinal number 212 greater than that of the element 21 recognized as element 31 with the ordinal number “1”, but at the same time less than that of the remaining unique elements 21. And so on, until all unique elements 21 receive their ordinal number as recognized elements 31.In the third stage, a search is performed among the elements 21 of the first data structure that are not yet recognized as elements 31 for those elements 21 whose systemic characteristics are identical to the elements 21 already identified as elements 31 in the second stage. The elements 21 thus identified are joined to the corresponding elements 31 (a group of elements 21 with systemic characteristics identical to the systemic characteristics of the identified element 21), thereby identifying all elements 21 of the first data structure with one or another element 31 formed in the second stage.
[0051] Without limitation, the identification of the system features of elements 21, if necessary, is carried out by organizing a request in the DBSP 20 and obtaining the values 2131 of the system characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1. In this case, without limitation, as described earlier, the system features of element 21 are at least the format and functional characteristics of the semantic parts 21 of the information objects 11 of the digitized document 1. Without limitation, the identification of the values 311 of the elements 31 is carried out after the identification of all the elements 31 of the second data structure, that is, after recognizing all the elements 21 of the first data structure 2 as one or another identified element 31 (an element 31 with one or another serial number and with unique or coinciding system features of the elements 21 of which the identified element 31 consists). In this case, the values 311 of the integrated semantic part 31 will be the values 211 of all the elements 21 of which the identified element 31 of the second data structure 3 consists. As an example, but not limitation, the determination of the serial numbers 312 of the elements 31 of the second data structure can be demonstrated as follows. In the first stage, for the element 31, which contains the semantic part 21 with the minimum value of the serial number 212, the serial number "1" will be set.In the second stage, the remaining unnumbered elements 31 are searched for to find an element 31 that contains a semantic part 21 with a serial number 212 greater than that of the element 31 with the serial number "1" but less than that of the other elements 31 whose serial numbers have not yet been identified. Such an element 31 is assigned the serial number "2". In the third stage, the analysis described in the second stage is repeated to identify the element 31 with the serial number "3", and so on until the number of remaining unnumbered elements 31 of the second document data structure equals "1". Then the last unnumbered element 31 is assigned a serial number equal to the maximum established serial number of the numbered elements 31 plus "1". Such a comparative analysis for identifying and generating elements 31 can be performed by any method known in the prior art and, accordingly, is not described in detail below.For example, but not limited to, such analysis can be performed traditionally by a linguist, or using a software algorithm in a linguistic (syntactic) processor, or using traditional programming by solving problems using the coding of immutable rules (a "rule-based" IT solution). Furthermore, given a sufficient number of examples, such analysis can be performed using a statistical processor (neural network, AI system) through the use of neural network (AI system) training technology.
[0052] In Fig. 9, as an example, but not limitation, a general diagram of the execution of the steps of step 1004 of forming the third data structure 4 of the SMD is shown. Step 1004 is characterized by: performing step 10041 of identifying and forming elements 41 of the third data structure 4, in which elements 41 of the third data structure 4 are identified and formed, representing linguistic constructions 41, representing integrated semantic parts 31 of information objects 11 of the digitized document 1 contained in the second data structure 3, wherein such integrated semantic parts 31 have systemic features of textual logical semantic parts, and also containing identification data of linguistic constructions 41 representing values 411 of linguistic constructions 41 and ordinal numbers 412 of linguistic constructions 41 in the digitized document 1, wherein the linguistic constructions 41 in the digitized document 1 are one of the following linguistic constructions 41 or a combination of the following linguistic constructions 41: ordinary linguistic constructions 41 of the third data structure 4, which are linguistic sentences, special linguistic constructions 41 of the third data structure 4, which are lists or enumeration headings, reconstructed linguistic constructions 41 of the third data structure 4, which are tables,containing at least two rows and two columns, wherein at least one row contains column headers and, accordingly, at least one column contains row headers; by performing step 10042 of forming a third data structure 4, in which the third data structure 4 is formed from the elements 41 of the third data structure 4 identified and formed in step 10041 and their identification data.
[0053] In Fig. 10, by way of example, but not limitation, the general structure of the generated third data structure 4 of the SMD is shown. The third data structure 4 contains elements 41 of the third data structure 4, which are linguistic constructions 41 and identification data of the linguistic constructions 41, which are values 411 of the linguistic constructions 41, and serial numbers 412 of the linguistic constructions 41 in the digitized document 1. Preferably, without limitation, the elements 41 of the third data structure 4 of the SMD are linguistic constructions 41, wherein the linguistic constructions 41 are understood to be integrated semantic parts 31 of the digitized (electronic) document 1, contained in the second data structure 3, with systemic features of textual logical semantic parts (textual-logical integrated semantic parts 31).Textual-logical integrated semantic parts, by way of example, but not limitation, may have the following systemic attributes of textual logical semantic parts: values 2131 of systemic characteristics 213 of semantic parts 21, constituting element 31: "Textual KIO. Listed machine-readable text. Special linguistic construction. Attributes of the main semantic information. Logical functional role." Preferably, without limitation, the value 411 of linguistic. constructions 41 is identical to the meaning of the textual-logical integrated semantic parts (integrated semantic parts 31 of the information objects 11 of the digitalized document 1, contained in the second data structure 3, and having the systemic features of textual logical semantic parts). Preferably, without limitation, the ordinal number 412 of the linguistic constructions 41 is the ordinal number of the textual integrated semantic part in the digitalized document 1.
[0054] Preferably, without being limited to, the identification and formation of element 41 of the third data structure 4 of the SMD within the framework of step 10041 is carried out by means of a comparative analysis of the values 2131 of the system characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1 included in the elements 31 of the second data structure 3. The object of comparison in the values 2131 of the system characteristics 21 are the format and functional roles of the semantic parts 21. In this case, without being limited to, all elements 21 included in element 31 have the same format and functional roles. Therefore, in order to conduct a comparative analysis of element 31, it is sufficient to conduct a comparative analysis of one element 21 included in element 31.In the event that the values of the system characteristics of the analyzed element 21 contain the "Text format" and "Logical functional role", then the analyzed element 31, which includes the analyzed element 21, is recognized as a textual-logical integrated semantic part and is identified as element 41 (linguistic construction) and replenishes the third structure 3 of the SMD. Preferably, without limitation, the identification of the system features of elements 21, if necessary, is carried out by organizing a request in the DBSP 10, formed within the framework of step 1003, consisting of the identification data of the semantic parts 21, and obtaining the values 2131 of the system characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1. In this case, as was described earlier, the system features of element 21 are at least the format and functional characteristics of the semantic parts 21 of the information objects 11 of the digitized document 1.Preferably, without limitation, the meaning 411 of each element 41 (linguistic construction) is identical to the meaning 311 of the integrated semantic part 31, which is recognized as a textual-logical integrated semantic part and identified as element 41 (linguistic construction).
[0055] In the data structure, elements 41, by way of example, but not limitation, may be referred to as "LK1", "LK2", "LK3", "LKp", where n>1 is the serial number 412 of element 41 in the digitized document 1. By way of example, but not limitation, the determination of the serial numbers 412 of elements 41 of the third data structure 4 may be demonstrated as follows. In the first stage, for element 41, formed from the integrated semantic part 31 with the minimum ordinal number 312, will be assigned the ordinal number "1". In the second stage, in the remaining unnumbered elements 41, such an element 41 is found, formed from the integrated semantic part 31 with the ordinal number 312 greater than that of the element 41 with the ordinal number "1", but less than that of the other elements 41, the ordinal number of which has not yet been identified. Such an element 41 receives the ordinal number "2". In the third stage, the analysis described in the second stage is repeated to identify the element 41 with the ordinal number "3", and so on until the number of remaining unnumbered elements 41 of the second document data structure is equal to "1". Then the last unnumbered element 41 will receive an ordinal number equal to the maximum established ordinal number of the numbered elements 41 plus "1".Such a comparative analysis for identifying and generating elements 41 can be performed by any method known in the art and, accordingly, is not described in detail below. For example, but not limited to, such an analysis can be performed traditionally by a linguist, or using a software algorithm of a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of unchangeable rules (GG-solving "by rules"). Furthermore, given a sufficient number of examples, such an analysis can be performed using a statistical processor (neural network, AI system) by employing neural network (AI system) training technology.The formation of the third data structure 4 of the SMD during step 10042 is carried out by combining in one data structure the elements 41 of the third data structure 4, as well as their identification data, according to principles and methods known from the prior art, which, accordingly, are not described in detail further.
[0056] In Fig. 11, by way of example, but not limitation, a general diagram of the execution of the steps of step 1005 of forming the fourth data structure 5 of the SMD is shown. Preferably, without limitation, step 1005 is characterized by: performing a step 10051 of identifying and forming the first elements 51 of the fourth data structure 5, in which the first elements 51 of the fourth data structure 5 are identified and formed, as well as the identification data of the first elements 51 of the fourth data structure 5, representing for each first element of the fourth data structure 5 the value 511 of the first element 51 of the fourth data structure 5 and the serial numbers 512 of the first element 51 of the fourth data structure 5 in the fourth data structure 5; wherein the first elements 51 of the fourth data structure 5 are linguistic sentences formed from the elements 41 of the third data structure 4, containing ordinary linguistic constructions 41, by identifying linguistic sentences of the fourth data structure 5 with ordinary linguistic constructions 41 of the third data structure 4; by performing the step 10052 of identifying and generating second elements 52 of the fourth data structure 5, in which the second elements 52 of the fourth data structure 5 are identified and generated, as well as the identification data of the second elements 52 of the fourth data structure 5, representing for each second element 52 of the fourth data structure 5 the value 521 of the second element 52 of the fourth data structure 5 and the ordinal numbers 522 of the second element 52 of the fourth data structure 5 in the fourth data structure 5;wherein the second element 52 of the fourth data structure 5 are linguistic sentences formed from the elements 41 of the third data structure 4 containing special linguistic constructions 41 by transforming the special linguistic constructions 41 into linguistic sentences of the fourth data structure 5; wherein the special linguistic construction 41 may be a list or an enumeration heading; by performing the step 10053 of identifying and forming the third elements 53 of the fourth data structure 5, in which the third elements 53 of the fourth data structure 5 are identified and formed, as well as the identification data of the third elements 53 of the fourth data structure 5, representing for each third element 53 of the fourth data structure 5 the value 531 of the third element 53 of the fourth data structure 5 and the serial numbers 532 of the third element 53 of the fourth data structure 5 in the fourth data structure 5;wherein the third element 53 of the fourth data structure 5 are linguistic sentences formed from the elements 41 of the third data structure 4 containing reconstructed linguistic constructions 41 by means of recreating individual linguistic sentences of the fourth data structure 5 from the data contained in the reconstructed linguistic constructions 41; wherein the reconstructed linguistic constructions are tables possessing the features of a logical basis for organizing the data; by performing step 10054 of forming the fourth data structure 5, in which the fourth data structure 5 is formed from the first elements 51 of the fourth data structure 5, the second elements 52 of the fourth data structure 5, the third elements 53 of the fourth data structure 5 and their identification data.
[0057] Fig. 12 shows, by way of example and not limitation, the general structure of the generated fourth data structure 5. Preferably, without limitation, the fourth data structure 5 (fourth SD 5) comprises a first element 51 of the fourth SD 5, a second element 52 of the fourth SD 5 and a third element 53 of the fourth SD 5, representing linguistic sentences of the fourth data structure 5 and identification data of the linguistic sentences in the fourth data structure 5, representing: for the first element 51 of the fourth SD 5, by way of example, but not limitation, the values 511 of the first element 51 of the fourth SD 5 and the ordinal numbers 512 of the first element 51 of the fourth SD 5 in the fourth SD 5; for the second element 52 of the fourth SD 5, by way of example, but not limitation, the values 521 of the second element 52 of the fourth SD 5 and the ordinal numbers 522 of the second element 52 of the fourth SD 5 in the fourth SD 5; for the third element 53 of the fourth SD 5, by way of example, but not limitation, the values 531 of the third element 53 of the fourth SD 5 and the ordinal numbers 532 of the third element 53 of the fourth SD 5 in the fourth SD 5.Preferably, but not limited to, the elements 51, 52, 53 of the fourth data structure 5 are linguistic sentences formed from various linguistic constructions 41 contained in the third data structure 4. Preferably, but not limited to, the linguistic sentences of the fourth SD 5, identified with the usual linguistic construction 41, are elements 51 of the fourth data structure 5. In this case, but not limited to, the usual linguistic construction 41 is considered to be a syntactically related group of words, that is, a linguistic sentence. Preferably, but not limited to, the linguistic sentences formed by transforming special linguistic constructions 41 into linguistic sentences are elements 52 of the fourth data structure 5.In this case, without being limited to, a special linguistic construction 41 is considered to be a linguistic construction 41 that combines the features of a conventional linguistic construction 41 (a syntactically related group of words, a linguistic sentence) and a data organization system (a list, a table of one row or column), which in practice represents, by way of example, but not limitation, a heading of enumerations. Preferably, without being limited to, the linguistic sentences obtained by recreating individual linguistic sentences from the data contained in the reconstructed linguistic constructions 41 represent elements 53 of the fourth data structure 5. Moreover, without being limited to, the reconstructed linguistic constructions 41 are tables that possess the features of a logical basis for data organization. Without being limited to, such features for tables generally include the presence in the table of at least two rows and two columns.At least one row and / or one column must contain not data, but rather the names of the data in the rows and / or columns (row headings and / or column headings, respectively). Such tables have data organization characteristics that can be described by the following logical relationships. formulas: "IF <...>, THEN <. ..>", or "(IF <...> AND, IF <. ..>), THEN <...>". By way of example, but not limitation, for the formation of element 51, the values of system characteristics 2131 for elements 21 that make up element 41 may be as follows: "Text KIO. Linear machine-readable text. Standard linguistic construction. Features of the main semantic information. Logical functional role." By way of example, but not limitation, for the formation of element 52, the values of system characteristics 2131 for elements 21 that make up element 41 may be as follows: "Text KIO. List machine-readable text. Special linguistic construction. Features of the main semantic information. Logical functional role." By way of example, but not limitation, for the formation of element 53, the values of system characteristics 2131 for elements 21 that make up element 41 may be as follows: "Textual KIO. Tabular machine-readable text. Reconstructable linguistic construction."Characteristics of basic semantic information. Logical functional role."
[0058] In the fourth data structure 5, the elements 51, by way of example but not limitation, may be referred to as "LPx1", "LPx2", "LPx3", "LPxn", where x is the ordinal number of the semantic part 41 in the third data structure 4 that contains the ordinary linguistic construction 41 identified with the linguistic sentence 51, and n> 1 is the ordinal number of the element 51 (the linguistic sentence identified with the ordinary linguistic construction 41) in such semantic part 41, starting with "1". In the fourth data structure 5, the elements 52, by way of example but not limitation, may be referred to as "LPx1", "LPx2", "LPx3", "LPxn", where x is the ordinal number of the semantic part 41 in the third data structure 4 that contains the special linguistic construction 41 from which the linguistic sentence 52 is formed, and n>1 is the ordinal number of the element 52 (the linguistic sentence 52 formed from the special linguistic construction 41) in such semantic part 41, starting with "1".In the fourth data structure 5, the elements 53, by way of example but not limitation, may be referred to as "LPx1", "LPx2", "LPx3", "LPxn", where x is the ordinal number of the semantic part 41 in the third data structure 4 that contains the reconstructed linguistic construction 41 from which the linguistic sentence 53 is recreated, and n>1 is the ordinal number of the element 53 (the linguistic sentence 53 recreated from the data of the reconstructed linguistic construction 41) in such semantic part 41, starting with "1". Preferably, without limitation, since the numbering of all linguistic sentences 51, 52, 53 of the fourth data structure 5 is performed one by one for all linguistic sentences 51, 52, 53 of the fourth data structure 5, outside. Depending on whether the first, second or third element of the fourth data structure 5 is a separate linguistic sentence 51, 52, 53, then, on the basis of the established preliminary ordinal numbers of the "xn" formats, the ordinal numbers of all linguistic sentences 51, 52, 53 of the fourth data structure 5 in the fourth data structure 5 are established in the format "LP1", "LP2", "LP3", "LPy", where y is the ordinal number of the element 51, 52, 53 of the fourth data structure 5 in the fourth data structure 5. In this case, without being limited to, the smallest number is received by such linguistic sentences 51, 52, 53 of the fourth data structure 5, which have the smallest values of the preliminary ordinal number in the "xn" format. When establishing the ordinal number, the index "x" is considered the main number, and the index "n" is considered an additional one. The smallest serial number of the element 51, 52, 53 of the fourth data structure 5 is received by such an element 51, 52, 53 of the fourth data structure 5, which has the smallest index "x", and if this index is equal for several elements 51, 52, 53 of the fourth data structure 5, the smallest serial number is received by the element 51, 52, 53 of the fourth data structure 5 with the smallest value of the index "n" in the preliminary serial number. Preferably, without limitation, the identification and formation of the elements 51 of the fourth data structure within the framework of step 10051 is carried out by identifying in the values 511 of the element 51 of the fourth data structure 5 the features of the end of a linguistic sentence and the features of the beginning of a linguistic sentence.Such features are formed and stored in a special user database (SUD) and represent a list of text symbols (text elements), the presence of which in conventional linguistic constructions 41 of the third data structure 4 is a feature of the beginning or end of a linguistic sentence. By way of example, but not limitation, symbols (text elements) that are features of the beginning of a sentence may include: a word (number) with a capital letter; the first word (number) in the semantic part, and the like. By way of example, but not limitation, symbols (text elements) that are features of the end of a sentence may include: punctuation marks (period, semicolon) followed by a space; the last word (number, punctuation mark) in the semantic part, and the like.
[0059] Preferably, without limitation, for identifying elements 51, such elements 41 are identified which, according to the values 2131 of the system characteristics 213 of the semantic parts 21 included in the element 41, correspond to the mentioned requirements for the first element 51 of the fourth data structure 5. Further, in the elements 41 corresponding to the mentioned requirements, the features of the end of a linguistic sentence and the features of the beginning of a linguistic sentence are identified. Based on the results of identifying the features The element 41 is divided into elements 51, which are linguistic sentences 51 contained in the element 41, based on the end of the linguistic sentence and the beginning features of the linguistic sentence. Such an analysis of the elements 41 for identifying and forming the elements 51 can be performed by any method known in the art and, accordingly, is not described in detail below. For example, without limitation, such an analysis can be performed traditionally by a linguist, or using a software algorithm of a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of unchangeable rules (an IT solution "by rules"). Moreover, if there are a sufficient number of examples, it is possible to perform such an analysis using a statistical processor (neural network, AI systems) by applying neural network (AI system) training technology.Preferably, without limitation, the identification of the system characteristics of the semantic parts that make up the elements 41 of the third data structure 4 of the SMD and their values, if necessary, is carried out by organizing a request to the Database of System Features 20 (DB SSF 20), formed within the framework of step 1003 consisting of the identification data of the semantic parts that make up the element 41, and obtaining the values 2131 of the system characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1, of which the element 41 consists. In this case, as was described earlier, the system features of the element 51 are at least the mentioned values 2131 of the format and functional characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1, of which the elements 41 consist, which correspond to the requirements of the system features of the elements 51.
[0060] Preferably, without limitation, the identification and formation of elements 52 of the fourth data structure within the framework of step 10052 is carried out by identifying in the values 521 of element 52 of the fourth data structure 5 the features of the first part of the composite linguistic sentence and the features of the second part of the composite linguistic sentence. Such features are formed and stored in a special user database (SUD) and represent a list of text symbols (text elements), the presence of which in the electronic text data array (logical data array), consisting of special linguistic constructions, is a feature of the first and second parts of the composite linguistic sentence. As an example, but not limitation, the symbols (text elements) that are features of the beginning of the first part of the composite sentence may be: a word (number) with a capital letter; the first word (number) in the semantic part, and the like. As By way of example, but not limitation, the symbols (text elements) that mark the end of the first part of a compound sentence may be: a punctuation mark (colon) followed by a space or a line feed. By way of example, but not limitation, the symbols (text elements) that mark the beginning of the second part of a compound sentence may be: a word (number) with a lowercase letter; the preceding symbol is a punctuation mark (colon or semicolon). By way of example, but not limitation, the symbols (text elements) that mark the end of the second part of a compound sentence may be: a punctuation mark (semicolon or period) followed by a space or a line feed. A practical example of forming compound linguistic sentences from elements 41 that represent a list or enumeration heading can be demonstrated in the following example.If the listing heading (element 41) has the following text: "In order to carry out the acceptance and transfer of the goods, the buyer is obliged to provide: a receipt for the payment of the goods; a power of attorney from the buyer; a document confirming the identity of the authorized person.", then the following elements 52 can be formed from it: "In order to carry out the acceptance and transfer of the goods, the buyer is obliged to provide a receipt for the payment of the goods"; "In order to carry out the acceptance and transfer of the goods, the buyer is obliged to provide a power of attorney from the buyer."; "In order to carry out the acceptance and transfer of the goods, the buyer is obliged to provide a document confirming the identity of the authorized person." Preferably, without limitation, for the identification of elements 52, such elements 41 are identified that, according to the values 2131 of the system characteristics 213 of the semantic parts 21 included in the element 41, correspond to the aforementioned requirements for the second element 52 of the fourth data structure 5.Further, in the elements 41 corresponding to the mentioned requirements, the features of the beginning of the first part of the composite linguistic sentence and the features of the end of the first part of the composite linguistic sentence, as well as the features of the beginning of the second part of the composite linguistic sentence and the features of the end of the second part of the composite linguistic sentence are identified. Based on the results of identifying all the mentioned features for identifying the elements 52, the element 41 is first divided into parts of elements 52, from which elements 52 are formed, representing composite linguistic sentences contained in the element 41. Such an analysis of the elements 41 for identifying and forming the elements 52 can be performed by any method known from the prior art and, accordingly, is not described in detail below. For example, without limitation, such an analysis can be performed traditionally by a linguist, or with the help of a software algorithm of a linguistic (syntactic) processor, or on.based on traditional programming by solving problems using coding of immutable rules (IT solution "by rules"). Moreover, if there are a sufficient number of examples, it is possible to perform such an analysis using a statistical processor (neural network, AI systems) by applying the technology of training the neural network (AI systems). Preferably, without limitation, the identification of the system characteristics 213 of the semantic parts 21 that make up the elements 41 of the third data structure 4 of the SMD and their values, if necessary, is carried out by organizing a request to the database of system features 20 (DB DSP 20), formed within the framework of step 1003, consisting of the identification data of the semantic parts 21 that make up the element 41, and obtaining the values 2131 of the system characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1 that make up the element 41.In this case, as was described earlier, the system features of element 52 are at least the mentioned values 2131 of the format and functional characteristics 213 of the semantic parts 21 of the information objects 11 of the digitalized document 1, of which elements 41 consist, which correspond to the requirements of the system features of elements 52.
[0061] Preferably, without limitation, the identification and formation of element 53 of the fourth data structure within the framework of step 10053 is carried out by identifying in the values 531 of element 53 of the fourth data structure 5 of the SMD the features of the first part of the reconstructed linguistic sentence and the features of the second part of the reconstructed linguistic sentence. Such features are formed and stored in a special user database (SUD) and represent a list of electronic page code symbols (text-table symbols), the presence of which in the electronic text data array (logical data array), consisting of reconstructed linguistic constructions, is a feature of the first and second parts of the reconstructed linguistic sentence.By way of example, but not limitation, the symbols (text-table symbols) that are features of a table name may be: a text line above the table; a cell of the first row of the table containing text corresponds in width to all cells in the second row; and the like. By way of example, but not limitation, the symbols (text-table symbols) that are features of the names of fields (columns) of a table may be: symbols indicating the number of fields in the first row of the table is more than one; symbols indicating the number of fields in the second row of the table is more than one (if there is a first row of the table name); symbols indicating the name of table fields located on several rows, and the like. By way of example, but not. Limitations, symbols (text-table symbols) that are attributes of table row names may include: symbols indicating the number of table rows; symbols indicating the number of table rows that contain the names of table fields; symbols indicating the number of table rows that contain the names of table rows, and the like. By way of example, but not limitation, symbols (text-table symbols) that are attributes of table values may include: symbols indicating table cells that are not related to either the table name or the names of table fields (columns) or rows, and the like.
[0062] A practical example of the formation of reconstructed linguistic sentences from elements 41, which constitute a data table, can be demonstrated using the following example. Let's assume that Table 1 (element 41, containing data with the following system attribute values for the semantic parts - "Tabular machine-readable text. Reconstructed linguistic construction.") has the following form (Table 1):
[0063] Table No. 1:
[0064] Then, the reconstructed linguistic sentences 53 can be as follows: "If the processing method is rough turning, then the roughness should be from 160 to 80 µm"; "If the processing method is rough turning, then the quality should be from 14 to 12"; "If the processing method is rough turning, then the allowance on the side should be from 1.5 to 3.5 µm"; "If the processing method is finish turning, then the roughness should be from 160 to 80 µm"; "If the processing method is finish turning, then the quality should be from 14 to 12"; "If the processing method is finish turning, then the allowance on the side should be from 1.5 to 3.5 µm"; "If the processing method is fine turning, then the roughness should be from 160 to 80 µm"; "If the processing method is fine turning, then the quality must be from 14 to 12"; "If the processing method is fine turning, then the allowance on the side must be from 1.5 to 3.5 microns."Preferably, without limitation, for identifying elements 53, such elements 41 are identified that, according to the values 2131 of the system characteristics 213 of the semantic parts 21 included in the element 41, correspond to the mentioned requirements for the third element 53 of the fourth data structure 5. Further, in accordance with the mentioned requirements. In elements 41, the table name features, the table field (column) name features, the table row name features, and the table value features are identified. Based on the identification of all the above features, for the identification of elements 53, element 41 is first divided into parts of elements 53, from which elements 53 are formed, which represent reconstructed linguistic sentences 53 contained in element 41. Such an analysis of elements 41 for the identification and formation of elements 53 can be performed by any method known in the prior art and, accordingly, is not described in detail below. For example, without limitation, such an analysis can be performed traditionally by a linguist, or using a software algorithm of a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of unchangeable rules (an IT solution "by rules").Moreover, if there are a sufficient number of examples, it is possible to perform such an analysis using a statistical processor (neural network, AI systems) by applying neural network training technology (AI systems). Preferably, without limitation, the identification of the system characteristics 213 of the semantic parts 21 that make up the elements 41 of the third data structure 4 of the SMD and their values, if necessary, is carried out by organizing a request to the database of system features 20 (DBSF 20), formed within the framework of step 1003, consisting of the identification data of the semantic parts that make up the element 41, and obtaining the values 2131 of the system characteristics 213 of the semantic parts 21 of the information objects 11 of the digitized document 1 that make up the element 41 of the third SD SMD.In this case, as was described earlier, the system features of element 53 are at least the mentioned values 2131 of the format and functional characteristics 213 of the semantic parts 21 of the information object 11 of the digitalized document 1, of which elements 41 consist, which correspond to the requirements of the system features of elements 53.
[0065] Preferably, without limitation, the formation of the fourth data structure 5 during step 10054 is carried out by combining in one data structure the elements 51, 52, 53 of the fourth data structure 5, as well as their identification data, according to principles and methods known from the prior art, which, accordingly, are not described in detail below.
[0066] In Fig. 13, by way of example, but not limitation, a general diagram of the execution of the steps of step 1006 of forming the fifth data structure 6 is shown. Preferably, without limitation, step 1006 is characterized by: performing step 1061 of identifying elements 61 of the fifth data structure 6, in which the elements 61 of the fifth are identified data structures 6, as well as identification data of elements 61 of fifth data structure 6, representing values 611 of elements 61 of fifth data structure 6 and ordinal numbers 612 of elements 61 of fifth data structure 6 in the corresponding linguistic sentence 51, 52, 53 of fourth data structure 5, wherein such linguistic sentences 51, 52, 53 are contained in linguistic constructions 41 of third data structure 4; wherein elements 61 of fifth data structure 6 are text elements 61 of linguistic sentence 51, 52, 53 of fourth data structure 5; performing step 10062 of forming fifth data structure 6, at which fifth data structure 6 is formed from identified elements 61 of fifth data structure 6 and their identification data.
[0067] In Fig. 14, by way of example, but not limitation, the general structure of the generated fifth data structure 6 is shown. The fifth data structure 6 (fifth SD 6) contains elements 61 of the fifth SD 6, which are text elements 61 of linguistic sentences 51, 52, 53 of the fourth data structure 5, contained in linguistic constructions 41 of the third data structure 4 SMD and identification data of text elements 61 of linguistic sentences 51, 52, 53 of the fourth data structure 5, contained in linguistic constructions 41 of the third data structure 4 SMD, which are by way of example, but not limitation: values 611 of text elements 61 in linguistic sentences 51, 52, 53 of the fourth data structure 5;ordinal numbers 612 of text elements 61 in linguistic sentences 51, 52, 53 of the fourth data structure 5. Preferably, without being limited to, text elements 61 of linguistic sentences 51, 52, 53 of the fourth data structure 5 as elements 61 of the fifth data structure 6 represent various isolated objects of the corresponding linguistic sentence 51, 52, 53, by way of example, but not limitation: text elements of the first type (primary text elements), which are, by way of example, but not limitation, words, numbers, digits, indices (constructions from digits and / or letters and / or signs), punctuation marks and the like, wherein the said objects of the linguistic sentence are isolated in the sentence by means of spaces, with the exception of punctuation marks, which do not have a space on at least one side (before the punctuation mark);Text elements of the second type (complex text elements), which are, by way of example, but not limitation, word forms or groups of words that are a single object according to the rules of morphology (complex word form or morphohomonym). For example, the words "in," "in accordance," and "with" represent several primary TEs, although at the same time they (the group of words) constitute a single complex text element "in accordance with." On; In practice, the user independently sets criteria for text elements in advance, specifying the type of text elements of a linguistic sentence that interests him. Preferably, but not limited to, the value 611 of the text element 61 is the set of all characters (letters, numbers, symbols, punctuation marks, spaces) that make up the element 61 in the linguistic sentence 51, 52, 53. Preferably, but not limited to, the ordinal number 612 of the text element 61 is the ordinal number 612 of the text element 61 in the linguistic sentence 51, 52, 53 of the fourth data structure 5. In the data structure, the elements 61, by way of example, but not limitation, may be referred to as "TE1.x", "TE2.x", "TE3.x" "TEp.x", where n>1 is the ordinal number of the text element 61 in the linguistic sentence 51, 52, 53 of the fourth data structure 5, and x>1 is the corresponding ordinal number 512, 522, 532 linguistic sentences 51, 52, 53 in the fourth data structure.
[0068] Preferably, but not limited to, the identification and formation of text elements 61 of the fifth data structure 6 during step 10061 is carried out by analyzing the text and identifying (highlighting) individual text elements 61 according to their type and description, which must be known in advance. For example, but not limited to, such an analysis can be carried out by highlighting words, numbers, or indices in a sentence, separated from each other by a space, as well as punctuation marks that are attached to said words, numbers, and indices. In this case, it is preferable that the last punctuation mark in the sentence is not taken into account and is not considered as a text element 61 of the linguistic sentence 51 of the fourth data structure 5.Preferably, without limitation, when identifying text elements 61 that are complex text elements 61 in the event that such a type of text elements 61 has been previously established, a request is made to separate databases (for example, without limitation, to a pluggable electronic morphological dictionary) to confirm the composition of the complex text element 61 for the purpose of its further identification as a text element 61 in the linguistic sentence 51 of the fourth data structure 5. Such an analysis of the linguistic sentences 51 of the fourth data structure 5 for the identification and formation of text elements 61 can be performed by any method known from the prior art and, accordingly, is not described in detail further.For example, but not limited to, such a complex analysis can be performed traditionally by a linguist, or using a software algorithm from a linguistic (syntactic) processor, or using traditional programming by solving problems using coding of immutable rules (a "rule-based" IT solution). Moreover, this is possible if a sufficient number of examples are available. It is possible to perform such an analysis using a statistical processor (neural network, AI systems) by employing neural network (AI system) training technology. Preferably, without limitation, the formation of the fifth data structure during step 10062 is performed by combining the elements 61 of the fifth data structure 6, as well as their identification data, into a single data structure according to principles and methods known in the prior art, which, accordingly, are not further described in detail.
[0069] In fig. 15, as an example, but not limitation, a general diagram of the execution of the stages of stage 1007 of the formation of a database of linguo-logical-subject features 30 (BDLL1P130), which is a database of linguistic, logical and subject features of text elements 61 of linguistic sentences 51, 52, 53 contained in linguistic constructions 41 of the third data structure 4, is shown. Preferably, without limitation, stage 1007 is characterized by: execution of stage 10071 of the formation of the first part of the linguo-logical-subject features of text elements 61 of linguistic sentences 51, 52, 53 of the fourth data structure 5, in which, for the linguistic analysis of text elements 61 contained in the fifth data structure 6 and classified as a word, identification data of such text elements 61 are provided and linguistic characteristics of such text elements 61 are obtained, as well as values the mentioned linguistic characteristics;by performing step 10072 of forming the second part of the linguistic-logical-subject features of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5, in which, for the logical analysis of the text elements 61 contained in the fifth data structure 6 and classified as a word, the identification data of such text elements 61 are provided, as well as the linguistic characteristics of such text elements 61 together with the mentioned values of the linguistic characteristics, and the logical characteristics of such text elements 61 of each linguistic sentence 51, 52, 53 are obtained, as well as the values of the mentioned logical characteristics;by performing step 10073 of forming the third part of the linguistic-logical-subject features of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5, in which, for the subject analysis of the text elements 61 contained in the fifth data structure 6 and classified as a word, the identification data of such text elements 61 are provided, as well as the linguistic characteristics of such text elements 61 together with the mentioned values of the linguistic characteristics, and the logical characteristics of such text elements 61 are provided, as well as the values of the mentioned logical characteristics, and the characteristics of such text elements 61 in the subject area are obtained, which are; subject characteristics, as well as the values of the said subject characteristics; performing step 10074 of forming a database of linguo-logical-subject features 20, at which a database of linguo-logical-subject features 30 of text elements 61 of linguistic sentence 51, 52, 53 of fourth data structure 5 is formed, wherein the linguo-logical-subject features of text elements 61 of linguistic sentence 51, 52, 53 of fourth data structure 5 are the said linguistic characteristics, the said logical characteristics and the said subject characteristics obtained for each text element 61 during steps 10071, 10072 and 10073, which have respectively the said values of linguistic characteristics, the said values of logical characteristics and the said values of subject characteristics.
[0070] In Fig. 16, by way of example, but not limitation, the general structure of the formed database of linguo-logical-subject features 30 (DB DLi 11111 30) is shown, which is a database of linguo-logical-subject features 30 of text elements 61 of linguistic sentence 51 of the fourth data structure 5. Preferably, without being limited, the practical purpose of DB DLLP 30 is the formation of two interrelated groups of characteristics: linguo-logical and logical-subject characteristics. Preferably, without being limited, the first group (linguistic characteristics) is necessary for the extraction from the linguistic sentence 51, 52, 53 of logical constructions (final judgments, simple judgments, elements of simple judgments (concepts, attributes of concepts, images)) for the transformation of the linguistic sentence into logical constructions, which represent, by way of example, but not limitation, simple judgments, final judgments.The second group (logical-subject characteristics) is necessary for the formation of subject-oriented structured information by correlating the logical objects of the aforementioned logical constructs with the subject objects of the subject constructs—objects and constructs accepted in a specific subject area and established in a formalized model of the structural elements of the subject area. By way of example, but not limitation, such a subject area could be legal science (law). In this case, the correlated subject construct would be a formalized model of the structural part of a legal norm (hypotheses, dispositions, or sanctions), and the subject objects would be the elements of the formalized model of the structural part of a legal norm (elements of the FSCHPN), which could include, by way of example, but not limitation, the subject of the legal relationship, the object of the legal relationship, the content of the legal relationship, and the like.
[0071] Preferably, without being limited to, the first part of the linguo-logical-subject features 613 of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 contains linguistic characteristics - morphological, syntactic and semantic characteristics. In this case, without being limited to, the set of values of all linguistic characteristics of the text element is for each text element 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 its distinctive (unique) linguistic feature in the linguistic sentence 51, 52, 53 of the fourth data structure 5. The morphological characteristics preferably indicate the morphological features of the text elements 61 of the linguistic sentence 51 of the fourth data structure, which can be classified, as an example, but not limitation, by the level of nesting (genus-species-subspecies).In this case, the morphological genera of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 are preferably a word, a digit, punctuation marks, and other signs; the morphological types are a part of speech (for words), a type of digit (Arabic, Roman), a type of punctuation mark (period, comma, and the like), a type of other sign; the morphological subtypes are gender, number, case of parts of speech, and the like for words, as well as a number, binary code, index, and the like for digits.The syntactic characteristics preferably indicate a set of syntactic features of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, among which the following syntactic features of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 can be distinguished, by way of example, but not limitation: syntactic role (predicate, subject, and the like); syntactic parent (syntactically main word); syntactic descendants (syntactically subordinate words); syntactic coordinative relationship (the presence of another text element having the same syntactic role and the same syntactic parent).The semantic characteristics preferably indicate the semantic features of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, among which the following semantic characteristics of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 can be distinguished, by way of example, but not limitation: a semantic group (a group of words that can be attributed to one class, genus, species or subspecies of objects or actions of the surrounding world when the features of the mentioned classes, genera, species or subspecies coincide), semantic status (the semantic meaning of a word or group of words within a phrase that refers to a certain conceivable image (object or action) - for example, but not limited to, conceivable. The image of “absence of a seller at the location of a consumer” consists of two top-level elements (terms): the first is “absence of a seller”, and the second is “location of a consumer”, which have the following semantic statuses: the first is the main one (defines the meaning of the term), the second is additional (clarifies the previously defined meaning of the main term)).
[0072] Preferably, without limitation, the second part of the linguistic-logical-subject characteristics 614 of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 contains logical characteristics. In this case, the set of values of all logical characteristics of the text element is for each text element 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 its distinctive (unique) logical feature in the linguistic sentence 51, 52, 53 of the fourth data structure.The logical characteristics preferably indicate the logical attributes of the text elements of the text elements of the linguistic sentence 51, 52, 53 of the fourth data structure 5, among which the following logical characteristics of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 can be distinguished, by way of example, but not limitation: the logical roles of each word, which is a text element 61 in the linguistic sentence 51, 52, 53 of the fourth data structure 5. The logical role of a word is understood to be the logical position of the word in the logical entities (logical objects) of the sentence, among which the following logical entities (logical objects) can be distinguished, by way of example, but not limitation: concept, attribute of a concept, term (part of an image), image (element of a simple judgment), simple judgment, complex judgment.Identifying the logical role of a word in the simplest logical objects (concept and concept attribute) is independent of the type of logical judgment construct and represents a label (index) that indicates what a given word is in these simple logical objects. For example, the word "law" is always a logical object "concept," and the word "federal" is a logical object "concept attribute." Identifying the logical role of a word in more complex (composite) logical objects (for example, a term and an image) depends on the formalized model of the logical judgment construct (FMLC), whose elements (the logical objects of the formalized model of the logical judgment construct) determine the logical role of the word.
[0073] Preferably, but not limited to, a necessary condition for the formation of linguistic and logical features is the presence of correlated objects (objects acceptable for comparisons). The question of the admissibility of comparing the text element 61 of sentence 51, 52, 53 with logical objects can be resolved on the basis of an analysis of the comparability of the logical objects of the FMLKS with the text elements 61 of sentence 51, 52, 53.This analysis shows that while text elements can be such logical objects as a “concept attribute” (for example, the word “established”), individual text elements of a sentence are in no way related to logical objects such as a “concept” expressed through a group of primary text elements 61 (for example, “violation of consumer rights”). And since an element (logical object) of a FMLKS can be not only a concept with a feature (for example, “established violation of consumer rights” - a group of four primary text elements), but also a significantly larger construction of syntactically related words (for example, the logical object “subject” can have the following linguistic form: “obligations to execute, within the limits of their authority, decisions of a judge to conduct operational-search activities”), it can be concluded that it is inadmissible to compare text elements of a sentence with logical objects.Preferably, but not limited to, the linguistic object of a sentence that can be compared with the logical object of the FMLCS may be a syntactic unit (SU). A syntactic unit is a word form or phrase (a syntactically related group of words). The flexible format of a syntactic unit allows it to correspond to any logical objects that make up the FMLCS. Thus, the identified relevant linguistic and logical characteristics are important from a practical point of view not so much for the description of individual text elements 61, but for relevant syntactic units (relevant SU), which consist of one or more text elements 61 of a sentence 51, 52, 53. Relevant SU are a current list of syntactic units correlated with the relevant logical objects of the relevant FMLCS.The current CE and the current FMLKS are installed in advance and recorded in the first user database (the first PBD), which is thus a database of current syntactic units (current CE), current logical objects (current LogO) and the current FMLKS, including a correlation table of current CE and current LogO.
[0074] Preferably, without limitation, the third part of the linguo-logical-subject characteristics 615 of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 contains subject characteristics. In this case, the set of values of all subject characteristics of the text element is for each text element 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 by its distinctive (unique) subject feature in the linguistic sentence 51, 52, 53. The subject characteristics preferably indicate the subject features of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, among which the following subject characteristics of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 can be distinguished, by way of example, but not limitation, in the subject area of law: the legal roles of each word, which is a text element 61 (correlated with a separate current CE or part of a current CE), in the linguistic sentence 51, 52, 53 of the fourth data structure 5.The legal role of a word is understood as the legal position of a word in the legal entities (legal objects) of a sentence, among which the following legal entities (legal objects) can be distinguished, by way of example but not limitation: legal concept, attribute of a legal concept, legal term, subject of legal relations, object of legal relations, element of the content of legal relations, hypothesis, disposition, sanction (structural part of a legal norm), and legal norm. Identifying the legal role of a word in the simplest legal objects (legal concept and attribute of a legal concept) does not depend on the type of legal construction and represents a label (index) indicating what a given word is in the aforementioned simple legal objects. For example, the word "law" is always a legal object "object of legal relations," and the word "federal" is an attribute of a legal object "attribute of an object of legal relations." The identification of the legal role of a word in more complex (composite) legal objects (for example, in a hypothesis, disposition or sanction) depends on the formalized model of the structural part of a legal norm, in relation to the elements of which (to the legal objects of the formalized model of the logical construction of a judgment) the legal role of a word will be established.
[0075] Preferably, but not limited to, a necessary condition for the formation of logical-subject features is the presence of correlated objects (objects permissible for comparison) in the logical and subject data arrays. If, by way of example, but not limitation, the subject area is understood to be the subject area of law, then the question of the permissibility of comparing the compared objects can be considered not speculatively, but practically. The practical question of the feasibility of comparing the logical objects of a sentence with legal objects can be resolved based on the results of scientific research in this subject area. It is reliably known from open sources that legal scholars rightly believe that, from a logical standpoint, any legal norm is a judgment. And the legal and logical constructions of a legal norm are reflected in a specific linguistic form. A construction in which linguistics plays the role of a legal-technical tool. In the text of a normative act, any judgment is expressed in the form of a declarative sentence. Moreover, the logical structure of the judgment corresponds, without limitation, to the grammatical structure of a complex sentence. In its meaning, the grammatical structure of the sentence coincides with the legal structure of the legal norm. Currently, the theory of legal norms fully allows us to consider any legal norm a judgment and confirms the existence of interrelations and the unity of legal norms as a unity of legal, logical, and grammatical constructions. Thus, the interrelations of legal and logical constructions provide the basis for the existence of interrelations (correlations) between the elements of legal and logical constructions. In other words, the logical objects of logical constructions and the legal objects of legal constructions are objects permissible for comparison.The result of such a comparison is a correlation (relationship) between specific legal objects and specific logical objects. Such correlation becomes practically possible with the presence of formalized models (legal and logical) containing the correlated objects. The priority formalized model is the legal (subject) formalized model (the formalized model of the structural part of a legal norm), since the number and type of objects (elements) of the legal (subject) formalized model determines the level of detail and depth of structuring of the logical formalized model (the formalized model of the logical structure of a judgment).The current formalized model of the basic design of the subject area (FMBKPO) is installed in advance and recorded in the second user database (second UDB), which is thus a database of current basic subject objects (current BPO) and FMBKPO, including a correlation table of current BPO and current logical objects (current LogO).
[0076] Preferably, without being limited to, the formation of the first part of the linguistic-subject characteristics - linguistic characteristics 613 and their values 6131 - for the text elements 61 of each linguistic sentence 51, 52, 53 of the fourth data structure 5 is preferably carried out at step 10071 by first complex analysis of each text element 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, which is an analysis of the said text elements 61, as an example, but not limitation, for example, an analysis based on the location of the text element in the structure of the sentence, its meaning, type, classification of its conceivable image and analysis of its connections with other text elements in the sentence, as well as information from the first PBD on the current CE. By According to the results of the first complex analysis, the linguistic characteristics 613 are preferably formed and entered at step 10074 into the BDLL1111 30 in the form of a list of linguistic characteristics 613 with the values of these characteristics 6131. For example, but not limited to, one of the linguistic characteristics 613 may be the "syntactic role" of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, with the value of this linguistic characteristic, for example, "subject", and may also be the "syntactic role" of the actual CE, consisting of both one text element 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, and of a group of the mentioned text elements 61, with the value of this linguistic characteristic, for example, "adverbial modifier of place". Such an analysis may be performed by any method known from the prior art and, accordingly, is not described in detail below.For example, but not limited to, such a complex analysis can be performed traditionally by a linguist, or using a software algorithm in a linguistic (syntactic) processor, or using traditional programming by solving problems using the coding of immutable rules (a "rule-based" IT solution). Furthermore, given a sufficient number of examples, such an analysis can be performed using a statistical processor (neural network, AI system) through the use of neural network (AI system) training technology.
[0077] Preferably, without being limited to, the formation of the second part of the linguistic-subject characteristics - logical characteristics 614 and their values 6141 - for the text elements 61 of each linguistic sentence 51, 52, 53 of the fourth data structure 5 is preferably carried out at step 10072 by a second complex analysis of each text element 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, as well as each current CE, representing as an example, but not limitation, an analysis of the said text elements 61 (or several of the said text elements 61 that make up the current CE) based on the location of the said text element 61 (group of the said text elements 61) in the structure of the sentence, its meaning, type, classification of its conceivable image and analysis of its connections with other text elements in the sentence, as well as an analysis of the identified linguistic characteristics 613 and their values 6131,as well as information from the first PBD on the current CEs, correlated with the current LogOs established in the FMLKS. Without limitation, based on the results of the second comprehensive analysis, the logical characteristics of 614 text elements of 61 linguistic sentences 51, 52, 53 of the fourth structure are preferably formed, data 5 and their input at step 10074 into BDLL1P130 in the form of a list of logical characteristics 614 with the values of these characteristics 6141. For example, but not limited to, one of the logical characteristics 614 of the said text element 61 may be the "logical role" of the said text element 61 with the value of this logical characteristic of "concept feature", or one of the logical characteristics 614 of a group of the said text elements 61 (actual CE) may be the "logical role" of a group of the said text elements 61 (actual CE), with the value of this logical characteristic of "subject of judgment". Such an analysis may be performed by any method known from the prior art and, accordingly, is not described in detail below.For example, but not limited to, such a complex analysis can be performed traditionally by a linguist, or using a software algorithm in a linguistic (syntactic) processor, or using traditional programming by solving problems using the coding of immutable rules (a "rule-based" IT solution). Furthermore, given a sufficient number of examples, such an analysis can be performed using a statistical processor (neural network, AI system) through the use of neural network (AI system) training technology.
[0078] Preferably, without being limited to, the formation of the third part of the linguistic-subject characteristics - subject characteristics 615 and their values 6151 - for the text elements 61 of each linguistic sentence 51, 52, 53 of the fourth data structure 5 is preferably carried out at step 10073 by means of a third complex analysis of each text element 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, as well as each current Logo, the analysis being, by way of example, but not limitation, an analysis of the said text elements 61 (or several of the said text elements 61) that make up the current Logo based on the logical characteristics 614 and their values 6141, its connections with other logical objects in the sentence, as well as information from the first PBD on the current Logo, correlated with the subject objects established in the FMBKPO, including a correlation table of the current BIO and the current Logo.Preferably, without being limited to, based on the results of the third complex analysis, the formation of subject characteristics 615 of text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 is preferably performed and they are entered at step 10074 into BDLL1P1 30 in the form of a list of subject characteristics 615 with the values of these characteristics 6151. For example, but not limited to, one of the subject characteristics 615 of the said text element 61 in the subject area of law (jurisprudence) may be the “legal role” of the said text element 61 co. The meaning of this subject characteristic 6151 is "subject of legal relations." Such an analysis can be performed by any method known in the art and, accordingly, is not described in detail below. For example, but not limited to, such a complex analysis can be performed traditionally by a linguist, or using a software algorithm of a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of immutable rules (a "rule-based" IT solution). Furthermore, given a sufficient number of examples, such an analysis can be performed using a statistical processor (neural network, AI system) through the use of neural network (AI system) training technology.
[0079] Preferably, without limitation, based on the identified first part of the linguo-logical-subject characteristics 613 of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5 and their values 6131, the second part of the linguo-logical-subject characteristics 614 of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5 and their values 6141, as well as the third part of the linguo-logical-subject characteristics 614 of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5 and their values 6141, a database of linguo-logical-subject features 20 is ultimately formed, which is a B DLi 1111130 of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5.In this case, the first part of the linguo-logical-subject characteristics 613 of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5 and their values 6131 form unique linguistic features of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5, the second part of the linguo-logical-subject characteristics 614 of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5 and their values 6141 form unique logical features of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5, and the third part of the linguo-logical characteristics 615 of the text elements 61 of the linguistic sentences 51, 52, 53 of the fourth data structure 5 and their values 6151 form unique subject features of the text elements 61 linguistic sentences 51, 52, 53 of the fourth data structure 5.
[0080] Fig. 17, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1008 of forming the sixth data structure 7. Preferably, without limitation, step 1008 is characterized by: executing step 1081 of forming the elements 71 of the sixth data structure 7, in which, based on the information, contained in the database of linguistic-logical-subject features 20, the fifth data structure 6, and also on the basis of data from the first user database containing data on current syntactic units and current logical objects, as well as on the current formalized model of the logical construction of a judgment, identify and form elements 71 of the sixth data structure 7, which are components of a simple judgment 71 of the corresponding linguistic sentences 51, 52, 53 of the fourth data structure 5, as well as identification data of the said components of a simple judgment 71, representing for each said component of a simple judgment 71 type 71.1, 71.x of such corresponding component of simple judgment 71, value 711 of such corresponding component of simple judgment 71 and ordinal number 712 of such corresponding component of simple judgment 71 in the corresponding linguistic sentence 51, 52, 53 of the fourth data structure 5; by performing step 10082 of forming the sixth data structure 7, in which the sixth data structure 7 is formed from the mentioned components of simple judgment 71 and their identification data. Moreover, the components of simple judgments 71 are components of the corresponding simple judgments, and the simple judgments are simple judgments of the corresponding linguistic sentences 51, 52, 53 of the fourth data structure 5.
[0081] In fig. 18, as an example, but not limitation, the general structure of the generated sixth data structure 7 is shown. Preferably, without limitation, the sixth data structure 7 (sixth SD 7) contains elements 71 of the sixth data structure, which are components of a simple judgment 71 of each simple judgment of each linguistic sentence 51, 52, 53 of the fourth data structure 5 and identification data of said components of a simple judgment 71, which are at least the type 71.1, 71.X of said components of a simple judgment 71, values 711 of said components of a simple judgment 71, ordinal numbers 712 of said components of a simple judgment 71 of the corresponding linguistic sentence 51, 52, 53 of the fourth data structure 5, constituting such components of a simple judgment 71 of a linguistic sentence 51, 52, 53 fourth data structure 5.Preferably, but not limited to, the components of simple proposition 71 of linguistic sentence 51, 52, 53 of the fourth data structure 5 are elements of a simple proposition. Preferably, according to the theory of logic, simple propositions are a set of structural elements of a simple proposition, by means of which something about the subject of the proposition is affirmed or refuted. In this case, preferably, the basic structural elements of a simple proposition are: are the subject and the predicate. According to logical theory, the subject of a simple judgment is preferably recognized as the concept expressing the object of the judgment, that is, what is stated in the given simple judgment. According to logical theory, the predicate of a simple judgment is preferably recognized as what is affirmed or denied about the object of the judgment (the subject of the simple judgment). For example, without limitation, the simple sentence "The authorized seller is obliged to transfer the goods to the buyer after payment" contains the structural element of the simple judgment "subject," meaning "authorized seller," as well as the structural element of the simple judgment "predicate," meaning "is obliged to transfer the goods to the buyer after payment." Moreover, the structural elements of simple judgments preferably consist of concepts that can have the characteristics of concepts.For example, without limitation, the element of the simple judgment “subject”, which has the meaning “authorized seller”, consists of the concept “seller” and the attribute of this concept “authorized”.Preferably, without limitation, for practical purposes related to the present transformation of a structured data array, which is a digitalized document 1, the selection in linguistic sentences 51, 52, 53 of simple judgments that have only two structural elements ("subject" and "predicate") is of no interest, since the main task of the logical formalization of linguistic sentences 51, 52, 53 of the fourth data structure 5 is the formation of such formalized objects - components of a simple judgment 71 - which in themselves (based primarily on the type of component of a simple judgment 71) explain the logical role of such formalized objects (concepts, concepts with a feature, or a group of concepts and features of concepts that represent a logical image of an element of a simple judgment) and their semantic function in a simple judgment, based on the practical problems that must be solved by formalizing objects with relevant semantic functions.For example, but not limited to, in the subject area of law, it is necessary to formalize the legal norms contained in the proposals of regulatory acts. To this end, preferably, but not limited to, specialists in the subject area of law create a formalized model of the structural part of a legal norm (FMSCHPN), which represents a formalized model of the hypothesis, disposition, or sanction of a legal norm. Such a FMSCHPN contains elements that possess unique semantic functions within the subject area of law, such as, but not limited to, the following FMSCHPN elements: legal rules and legal facts (events and circumstances influencing the practical application of a legal rule). In this case, preferably, a legal rule consists of subelements (nested elements), representing the subject of legal relations. The object of legal relations and the content of legal relations. Moreover, the sub-element "content of legal relations" preferably also consists of sub-elements (nested elements), namely, the regulation method, influencing objects, and definition (a defining expression that reveals the meaning of the defined concept or term or establishes the meaning of the concept or term, which in practice means another name for the subject or object of legal relations, or a definition of the subject or object of legal relations). Thus, preferably, based on the practical problem and the aforementioned formalized model formed to solve it, the number and types of elements of the aforementioned formalized model (elements that do not contain sub-elements and sub-sub-elements) are determined. Moreover, preferably, each element of such a formalized model possesses unique semantic functions of the subject area, representing the relevant elements of such a formalized model.Preferably, such a current formalized model serves as a guideline for the formation of a current formalized model of a simple judgment (a formalized model of the logical construction of a judgment), and in such a simple judgment, the structural element "predicate" should be divided into a number of subelements that is at least equal to the number of current elements of the formalized model of the subject area minus one - since among the current elements of the formalized model of the subject area, there must be one current element that corresponds to (correlates with) the structural element "subject." Moreover, such a current formalized model of a simple judgment (a formalized model of the logical construction of a judgment) preferably contains current elements of the simple judgment with current semantic functions.
[0082] As an example, but not a limitation, Table 2 demonstrates the logical roles and semantic functions of the structural elements of simple judgments (logical objects) of the sentence “The authorized seller is obliged to transfer the goods to the buyer after payment” without taking into account the solution of a practical problem and the current formalized model of the subject area.
[0083] Table No. 2:
[0084] Thus, the example provided preferably demonstrates that a simple proposition has a minimum number of elements equal to two (subject, predicate). Moreover, the maximum number of elements of a simple proposition preferably depends solely on the current formalized model of the simple proposition. As an example, but not limitation, Table 3 demonstrates the logical roles and semantic functions of the current elements of simple propositions (logical objects) of the sentence "The authorized seller is obligated to deliver the goods to the buyer after payment" based on the solution to a practical problem and taking into account the current formalized model of the logical structure of the proposition and the current formalized model of the subject area.
[0085] Table No. 3:
[0086] Thus, preferably, the components of the simple judgment 71 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 are all elements of the simple judgment established in the current formalized model of the simple judgment (the formalized model of the logical construction of the judgment), and such current formalized model of the simple judgment is contained in the first user database (the first PDB). Preferably, without limitation, information about which text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 constitute individual components of the simple judgment 71 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 is contained in the database of linguistic-logical-subject features 30 (B DLILI 30), formed during step 1007. Moreover, preferably, the type of the mentioned components of the simple judgment 71 and the method of correlation (relationship) of the actual logical objects are also contained in the BDLLPP 20.
[0087] Preferably, without limitation, the components of simple judgment 71 (CPS 71) may be of at least two types - the first CPS 71.1 and the second CPS 71.x. Preferably, without limitation, the number of second CPS 71.x will correspond to the number of elements of the simple judgment in the model of the simple judgment (the formalized model of the logical construction of the judgment). In this case, without limitation, the first CPS 71.1 (element "1" in the current formalized model of the logical construction of the judgment) will always be the subject. When identifying the second CPS 71.x, each identified CPS 71.x will receive an index value "x" that corresponds to the ordinal number of the element of the simple judgment, represented in the current formalized model of the logical construction of the judgment. Preferably, without being limited to, the values 711 of the components of the simple judgment 71 of the linguistic sentences 51, 52, 53 of the fourth data structure are the values 711 of the components of the simple judgment 71 of all types (the first CPS 71.1 and the second CPS 71.x), which make up the component of the simple judgment 71. In this case, preferably, the value 711 of the components of the simple judgment 71 means the values 711.1 and 711.x of all the components of the simple judgment 71.1 and 71.x identified in the linguistic sentence 51, 52, 53.Preferably, without limitation, the ordinal number 712 of the component of the simple proposition 71 are the ordinal numbers 712 of the components of the simple proposition 71 in the linguistic sentence. 51, 52, 53 of the fourth data structure 5. In the sixth data structure 7, the elements 71, by way of example, but not limitation, may be referred to as "KPS1", "KPS2", "KPS3" "KPSp", where n> 1 is the ordinal number of the component of the simple proposition 71 in the linguistic sentence 51, 52, 53 of the fourth data structure 5. In this case, without limitation, the first in order component of the simple proposition 71 in the linguistic sentence 51, 52, 53 receives the ordinal number "1", the next one receives the ordinal number "2", and so on until the last component of the simple proposition in the linguistic sentence 51, 52, 53, which receives the last ordinal number. Moreover, preferably, the sequence of components of the simple proposition 71 in the linguistic sentence 51, 52, 53 is determined by the ordinal number of the first text element 61, from which the linguistic construction 41 is formed, which is the source of data for the formation of the linguistic sentence 51, 52, 53. In other words, preferably, the sequence of components, based on the ordinal number of its first text element, may not be of the form "1-2-3-4" and so on, but, for example, without limitation, of the form "3- 7-11-12-14-20" and so on, that is, based on the numbers of the first text elements 61 of the components of the simple judgment 71 in the linguistic sentence 51, 52, 53 of the fourth data structure 5. In this case, preferably, the values 712 of the components of the simple judgment 71 mean the values 712.1 and 712.x of all components of the simple judgment 71.1 and 71.x identified in the linguistic sentence 51, 52, 53 of the fourth data structure 5. Preferably, without limitation, different types 71.1, 71.x components of a simple judgment 71 linguistic sentence 51, 52, 53 of the fourth data structure 5 are identified and formed on the basis of data from the database. linguo-logical-subject features 30 (LLSSF 30), containing both information on the types of elements of a simple judgment and information on the content of individual elements of a simple judgment (that is, which text elements 61 each component of a simple judgment 71 in a linguistic sentence 51, 52, 53 of the fourth data structure 5 consists of). Preferably, without limitation, the identification and formation of components of a simple judgment 71 of a linguistic sentence 51, 52, 53 of the fourth data structure 5 during step 10081 is carried out by means of a comprehensive linguo-logical analysis of elements 61 of the fifth data structure 6 - the aforementioned text elements 61 and their identification data. Such a comprehensive analysis of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 is carried out using information about the text elements 61 from the generated BDLLPP 20, as well as on the basis of the data of the current formalized model of the logical construction of judgment (FMLCS).In this case, the current formalized model of the logical construction of a judgment preferably contains at least two types of components—the first CPS 71.1 and the second CPS 71.x. Thus, a formalized model of the logical construction of a judgment is preferably a system for describing a simple judgment that has at least two of the aforementioned components. The goal of this comprehensive analysis is to identify all components of a simple judgment in a linguistic sentence, as defined by the formalized model of the logical construction of a judgment. Preferably, but not limited to, the identification and formation of the components of the simple judgment 71 of the sixth data structure 7 during step 10081 is performed iteratively. The number of steps of step 10081 depends on the currently used formalized model of the logical construction of the judgment. Preferably, but not limited to, said model contains a fixed number of types of elements of the simple judgment, and the number of steps of step 10081 is determined in accordance with this number of types, since within the framework of one step it is possible to identify one type of element of the simple judgment and to form only one type of component of the simple judgment 71. Moreover, without limitation, since the formalized model of the logical construction of the judgment can minimally contain no less than two components, the minimum number of steps will also be equal to two.As an example, but not limitation, Table No. 4 shows an example of identifying and forming the components of a simple judgment 71 in the linguistic sentence “The authorized seller is obliged to transfer the goods to the buyer after payment” in accordance with the eight-element actual formalized model of the logical construction of the judgment indicated in Table No. 4.
[0088] Table No. 4: 0089] Preferably, without limitation, the naming of the actual elements of the formalized model of a simple judgment, the identification of the actual elements of the formalized model of a simple judgment in a linguistic sentence, the identification of the type of the CPS 71, the identification of the meaning of the CPS 71, if necessary, are carried out by organizing a request to the database of linguo-logical-subject features 30 (BDLL1P130) of text elements 61 of linguistic sentences 51, 52, 53 of the fourth data structure 5, formed within the framework of step 1007, consisting of the identification data of the text elements 61 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, and obtaining a list of elements of the actual formalized model of a simple judgment, as well as information on each text element 61 and on which elements of a simple judgment the said text elements 61 are included.Preferably, but not limited to, such identification and generation of elements of the CPS 71 can be accomplished by any method known in the art and, accordingly, is not described in detail further. For example, but not limited to, such a comprehensive analysis can be performed traditionally by a linguist, or with. Using a software algorithm from a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of immutable rules (a "rule-based" IT solution). Furthermore, given a sufficient number of examples, such analysis can be performed using a statistical processor (neural network, AI system) by applying neural network (AI system) training technology.
[0090] Preferably, without limitation, the formation of the sixth data structure 7 of the SMD during step 10082 is carried out by combining in one data structure the elements 71 (KIS 71) of the sixth data structure 7, as well as their identification data, according to principles and methods known from the prior art, which, accordingly, are not described in detail below.
[0091] Fig. 19, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1009 for forming the seventh data structure 8.Preferably, without limitation, step 1009 is characterized by: performing step 10091 of forming elements 81 of the seventh data structure 8, at which, on the basis of information contained in the database of linguistic-logical-subject features 30 and the sixth data structure 7, from the formed components of simple judgments 71 in accordance with the current formalized model of the logical construction of judgment, elements 81 of the seventh data structure 8 are formed, which are simple judgments 81, as well as identification data of simple judgments 81, which represent values 811 of simple judgments and ordinal numbers 812 of simple judgments in linguistic sentences 51, 52, 53 of the fourth data structure 5; performing step 10092 of forming the seventh data structure 8, at which the seventh data structure 8 is formed from the formed simple judgments 81 and their identification data.
[0092] In Fig. 20, by way of example, but not limitation, the general structure of the generated seventh data structure 8 is shown. Preferably, without limitation, the seventh data structure 8 (seventh SD 8) contains elements 81 of the seventh data structure 8, which are simple judgments 81 of each linguistic sentence 51, 52, 53 of the fourth data structure 5 and identification data of said simple judgments 81, which are, by way of example, but not limitation, values 811 of said simple judgments 81, ordinal numbers 812 of simple judgments 81 in the corresponding linguistic sentence 51, 52, 53 of the fourth data structure 5.
[0093] Preferably, but not limited to, the simple propositions 81 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 represent a set of elements of a simple proposition. Preferably, such elements of a simple proposition 81, according to the theory of logic, represent the structural elements of a simple proposition—the subject and the predicate. In this case, the subject is preferably the object about which something is asserted or refuted in the simple proposition, and the predicate is that which is stated in this simple proposition. Moreover, from a practical point of view, the element of a simple proposition 81 is preferably the subject and a certain number of subelements (nested elements) of the predicate, wherein such a breakdown of the predicate into subelements is justified solely for practical purposes within the framework of solving the current problem.The solution of actual problems determines actual models of the logical construction of judgment, which contain actual elements of simple judgment 81, on the basis of which, at step 1008, the components of simple judgment 71 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 are identified and formed. Preferably, without limitation, simple judgment 81 represents the primary logical construction of thinking with the help of which the idea is formed and conveyed that something (the predicate of the judgment) is affirmed or refuted about the subject of the judgment (the subject of the judgment). Preferably, from a linguistic point of view, simple judgments 81 are simple sentences, or simplified simple sentences.In this case, various variants of simple sentences that can be considered simple judgments 81 are preferably possible, for example, but not limited to this example, the following types of simple sentences can be given: simple sentences in their original, untransformed form; as well as simple sentences in a transformed (simplified) form, for example, but not limited to: without participial or adverbial participial phrases, which themselves form simplified simple sentences that are influencing simple judgments; without homogeneities (non-homogeneous, without series of homogeneous members, which leads to the formation of homogeneous simplified simple sentences); without insertions (without text in brackets); without conditional names (without text in quotation marks); without circumstances (conditions); and the like, including combinations of the above and unspecified types.Preferably, without being limited to, simple judgments 81, from the point of view of the presence of syntactic connections between the words of a sentence containing the said simple judgments 81, can be both main simple judgments and influencing simple judgments.
[0094] Preferably, without limitation, the simple propositions 81 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 have the identification data: the value 811 of the simple proposition 81 and the ordinal number 812 of the simple proposition 81. Preferably, but not limited to, the value 811 of the simple proposition 81 is the set of values of the text elements 61 of all components of the simple proposition 71 that make up the simple proposition 81 of the linguistic sentence 51, 52, 53 of the fourth data structure. Preferably, but not limited to, the ordinal number 812 of the simple proposition 81 is the ordinal number of the simple proposition 81 in the linguistic sentence 51 of the fourth data structure. In the data structure, the elements 81, by way of example, but not limitation, may be referred to as "PS1", "PS2", "PS3", "PSp", where n>1 is the ordinal number of the simple proposition 81 in the linguistic sentence 51, 52, 53 of the fourth data structure 5.Preferably, without being limited to, the formation of simple judgments 81 of the linguistic sentence 51 of the fourth data structure 5 during step 10091 is carried out on the basis of the components of the simple judgment 71 formed in step 1008 in the linguistic sentence 51, 52, 53 of the fourth data structure 5 by combining the said components of the simple judgment 71 in accordance with the current formalized model of the logical construction of the judgment and taking into account information from the database of linguistic-logical-subject features 30 of text elements 61 on the presence of syntactic links between text elements 61 included in the various components of the simple judgment 71. A non-limiting example of the formation of the simple judgment 81 of the linguistic sentence 51, 52, 53 of the fourth data structure is given in Table No. 5 (the footnote under the symbol “*” means an ellipsis).
[0095] Table No. 5:
[0096] Preferably, without limitation, the identification of the value 811 of the simple proposition 81 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 during step 10091 is carried out by identifying the value 811 of the simple proposition 81 with the values 711 of all components of the simple proposition 71 that form this simple proposition 81. Preferably, without limitation, the identification of the ordinal numbers 812 of the simple proposition 81 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 during step 10091 is carried out by comparing the ordinal numbers of each simple proposition 81 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 with the ordinal numbers of other simple propositions 81 of the same linguistic sentence 51, 52, 53 of the fourth data structure 5. Preferably, not limiting itself, such a simple judgment 81, which has a component of the simple judgment 71 with the minimum ordinal number, will have the ordinal number “1”.If there is more than one such simple judgment 81, then the next component of the simple judgment 71, which has the next highest ordinal number, should be checked for such simple judgment 81, and without limitation, such simple judgment 81 ultimately receives the ordinal number "1". Preferably, without limitation, the next highest ordinal numbers are established according to the same rules, and without limitation, such simple judgments 81 that have received the ordinal number 812 no longer participate in the aforementioned comparison. Such identification and formation of the elements 81 of the seventh data structure 8 can be performed by any method known from the prior art and, accordingly, are not described in detail below.For example, but not limited to, such a complex analysis can be performed traditionally by a linguist, or using a software algorithm in a linguistic (syntactic) processor, or using traditional programming by solving problems using the coding of immutable rules (a "rule-based" IT solution). Furthermore, given a sufficient number of examples, such an analysis can be performed using a statistical processor (neural network, AI system) through the use of neural network (AI system) training technology.
[0097] Preferably, without limitation, the formation of the seventh data structure 8 during step 10092 is carried out by combining in one data structure the elements 81 of the seventh data structure 8 and their identification data according to principles and methods known from the prior art, which, accordingly, are not described in detail below.
[0098] Fig. 21, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1010 of forming the eighth data structure 9. Preferably, without limitation, step 1010 is characterized by: executing step 10101 forming elements 91 of the eighth data structure 9, wherein, on the basis of information contained in the database of linguistic-logical-subject features 30 and in the seventh data structure 8, as well as in accordance with the current formalized model of the logical construction of the judgment, elements 91 of the eighth data structure 9 are identified and formed, which are the final judgments 91 of the corresponding linguistic sentences 51, 52, 53 of the fourth data structure 5, as well as the identification data of the final judgments 91, which are the values 911 of the final judgments 911 and the ordinal numbers 912 of the final judgments 91 in the eighth data structure 9; performing step 10102 of forming the eighth data structure 9, wherein the eighth data structure 9 is formed from the identified final judgments 91 and their identification data.
[0099] In Fig. 22, by way of example, but not limitation, the general structure of the generated eighth data structure 9 is shown. Preferably, without limitation, the eighth data structure 9 (eighth SD 9) contains elements 91 of the eighth data structure 9, which are the resulting judgments 91 of each linguistic sentence 51, 52, 53 of the fourth data structure 5 and the identification data of said resulting judgments 91, which are, by way of example, but not limitation, the values 911 of said resulting judgments 91 and the ordinal numbers 912 of said resulting judgments 91 in the eighth data structure 9. Preferably, without limitation, the resulting judgments 91 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 are, from the point of view of cause-and-effect relationships, conditioned judgments and / or unconditional judgments.Unconditional judgments are assertions or refutations about the subject of the judgment that do not imply the presence of any conditions of the assertion or refutation. In other words, if a judgment has the characteristic of being unconditional, it is not preferable to have the form of a simple judgment. Preferably, such a single simple judgment within the final judgment is the principal simple judgment (rule). Conversely, conditional judgments preferably imply the presence of a certain condition or set of conditions under which assertions or refutations about the subject of the judgment are valid or true. Preferably, any conditioned proposition has the form of a complex proposition, that is, a group of simple propositions whose elements are linked to one another by a subordinate syntactic relation. Moreover, preferably, a simple proposition whose elements do not contain words that have the role of a syntactic descendant (a subordinate word in a word pair with a subordinate syntactic relation) is the principal simple proposition. (rule), and a simple judgment, the elements of which contain a word or words that have the role of a syntactic descendant, is an influencing simple judgment (conditionality).
[0100] Preferably, without limitation, based on the fact of the presence of unconditional and conditioned judgments, the final judgments 91 may consist of simple judgments 81 of two types - first simple judgments 81.1 (FSJ 81.1), which include the main simple judgments (rules), and second simple judgments 81.Y (SSJ 81.Y) (where Y>2 is the ordinal index of the SSJ 81.Y in the final judgment 91), which include the influencing simple judgments (conditionalities). Moreover, both types of simple judgments 81 (1111C 81.1 and IPC 81.Y) are formed at step 1009 from the components of the simple judgment 71. In this case, preferably, unconditional judgments contain only one IPC 81.1, and conditioned judgments contain one IPC 81.1, but also, in addition, one or more IPC 81.Y. Preferably, the ordinal index “Y” of the IPC 81.Y as part of the final judgment 91 is determined by the ordinal number 612 of the text element 61 in the element of the final judgment 91 of the IPC 81.Y, which has the role of a syntactic descendant.Preferably, the element of the final judgment 91 of the VPS 81.Y, which has the role of a syntactic descendant with the minimum ordinal number 612 of the text element 61 among all the elements of the final judgment 91 of the VPS 81.Y of one conditioned judgment, receives the index "2" (i.e., VPS 81.2). Preferably, the remaining unnumbered elements of the VPS 81.Y of the same final judgment 91 are then numbered according to the same principle, receiving further ordinal indices "Y" equal to "3", "4", and so on. As an example, but not limitation, the first and second simple judgments in the linguistic sentence 51, 52, 53 can be demonstrated. For example, without limitation, the linguistic sentence "The police immediately come to the aid of everyone who needs their protection from criminal and other unlawful attacks.» contains the following simple judgments 81: «The police immediately come to the aid of everyone»; «Who needs its protection from criminal attacks»; «Who needs its protection from other illegal attacks»; from the simple judgments 81 of the above-mentioned linguistic sentence, two complex judgments are formed; in this case, the syntactic parent - the main text element 61 in the pair of text elements 61 with a subordinate syntactic connection - is the text element 61 «to everyone», and the syntactic descendant is the text element 61 «who»; in this case, accordingly, the first simple judgment 81.1 (IP1C 81.1.) is the simple judgment «The police immediately come to the aid of everyone»; «Who needs its protection from criminal attacks», and the second simple judgments are the simple judgment 81.2 (VPS. 81.2) "Who needs its protection from criminal attacks" and the second simple judgment 81.3 (VPS 81.3) "Who needs its protection from other illegal attacks." Preferably, in this way two final judgments 91 are obtained, which are a complex judgment by the composition of their elements. One final judgment 91, which is a complex judgment, consists of 1П1С 81.1 and VPS 81.2 ("The police immediately come to the aid of everyone who needs its protection from criminal attacks"), and the other final judgment 91, also a complex judgment, consists of 1П1С 81.1 and VPS 81.3 ("The police immediately come to the aid of everyone who needs its protection from other illegal attacks"),
[0101] Preferably, but not limited to, the resulting judgments 91 of the linguistic sentence are, from the point of view of the interrelationship of judgment rules, independent judgment rules and interdependent judgment rules; wherein the independent judgment rules contain one IP1C 81.1, and the interdependent judgment rules contain two or more IP1C 81.1. Preferably, the interdependent judgment rules are formed on the basis of the dependency features of simple judgments established in the judgment formation requirements for the current formalized model of the logical construction of a judgment, contained in the second user database. As an example, but not limitation, such requirements may concern the presence of a coordinating relation of the logical "AND" type between individual IP1C 81.1.Without limitation, such a relationship indicates the interdependence of such judgment rules, or, in other words, the distortion of the essence of judgment when using individual judgment rules that are interdependent judgment rules. Preferably, the presence of "interdependent rules" can be demonstrated, by way of example, but not limitation, in the summary judgments 91 of the following sentence: "Exceeding the established speed limit of a vehicle by more than 20 but not more than 40 kilometers per hour shall entail the imposition of an administrative fine in the amount of five hundred rubles." This sentence contains two simple judgments: "Exceeding the established speed limit of a vehicle by more than 20 kilometers per hour shall entail the imposition of an administrative fine in the amount of five hundred rubles"; and "Exceeding the established speed limit of a vehicle by no more than 40 kilometers per hour shall entail the imposition of an administrative fine in the amount of five hundred rubles."At the same time, without limitation, the use of a linguistic tool that defines the boundaries of something (“from. . . to”; “more. . . , but not more”; “less. . . , but not less”; “from. . . to”; and the like) speaks of the interconnectedness and inseparability of the simple judgments formed, based on the logical “AND”. Preferably, therefore, those indicated in such. In this example, two simple judgments are interdependent judgments and, from the point of view of the final judgments 91, represent one final judgment 91. Another example, but not a limitation, of identifying between IP1C 81.1. a coordinating connection of the logical "AND" type is the presence of an adverbial turnover in a sentence, when the adverbial participle is a syntactic descendant of a word from IP1C 81.1. For example, in the sentence "The driver must drive a vehicle at a speed not exceeding the established limit, taking into account the traffic density," three simple judgments 81 are identified: "The driver must drive a vehicle at a speed"; "The driver taking into account the traffic density"; "At a speed not exceeding the established limit."In this case, it is preferable that the simple judgments “The driver must drive the vehicle at a speed” and “The driver takes into account the traffic intensity” are interdependent judgments, and the simple judgment “At a speed not exceeding the established limit” is a condition of the simple judgment “The driver must drive the vehicle at a speed”.
[0102] Preferably, without being limited to, the value 911 of the final judgment 91 are the values 811 of the simple judgments 81 from which the final judgment 91 is formed. In this case, the values 811 of the simple judgments 81 are the value 811.1 of the first simple narrowing (FSN 81.1.) and the value 811.Y of the second simple judgments (SSJ 81.Y) from which the final judgment 91 is formed. Preferably, without being limited to, the ordinal number 912 of the final judgment 91 are the ordinal numbers of the final judgment 91 in the eighth data structure 9. In the data structure, the final judgments 91, by way of example, but not limitation, can be named as “IS1”, “IS2”, “ISZ”, “ISp”, where n>1 is the ordinal number of the element 91 in the eighth data structure 9.Preferably, without limitation, the assignment of ordinal numbers 912 to the final judgments 91 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 is carried out as follows: the ordinal number “1” is assigned to the final judgment 91 formed from the linguistic sentence 51, 52, 53 with the ordinal number “1” and consisting of the simple judgment 81 with the ordinal number “1”. Preferably, if in the linguistic sentence 51, 52, 53 with the ordinal number "1" the simple judgment 81 with the ordinal number "1" belongs to the conditioned judgment, then the ordinal number "1" is assigned to such a final judgment 91 that has the minimum number of second simple judgments 81.Y with the minimum ordinal number 811 of the simple judgment 81. Preferably, the ordinal number "2" is assigned to such a final judgment 91 in which the elements of the final judgment 91 of the ISS 81.Y have a higher ordinal number of the simple judgment 81 than in the final judgment 91 with the ordinal number "1", or, if the same. there are no second simple judgments 81.Y, then such a final judgment 91 in which the elements of the final judgment 91 IP1C 81.1 have a greater ordinal number of the simple judgment 81 than in the final judgment 91 with the ordinal number "1". Preferably, the same rule also extends further to determining the ordinal numbers 912 of all elements 91 of the eighth data structure 9. Preferably, without limitation, in the case where IP1C 81.1 of the final judgment 91 represents a group of interrelated judgment rules, then the numbering of individual simple judgments included in the group of interrelated judgment rules is not performed, since these simple judgments already have unique ordinal numbers, like the simple judgment 81.
[0103] Preferably, but not limited to, the identification and formation of final judgments 91 of the eighth data structure 9 is carried out iteratively during step 10101. In the first step of step 10101, the rules of final judgments 91 (first simple judgments 81.1) are identified. In the second step of stage 10101, the identification of the conditions of the final judgments 91 (the second simple judgments 81.Y) is carried out for the identified rules of the final judgments 91. In the third step of stage 10101, the combination of the identified rules of the final judgments 91 and the conditions of the final judgments 91 is carried out to form the final judgments 91. Preferably, without limitation, the identification of the rules of the final judgments 91 (the first simple judgments 81.1) in the first step of stage 10101 is carried out by means of a third complex analysis of the elements 81 of the seventh data structure 8 - the simple judgments 81 and their identification data.Such analysis of simple judgments 81 is carried out using information about text elements 61 and using information from the generated BDL 1111120, as well as taking into account the requirements for a simple judgment 81 as a simple judgment of the first type, that is, a simple judgment whose components do not contain text elements that are syntactic descendants. Preferably, the objective of said third complex analysis is to identify among the elements 81 of the seventh data structure 8 such simple judgments 81 that meet the requirements for the first simple judgment 81.1. Preferably, without limitation, the identification of the conditions of the final judgments 91 (second simple judgments 81.Y) at the second step of stage 10101 is carried out by means of a fourth complex analysis of the elements 81 of the seventh data structure 8 - simple judgments 81 and their identification data.Preferably, without limitation, such analysis of simple judgments 81 is carried out using information about text elements 61 and using information from the generated BDLL1P120, as well as taking into account the requirements for a simple judgment 81 as a simple judgment of the second type, that is, a simple judgment whose components contain. text elements that are syntactic descendants. Preferably, the purpose of said fourth complex analysis is to identify, for each identified first simple judgment 81.1, among the elements 81 of the seventh data structure 8, such simple judgments 81 that meet the requirements for the second simple judgment 81.Y. Preferably, without limitation, the combination of the identified rules of the resulting judgments 91 (the first simple judgments 81.1) and the conditions of the resulting judgments 91 (the second simple judgments 81.Y) for forming the resulting judgments 91 in the third step of stage 10101 is performed by means of a fifth complex analysis of the identified rules of the judgments 91 and the conditions of the resulting judgments 91 and their identification data.Preferably, but not limited to, such analysis is performed using information about the text elements 61 and using information from the generated B DLILI 20, as well as taking into account the requirements for assembling the final judgments 91 from the first and second simple judgments. Preferably, the purpose of said fifth complex analysis is the identification and formation of elements 91 of the eighth data structure 9. By way of example, but not limitation, the requirements for forming the final judgments 91 from the first and second simple judgments contain, at a minimum, the following conditions: if no second simple judgment 81.Y is identified for the first simple judgment 81.1, then the final judgment 91 is formed only from one simple judgment 81 - 1П1С 81.1; if for the first simple judgment 81.1 one second simple judgment 81.Y is identified, then the final judgment 91 is formed from two simple judgments 81 - from 1П1С 81.1 and ВПС 81.Y; if for the first simple judgment 81.1 more than one second simple judgment 81.Y is identified, then in order to form the final judgment 91 it is necessary to establish syntactic subordinate connections between the identified second simple judgments 81.Y, after which it is necessary to form the final judgment 91 from the first simple judgment 81.1 and the second simple judgments that have a subordinate syntactic connection between themselves; without being limited to, as a result one of three options for forming the final judgment 91 can be realized: according to the first option, in which all identified second simple judgments 81.Y will have a subordinate syntactic connection between themselves, one element 91 will be formed from simple judgments 81 - from the PPS 81.1 and all the VPS 81.Y, located in the sequence of subordinate syntactic connections between themselves in accordance with the ordinal number of the index "Y"; according to the second option, in which all identified second simple judgments 81.Y will not have a continuous subordinate syntactic connection between themselves, as many elements 91 from the VPS 81.Y will be formed as there are second simple ones. judgments 81.Y will have the same syntactic descendant; according to the third option, if some of the identified second simple judgments 81.Y have a continuous subordinate syntactic relationship between themselves, and some of the identified second simple judgments 81.Y do not have a continuous subordinate syntactic relationship between themselves, then such a number of elements 91 will be formed that corresponds to the rules for forming judgments according to the first and second options. As an example, but not limitation, the following sentence may be considered: "The police are obliged to ensure that every citizen has the opportunity to become familiar with documents and materials directly affecting their rights and freedoms, unless otherwise provided by federal law."From the sentence under consideration, the following simple judgments 81 (PS 81) were formed with markings of individual text elements 61 (words) by their syntactic roles (СРп is the syntactic parent and СПп is the syntactic descendant, where п>1 is the ordinal number of the syntactic parent or syntactic descendant in this example), defining the syntactic subordinate relationship between the simple judgments of the sentence under consideration (Table No. 6):
[0104] Table No. 6:
[0105] During the first step of stage 10101 the first simple judgments are identified 81.1 (PPS 81.1) (table No. 7):
[0106] Table No. 7:
[0107] During the second step of stage 10101, for the identified first simple judgments 81.1, all second simple judgments 81.Y (SSP 81.Y) associated with them by syntactic subordination were identified (table no. 8):
[0108] Table No. 8:
[0109] During the third step of stage 10101 for forming final judgments 91 on the basis of each first simple judgment 81.1, a variant of forming the final judgment 91 is established. For both first simple judgments 81.1, it is established that, firstly, all second simple judgments 81.Y do not have any syntactic subordination relationship with each other, and, secondly, the two second simple judgments - VPS 81.3 and VPS 81.4 for 1П1С 81.1 with the ordinal number "1"; VPS 81.5 and VPS 81.6 for 1П1С 81.1 with the ordinal number "2" - have one syntactic descendant: SP2 - for VPS 81.3 and VPS 81.4 for 1П1С 81.1 with the ordinal number "1"; SPZ - for VPS 81.5 and VPS 81.6 for 1П1С 81.1 with ordinal number “2”), In this regard, on the basis of each first simple judgment 81.1 with ordinal numbers “1” and “2”, two final judgments 91 (IS 91) are formed (table No. 9):
[0110] Table No. 9
[0111] Preferably, such identification and generation of elements 91 of the eighth data structure 9 can be performed by any method known in the art and, accordingly, is not described in detail further. For example, without limitation, such a complex analysis can be performed traditionally by a linguist, or using a software algorithm of a linguistic (syntactic) processor, or based on traditional programming by solving problems using the coding of immutable rules (an IT solution "by rules"). Moreover, if a sufficient number of examples are available, such an analysis can be performed using a statistical processor (neural network, AI systems) by applying neural network (AI system) training technology.
[0112] Preferably, without limitation, the formation of the eighth data structure 9 during step 10102 is carried out by combining in one data structure the elements 91 of the eighth data structure 9 of the SMD, as well as their identification data, according to principles and methods known from the prior art, which, accordingly, are not described in detail below.
[0113] Fig. 23, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1011 of forming the ninth data structure 10. Preferably, without limitation, step 1011 is characterized by: performing step 10111, forming elements 92 of the ninth data structure 10, in which, on the basis of information contained in the database of linguistic-logical-subject features 30, in the second user database, in the sixth data structure 7, in the seventh data structure 8 and in the eighth data structure 9, as well as in accordance with the current formalized model of the basic structure of the subject area, elements 92 of the ninth data structure 10 are identified and formed, which are the basic structures of the subject area 92, as well as identification data of the basic structures of the subject area 92, representing the values 921 of the basic structures of the subject area and the serial numbers 922 of the basic structures of the subject area in the ninth data structure 10;performing step 10112 of forming the ninth data structure 10, in which the ninth data structure 10 is formed from the identified basic constructs of the subject area 92 and their identification data.
[0114] In Fig. 24, by way of example, but not limitation, the general structure of the generated ninth data structure 10 is shown. Preferably, without limitation, the ninth data structure 10 (ninth SD 10) contains elements 92 of the ninth data structure 10, which are basic constructions of the subject area 92 (BCDO 92) of each linguistic sentence 51, 52, 53 of the fourth data structure 5 and identification data of said BCDO 92, which are, by way of example, but not limitation, values 921 of said BCDO 92 and ordinal numbers 922 of said BCDO 92 in the ninth data structure 10. Preferably, without limitation, the basic constructions of the subject area 92 of the linguistic sentence 51, 52, 53 of the fourth data structure 5 represent, from a logical point of view, final judgments.In other words, if the eighth SD 9, containing elements 91, which are final judgments 91, is an exclusively logical construction regardless of the domain, then the ninth SD 10, containing elements 92, which are basic constructions of the subject area 92, is already a logical construction of the subject area. In this case, preferably, both the exclusively logical construction regardless of the domain and the logical construction of the subject area, from a logical point of view, consist of conditioned and / or unconditioned judgments. Preferably, the difference between the exclusively logical construction regardless of the domain and the logical construction of the subject area consists, from a logical point of view, in that the elements of simple judgments in the exclusively logical construction regardless of the domain are components of a simple judgment 71, containing logical objects. established in the current formalized model of the logical construction of judgments (FMLCS), and the elements of simple judgments in the logical construction of the subject area are the components of the basic subject construction 72 (BSC 72), containing the basic subject objects established in the current formalized model of the basic construction of the subject area (FMBCDO). In this case, the elements 71 in the FMLCS may be identical or not identical to the elements 72 in the FMBCDO. The rules for converting elements 71 into elements 72 are established in the correlation table of logical and subject objects contained in the second user database. In this case, also, that is, according to the same rules as the first simple judgments 81.1 and the second simple judgments 81.Y are formed from the components of the simple judgment 71 exclusively by the logical construction, regardless of the domain, the first simple judgments 82 are formed from the components of the basic subject construction 72.1 logical construction of the subject area (the first basic subject constructions 82.1) and the second simple judgments 82. Y of the logical construction of the subject area (the second basic subject constructions 82. Y). In this case, the first basic subject construction 82.1 and the second basic subject constructions 82. Y are the basic subject constructions 82 of the corresponding BCPO 92.
[0115] Preferably, without limitation, from the first basic subject constructions 82.1 and the second basic subject constructions 82.Y, a basic construction of the subject area 92 is formed, equivalent to the final judgment 91 in terms of the composition and content of the simple judgments contained in the final judgment 91, doing this in a similar manner as when forming the final judgment 91 from the first simple judgments 81.1 and the second simple judgments 81.Y, the process of which is described in detail with reference to step 1010.
[0116] Preferably, without being limited to, the basic constructions of the subject area 92 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, consisting of elements of two types - elements 82.1 and elements 82. Y - have identification data of the BCPO 92: as an example, but not limitation, the values 921 of the BCPO, consisting of the values 821.1 of the elements 82.1 and the values 821. Y of the elements 82. Y, and the ordinal number 922 of the BCPO 92 in the ninth data structure 10.
[0117] Preferably, without being limited to, the values 921 of the BCPO 92 are the values 821 of the basic subject constructions 82 from which this BCPO 92 is formed. In this case, the values 821 of the basic subject constructions 82 are the corresponding values 821.1 of the first basic subject construction 82.1 and the values 821. Y of the second basic subject constructions 82. Y from which the BCPO 92 is formed.
[0118] Preferably, without being limited, the ordinal numbers 922 of the BKPO 92 are the ordinal numbers of the BKPO 92 in the ninth data structure 10. In the data structure, the BKPO 92, by way of example, but not limitation, may be named as "BKPO1", "BKPO1", "BKPOZ", "BKPOn", where n>1 is the ordinal number of the element of the BKPO 92 in the ninth data structure 10. In this case, preferably, the ordinal numbering of the elements 92 in the ninth data structure 10 completely corresponds to the ordinal numbering of the elements 91 in the eighth data structure 9.
[0119] Preferably, without limitation, the identification and formation of elements 92 of the ninth data structure 10 is performed in the course of step 10111 in a step-by-step manner. In the first step of step 10111, the components of the basic subject construction 72 (BSC 72) are identified and formed from the elements of the CPS 71 of the sixth data structure 7. In the second step of step 10111, the basic subject constructions 82 are formed from the elements of the BSC 72. In the third step of step 10111, the first basic subject constructions 82.1 are identified and the second basic subject constructions 82. Y are identified. In the fourth step of step 10111, the identified first basic subject constructions 82.1 and the second basic subject constructions 82. Y are combined to form the BSSC 92 of the ninth data structure 10.
[0120] Preferably, without being limited to, the identification and formation of the KBPC 72 at the first step of stage 10111 is performed based on the data from the correlation table of the current basic subject objects (current BPO) and the current logical objects (current LogO), contained in the second user database (second PBD). In this case, the logical objects are the KPS 71, and the basic subject objects are the KBPC 72. Preferably, without being limited to, the composition of the elements of the KBPC 72 can be one or more KPS 71; in this case, the exact composition of each KBPC 72 of a unique name is established in the correlation table of the current BPO and the current LogO. Preferably, without limitation, the formation of basic subject constructions 82, in the second step of stage 10111, is carried out from the CPC 72 in a similar manner to the formation of simple judgments 81 from the components of simple judgments 71, the process of which is described in detail with reference to stage 1009.Preferably, without limitation, the identification of the first basic subject constructions 82.1 and the second basic subject constructions 82.Y in the third step of stage 10111 is carried out similarly to the identification of the rules of the final judgments 91 (the first simple judgments 81.1) and the conditions of the final judgments 91 (the second simple judgments 81.Y), the identification process of which is described in detail with reference to stage 1010. Preferably, without limitation, the combination of the identified first basic subject constructions 82.1 and second basic subject constructions 82. Y for the formation of basic constructions of the subject area 92 and their identification data in the fourth step of stage 10111 is carried out in a similar manner as the formation of final judgments 91 and their identification data, the process of which is described in detail with reference to stage 1010.
[0121] By way of example, but not limitation, for a legal subject area, a judgment may be related to the basic subject construct of the "structural part of a legal norm." Specifically, by way of example, but not limitation, it may be related to a disposition (i.e., a rule that must be observed), a sanction (i.e., a rule that determines the extent of liability for violating the rules), or a hypothesis (i.e., a conditionality of a rule reflecting some preliminary action, situation, or condition). These legal objects—hypothesis, disposition, and sanction—are contained, among other things, in simple sentences of regulatory acts. To transform a judgment in a legal subject area, it is necessary to create a formalized model of the basic construct of the subject area—a formalized model of the structural part of a legal norm (FMSCLN). Several different FMSCLNs may be formulated within the framework of professional discussion.To form the ninth data structure of the 10 SMD, it is necessary to create an up-to-date FMSCHPN, as well as a correlation table between the elements of the current formalized model of the logical construction of a judgment (FMLCS) and the elements of the current formalized model of the structural part of a legal norm. Moreover, without limitation, specialists in this technical field must clearly understand the rigid connection between a simple logical judgment and a part of a legal norm (hypothesis, disposition, sanction). This is demonstrated by way of example, but not limitation, in the following examples within the framework of some formalized models of the logical construction of a judgment, a formalized model of the structural part of a legal norm, and a correlation table, all generated solely for the purpose of illustration.By way of example, but not limitation, the following sentence from the normative act, the Law on Police, is considered: "When addressing a citizen, a police officer is obliged, in the event of the application of measures restricting his rights and freedoms, to explain to him the reasons and grounds for the application of such measures, as well as the rights and obligations of the citizen arising in connection therewith." By way of example, but not limitation, the current formalized model of the logical construction of a judgment may contain the following elements of a simple judgment - components of a simple judgment 71 (Table No. 10):
[0122] Table No. 10: Components of a simple judgment (KPS 71)
[0123] As an example, but not limitation, the current formalized model of the structural part of a legal norm may contain the following elements - components of the basic subject construction 72 (BSC 72) (Tables No. 11 and No. 12):
[0124] Table No. 11:
[0125] Table No. 12:
[0126] By way of example, but not limitation, at step 1010, an eighth SD 9 was formed, representing the final judgments of the 91 sentences under consideration (table no. 13):
[0127] Table No. 13:
[0128] As an example, but not limitation, at stage 1011 the tenth SD 10 was formed, representing the basic design of the subject area, which in the subject area of law is a formalized model of the structural part of a legal norm (tables No. 14 and No. 15):
[0129] Table No. 14:
[0130] Table No. 15: Components of the 2nd part of the formalized model of the structural part of a legal norm (KBPC 72) Components of the 2nd part of the formalized model of the structural part of a legal norm (KBPC 72)
[0131] Preferably, without being limited, such identification and generation of elements 92 of the ninth SD 10 can be performed by any method known from the prior art and, accordingly, are not described in detail further. For example, without being limited, such identification and generation can be performed traditionally by a legal specialist, or based on traditional programming by solving problems using the coding of immutable rules (IT solution "by rules") using A software algorithm for a linguistic (syntactic) processor. Moreover, given a sufficient number of examples, such analysis can be performed using a statistical processor (neural network, AI system) by applying neural network (AI system) training technology.
[0132] Preferably, without limitation, the formation of the ninth data structure 10 during step 10112 is carried out by combining elements in one data structure 92 ninth data structure 10, as well as their identification data according to principles and methods known from the prior art, which, accordingly, are not described in detail further.
[0133] Fig. 25, by way of example and not limitation, shows a general diagram of the execution of the steps of step 1012 for forming the final data structure 12.Preferably, without limitation, step 1012 is characterized by: performing step 10121 of forming elements 93 of the final data structure 12, at which, on the basis of information contained in the third user database and in the ninth data structure 10, as well as in accordance with the current formalized model of the target construction of the subject area, elements 93 of the final data structure 12 are identified and formed, which are target constructions of the subject area 93, as well as identification data of the target constructions of the subject area 93, representing: values 931 of the target constructions of the subject area 93, and serial numbers 932 of the target constructions of the subject area 93 in the final data structure 12; performing step 10122 of forming the final data structure 12, at which the final data structure 12 is formed from the identified target constructions of the subject area 93 and their identification data.
[0134] In Fig. 26, by way of example, but not limitation, the general structure of the generated final data structure 12 is shown. Preferably, without limitation, the final data structure 12 (final SD 12) contains elements 93 of the final data structure 12, which represent target constructions of the subject area 93 (TCS 93) of each linguistic sentence 51, 52, 53 of the fourth data structure 5 and identification data of said target constructions of the subject area 93, which represent, by way of example, but not limitation, values 931 of said elements 93 and ordinal numbers 932 of said elements 93 in the final data structure 12.
[0135] Preferably, but not limited to, domain-specific target constructs 93 (TsKPO 93) linguistic sentences 51, 52, 53 of the fourth data structure 5 represent single-element, multi-element or mixed constructions. Preferably, single-element constructions are such elements 93 of the final data structure 12, which are actually identical to the elements 92 of the ninth data structure 10 - the basic constructions of the subject area 92 (BCS 92); in other words, the type (composition of elements) of the target construction of the subject area 93 in single-element constructions coincides with the type (composition of elements) of the basic construction of the subject area 92. In this case, preferably, all identified elements of the BCS 93 have the same unique functional names.
[0136] Preferably, but not limited to, multi-element structures imply the presence of at least two elements of the CKPO 93 with different unique functional names, as well as the fulfillment of conditions under which individual BKPO 92 are identified as elements of the CKPO 93 with different unique functional names, as well as the fulfillment of rules for combining identified BKPO 92 with different unique functional names into a multi-element structure. Preferably, but not limited to, mixed structures imply the presence of both single-element structures and multi-element structures in the CKPO 93.
[0137] Preferably, without limitation, in connection with the presence of the above-described types of structures of elements 93 (single-element, multi-element, mixed structures), the target structures of the subject area 93 can consist of BCPO 92 of two types - the main BCPO 92.1 and additional BCPO 92. Y; wherein the single-element structures contain only one main BCPO 92.1, and the multi-element structures contain one main BCPO 92.1, but also additionally contain one or more additional BCPO 92. Y, where Y>2 is the ordinal index of the BCPO 92 of the unique name in the composition of the CCPO 93.
[0138] By way of example, but not limitation, it is possible to demonstrate the main BKPO 92.1 and additional BKPO 92. Y in the subject area of law; it is necessary to clarify that in the subject area of law the target construction of subject area 93 (TCPO 93), by way of example, but not limitation, can be the construction of a legal norm. Unlike the composition of the structural elements of a legal norm (hypothesis, disposition, sanction), the composition of the elements of a legal norm (the construction of a legal norm) is a controversial issue in the legal community. There are many concepts concerning the composition of legal norms, among which the following main groups can be distinguished according to the constructions (composition) of a legal norm: legal norms regulating legal relations in everyday life, which are two-element constructions consisting, at a minimum, of the structural parts of a legal norm "disposition" and "sanction"; legal norms establishing normative definitions individual subjects or objects of legal relations, representing single-element constructions consisting, at a minimum, of the structural part of the legal norm "disposition"; legal norms establishing principles, guarantees, declarations, representing single-element constructions consisting, at a minimum, of the structural part of the legal norm "disposition"; legal norms having a single-element or two-element construction and containing the structural part of the legal norm substantiating the disposition, wherein such disposition represents the result or consequence of the execution of the rules and / or conditions established in other normative rules confirming the status of the disposition as relevant (valid and permissible) for its application in a specific branch (institution, sub-institution) of law, on the basis of specific legal principles, guarantees and declarations, in specific legal circumstances, in specific territories, in a specific time period.Such other normative rules will constitute an integral structural part of the aforementioned dispositions—a hypothesis. The number of such hypotheses may not be limited to one, but rather represent a construct of hypotheses, among which the relevance of the disposition will be substantiated by not one, but several. Furthermore, such a construct of hypotheses may include not only hypotheses substantiating the relevance of the disposition, but also hypotheses substantiating the relevance of the hypotheses included in the construct of hypotheses. Preferably, and without limitation, the 93rd data structure includes elements 92.1 and 92.Y of the ninth data structure 10.
[0139] Preferably, without being limited, the target constructions of the subject area 93 of the linguistic sentence 51, 52, 53 of the fourth data structure 5, consisting of elements 92 of two types - 92.1 and 92.Y, have identification data of the CPC 93: as an example, but not limitation, values 931 of the CPC 93, consisting of values 921.1 and 921.Y of the elements 92.1 and 92. Y, and ordinal numbers 932 of the CPC 93, which are ordinal numbers 932 of the CPC 93 in the final data structure 12.
[0140] Preferably, without being limited to, the values 931 of the CPC 93 are the values 921.1 and 921. Y of the corresponding elements 92.1 and 92. Y of the corresponding BCPO 92 of the ninth data structure 10, from which the corresponding CPC 93 is formed.
[0141] Preferably, but not limited to, the ordinal numbers 932 of the CPC 93 are the ordinal numbers 932 of the CPC 93 in the final data structure 12. In the data structure of the CPC 93, as an example, but not limitation, each element 93 can be referred to as "CPC1", "CPC2", "CPC3", "CPCn", where n>1 is the ordinal number of the element of the CPC 93 in the final data structure 12. Preferably, but not limited to, the ordinal numbering of the elements 93 in the array of target constructions of the subject area linguistic sentence 51, 52, 53 of the fourth data structure 5 is performed as follows: the ordinal number "1" is assigned to the CKPO 93 formed from the linguistic sentence 51, 52, 53 with the ordinal number "1", consisting of the BKPO 92 with the ordinal number "1"; if in the linguistic sentence 51, 52, 53 with the ordinal number "1" the element 92 with the ordinal number "1" refers to the additional BKPO 92 (92. Y), then the ordinal number "1" is assigned to such CKPO 93, which has the minimum number of BKPO 92. Y with the minimum ordinal number; The ordinal number "2" is assigned to such a CKPO 93 in which the elements of the BKPO 92. Y have a higher ordinal number of the BKPO 92. Y than in the CKPO 93 with the ordinal number "1", or, if there are no such BKPO 92. Y, then to such a CKPO 93 in which the elements of the BKPO 92.1 have a higher ordinal number of the BKPO 92.1 than in the CKPO 93 with the ordinal number "1"; and so on, according to the same principle, all elements 93 of the final data structure 12 SMD are numbered with an ordinal number.
[0142] Preferably, but not limited to, the identification and formation of the final data structure CKPO 93 is performed in the course of step 10121 in a step-by-step manner. In the first step of step 10121, the main BKPO 92.1 of the CKPO 93 is identified. In the second step of step 10121, the additional BKPO 92. Y of the CKPO 93 is identified for the identified elements of the BKPO 92.1 of the multi-element constructions of the CKPO 93. In the third step of step 10121, the identified main BKPO 92.1 of the CKPO 93 and the additional BKPO 92. Y of the CKPO 93 (if such were identified for the corresponding BKPO 92.1) are combined to form the CKPO 93 of the linguistic sentence 51, 52, 53 of the fourth data structure 5.
[0143] Preferably, without being limited to, the identification of the main BKPO 92.1 TsKPO 93 at the first step of stage 10121 is carried out by means of the sixth complex analysis of the elements of the ninth data structure 10 SMD - BKPO 92 and their identification data; such analysis of BKPO 92 is carried out using information about the text elements 61 and using information from the generated BDLZ 1111130, as well as taking into account the requirements for BKPO 92 as the main BKPO 92.1, obtained, for example, without being limited to, from the data of the correlation table of the basic constructions of the subject area of the ninth SD 10 and the current elements of the target constructions of the subject area, contained in the formalized model of the target construction of the subject area (FMTSKPO); the purpose of the said sixth complex analysis is to identify among the elements 92 of the ninth data structure 10 such BCPO 92 that meet the requirements for the main BCPO 92.1 of the CKPO 93.
[0144] Preferably, without limitation, the identification of additional BKPO 92. Y TsKPO 93 in the second step of stage 10121 is carried out by the seventh complex analysis of the elements of the ninth data structure 10 SMD - BKPO 92 and their identification data; such analysis of BKPO 92 is performed using information about the text elements 61 and using information from the generated BDLZ 11111 30, as well as taking into account the requirements for BKPO 92, as an additional BKPO 92. Y, obtained, for example, without limitation, from the data of the correlation table of the basic constructs of the subject area of the ninth SD 10 and the current elements of the target constructs of the subject area contained in the FMCKPO; the purpose of the aforementioned seventh complex analysis is to identify among the elements of the ninth data structure 10 such BKPO 92 that meet the requirements for additional BKPO 92. Y CKPO 93.
[0145] Preferably, without limitation, the combination of the identified main BKPO 92.1 TsKPO 93 and additional BKPO 92. Y TsKPO 93 to form the TsKPO 93 at the third step of stage 10121 is carried out by means of the eighth comprehensive analysis of the identified main BKPO 92.1 TsKPO 93 and additional BKPO 92. Y TsKPO 93 and their identification data; such analysis is carried out using information about the text elements 61 and using information from the generated BDLZ 1111130, as well as taking into account the requirements for the formation of the TsKPO 93 from the main BKPO 92.1 and additional BKPO 92. Y contained in the FMCKPO; the purpose of the mentioned eighth complex analysis is the identification and formation of the CKPO 93 of the linguistic sentence 51, 52, 53 of the fourth data structure 5. As an example, but not limitation, the requirements for the formation of the CKPO 93 from the main BCPO 92.1 and additional BCPO 92. Y contain, at a minimum, the following conditions: if no additional BCPO 92 is identified for the main BCPO 92.1.Y, then the CKPO 93 is formed from only one element of the CKPO 93 - from the BKPO 92.1; if one additional BKPO 92. Y is identified for the main BKPO 92.1, then the CKPO 93 is formed from two elements of the CKPO 93 - from the main BKPO 92.1 and BKPO 92. Y; if more than one additional BKPO 92. Y is identified for the main BKPO 92.1, then in order to form the CKPO 93, it is necessary to establish a unique name for the additional BKPO 92. Y, after which it is necessary to form the CKPO 93 from the main BKPO 92.1 and additional BKPO 92. Y that have a unique name, in accordance with the requirements for the formation of the CKPO 93 based on the formalized model of the target design of the subject area (FMTSKPO).As an example, but not a limitation, the following sentences can be considered, on the basis of which it is possible to demonstrate the stage of formation of the Central Traffic Regulation 93: first: “The driver must drive the vehicle at a speed not exceeding the established limit, taking into account the traffic intensity and the condition of the vehicle”; second: “Exceeding the established speed limit. "A vehicle speeding more than 20 but not more than 40 kilometers per hour shall entail the imposition of an administrative fine in the amount of five hundred rubles."
[0146] From the first sentence under consideration at stage 1010, the following final judgments 91 were formed (table No. 16):
[0147] Table No. 16:
[0148] From the second sentence under consideration at stage 1010, the following final judgments 91 were formed (table No. 17):
[0149] Table No. 17:
[0150] From the first sentence under consideration at stage 1011, the following basic constructions of subject area 92 were formed, which represent structural parts of legal norms in the subject area of law (tables No. 18 and No. 19):
[0151] Table No. 18:
[0152] Table No. 19:
[0153] From the second sentence under consideration, at stage 1011, the following BKPO 92 were formed, which represent structural parts of legal norms in the subject area of law (tables No. 20 and No. 21):
[0154] Table No. 20:
[0155] Table No. 21:
[0156] From the first and second sentences considered at stage 1012, the following CCPO 93 were formed, which represent legal norms in the subject area of law, which, according to the formalized model of the target design of the subject area, are two-element constructions of a legal norm, consisting of the structural parts of the legal norm “disposition” and “sanction” (tables No. 22, No. 23, No. 24, No. 25):
[0157] Table No. 22: DISPOSITION OF THE LEGAL NORMA 0158] Table No. 23: 0159] Table No. 24: SANCTION OF LEGAL NORMS 0160] Table No. 25
[0161] Such identification and generation of the final SD 12 dataset can be accomplished by any method known in the art and, accordingly, is not described in detail below. For example, but not limited to, such identification and generation can be performed traditionally by a legal professional, or based on traditional programming by solving problems through the coding of immutable rules (an IT solution "by rules") using the software algorithm of a linguistic (syntactic) processor. Furthermore, given a sufficient number of examples, such analysis can be performed using a statistical processor (neural network, AI systems) through the use of neural network (AI system) training technology.
[0162] Preferably, without limitation, the formation of the final data structure 12 during step 10122 is carried out by combining in one data structure the elements 93 (DC 93) of the final data structure 12, as well as their identification data, according to principles and methods known from the prior art, which, accordingly, are not described in detail below.
[0163] In Fig. 27, by way of example and not limitation, an exemplary diagram of a structured data array transformation system 2000 is illustrated, which in a preferred embodiment comprises at least one or more computer devices 2001 for transforming a structured data array, comprising at least one or more processors 20011 and memory 20012. The said devices 2001 for converting a structured data array may be, but are not limited to, a personal computer, a laptop, a tablet computer, a pocket computer, a smartphone, a phablet, and the like. The memory (machine-readable storage medium) 20012 of the device 2001 for converting a structured data array contains program code that, when executed, causes the said one or more processors 20011 of the said device 2001 to perform the actions of the previously described methods for converting a structured data array.In some cases, the computing device 2001 may be a server computing device associated with a user computing device configured to transmit to the server computing device 2001 a command or commands causing the processor or processors 20011 of the server computing device to execute program code, which, when executed by the processor or processors of the server computing device 20011, causes the processor or processors 20011 of the server computing device to perform actions of any of the previously described conversion methods. structured data array. User computing device 2002 may be, but is not limited to: a personal computer, a laptop computer, a tablet computer, a pocket computer, a smartphone, a phablet, a thin client, and the like. User computing device 2002 may be connected to server computing device 2001 via a wired or wireless connection. Said memory 20012 of computing device 2001 (server computing device 2001) contains one or more structured data arrays to be converted, containing at least a linguistic sentence, and may also contain any of the previously described data structures for any of the previously described methods for converting a structured data array.Moreover, one or more structured data arrays, user databases, other databases, data models and tables, and other data to be converted may be loadable and stored, in particular, in the database 2003 of the structured data array conversion system. By way of example, but not limitation, the machine-readable storage medium (memory 20012) may include random access memory (RAM); read-only memory (ROM); electrically erasable programmable read-only memory (EEPROM); flash memory or other memory technologies; CDROM, digital versatile disc (DVD) or other optical or holographic data carriers; magnetic cassettes, magnetic tape, magnetic disk storage device or other magnetic storage devices, carrier waves or other data carrier that can be used to encode the desired information and that can be accessed by device 2001. Memory includes a data carrier based on a computer storage device in the form of volatile or non-volatile memory, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, and so on. The memory stores an exemplary environment in which, using computer commands or codes stored in the device's memory, a procedure for converting a structured data array can be performed.The device comprises one or more processors 20011, which are designed to execute computer instructions or codes stored in the device's memory to facilitate the execution of a structured data array conversion procedure. The computer instructions or codes stored in the memory are designed to perform the structured data array conversion. System 2000 may also include a database (DB) 2003. DB 2003 may comprise, but is not limited to: a hierarchical DB, a network DB, a relational DB, an object DB, an object-oriented DB, an object-relational DB, a spatial DB, a combination of two or more of these DBs, and the like. DB 2003 stores data in memory, which may be, but is not limited to: a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory, CDROM, digital versatile disk (DVD) or other optical or holographic storage media; magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, carrier waves or other storage media that can be used to store the required information and which can be accessed by means of the structured data array conversion device 2001.The DB 2003 serves to store data that represent at least commands for performing the steps of the previously described methods for transforming a structured data array; one or more structured data arrays to be transformed, containing at least a linguistic sentence, or one of the previously described source data structures for any transformation method, which can be loaded into the memory 20012 of the device 2001 for transforming a structured data array; and other data necessary for the functioning of the system. The exemplary system 2000 for transforming a structured data array may additionally comprise a server computer device 2001, which, in addition to the previously described functions, stores and facilitates the manipulation of computer commands or codes previously described in this document, which, accordingly, are not further described.The server computer device 2001, in addition to the functions described earlier, can provide for the regulation of the data exchange in the system 2000 for converting a structured data array, and also provides for the processing of data provided that one or more user computer devices 2002 are connected to it. In this case, all the computing power necessary for ensuring the execution of the procedure for converting a structured data array is located on the server computer device 2001. The system 2000 can also contain one or more data transmission networks 2004. The data transmission networks 2004 can include, but are not limited to, one or more local area networks (LAN) and / or wide area networks (WAN), or can represent an information telecommunications network, the Internet, or an Intranet, or a virtual private network (VPN), or a combination thereof, and the like.The server computing device 2001 also has the ability to provide a virtual computing environment (Virtual. Machine) for providing interaction between the user computer device 2002 and the database 2003. The network 2004 serves to provide interaction between the computer device 2001, the database 2003 and the user computer device 2002 of the system 2000 for converting a structured array of data. In this case, the user computer device 2002 can be connected to the server computer device 2001 directly, using wired and wireless communication methods and techniques known from the prior art, which, accordingly, are not described in more detail below. The mentioned devices 2001, 2002, by way of example, but not limitation, can be equipped with input-output (i / o) devices suitable for providing the user with the results of performing one or another of the previously described steps of any of the claimed methods described with reference to Figs. 1-26.
[0164] The present description of the implementation of the claimed invention demonstrates only particular embodiments and does not limit other embodiments of the claimed invention, since possible other alternative embodiments of the claimed invention, not going beyond the scope of the information set out in this application, should be obvious to a specialist in the given field of technology, having the usual qualifications for whom the claimed invention is intended.
Claims
Invention formula 1. A method, executed by a processor or processors of a computer device, for transforming a structured data array containing at least information objects of a digitized document, which are separate blocks of the information content of the digitized document, which are textual information objects, and / or which are visual information objects, and / or which are textual and visual information objects; the method is characterized by: performing step 1001 of forming a first data structure, in which a first data structure is formed, containing elements of the first data structure, which are semantic parts of the information objects of the digitized document, as well as identification data of such semantic parts, which are the values of such semantic parts and the ordinal numbers of such semantic parts in the digitized document;by performing step 1002 of forming a database of system features of semantic parts, in which system features of the semantic parts contained in the first data structure are identified, representing the format system characteristics of such semantic parts and the functional system characteristics of such semantic parts, as well as the values of the corresponding mentioned system characteristics, in order to identify semantic parts with structural system features, and / or to identify semantic parts with logical system features, and / or to identify semantic parts with information system features, and / or to identify semantic parts with requisite system features; and a database of system features of semantic parts is formed from such identified system features;performing step 1003 of forming a second data structure, in which a second data structure is formed, containing elements of the second data structure, which are integrated semantic parts of information objects of the digitized document, which are grouped semantic parts contained in the first data structure with matching system features or grouped semantic parts contained in the first data structure with unique system features, and also containing identification data of the integrated semantic parts, which are unique varieties of the said semantic parts with matching system features; or said semantic parts with unique system features, the values of said semantic parts with matching system features or said semantic parts with unique system features, and the serial numbers in the digitized document of said semantic parts with matching system features or said semantic parts with unique system features, wherein such said semantic parts with matching system features or such said semantic parts with unique system features constitute said integrated semantic parts; by performing step 1004 of forming a third data structure, in which a third data structure is formed containing, as elements of the third data structure, linguistic constructions representing the integrated semantic parts of the information objects of the digitized document contained in the second data structure,wherein such integrated semantic parts have systemic features of textual logical semantic parts, and also contain identification data of linguistic constructions, which are the values of linguistic constructions and ordinal numbers of linguistic constructions in the digitized document, wherein the linguistic constructions in the digitized document are one of the following linguistic constructions or a combination of the following linguistic constructions: ordinary linguistic constructions of the third data structure, which are linguistic sentences, special linguistic constructions of the third data structure, which are lists or enumeration headings, reconstructed linguistic constructions of the third data structure, which are tables containing at least two rows and two columns, wherein at least one row contains column headings and / or, respectively, at least,one column contains row headers; performing step 1005 of forming a fourth data structure, in which a fourth data structure is formed, containing as elements of the fourth data structure linguistic sentences of the fourth data structure, formed from elements of the third data structure and representing either linguistic sentences that are ordinary linguistic constructions of the third data structure, or linguistic sentences obtained by transforming special linguistic constructions of the third data structure, or linguistic sentences obtained by recreating from reconstructed linguistic constructions of the third data structure; wherein the fourth data structure also includes identification data of the linguistic sentences, a fourth data structure, representing the values of the linguistic sentences of the fourth data structure and the ordinal numbers of the linguistic sentences of the fourth data structure in the fourth data structure; performing step 1006 of forming the fifth data structure, in which a fifth data structure is formed, containing as elements of the fifth data structure the text elements of the linguistic sentences of the fourth data structure, as well as the identification data of the text elements of the linguistic sentences of the fourth data structure, representing the values of the corresponding text elements of the corresponding linguistic sentences of the fourth data structure and the ordinal numbers of the corresponding text elements in the corresponding linguistic sentences of the fourth data structure;performing stage 1007 of forming a database of linguistic-logical-subject features, in which linguistic, logical and subject features of text elements of linguistic sentences of the fourth data structure are identified and a database of linguistic-logical-subject features is formed from the identified features;performing step 1008 of forming a sixth data structure, in which a sixth data structure is formed, containing elements of the sixth data structure, which are components of a simple judgment, which are components of corresponding simple judgments, wherein the simple judgments are simple judgments of corresponding linguistic sentences of the fourth data structure, and also containing identification data of said components of a simple judgment, which are, for each corresponding component of a simple judgment from the sixth data structure, the type of such corresponding component of a simple judgment, the value of such corresponding component of a simple judgment and the ordinal number of such corresponding component of a simple judgment in the corresponding linguistic sentence;performing step 1009 of forming a seventh data structure, in which a seventh data structure is formed, containing elements of the seventh data structure, representing simple judgments of the corresponding linguistic sentences of the fourth data structure, and also containing identification data of such simple judgments, representing the values of the corresponding simple judgments and the ordinal numbers of the corresponding simple judgments in the corresponding linguistic sentences of the fourth data structure; performing step 1010 of forming an eighth data structure, in which an eighth data structure is formed, containing elements of the eighth data structure, representing final judgments of the corresponding linguistic sentences of the fourth data structure, formed from the mentioned simple judgments of the corresponding linguistic sentences of the fourth data structure, and also containing identification data of the final judgments, representing the values of the final judgments and the ordinal numbers of the final judgments in the eighth data structure;performing step 1011 of forming a ninth data structure, in which a ninth data structure is formed, containing elements of the ninth data structure, representing basic constructs of the subject area, formed from data, including data of the said sixth data structure formed as a result of performing step 1008, wherein the formation of basic constructs of the subject area is carried out on the basis of data on a formalized model of the basic construct of the subject area and data on a formalized model of the logical construct of the judgment, and also containing identification data of the basic constructs of the subject area, representing the values of the basic constructs of the subject area and the ordinal numbers of the basic constructs of the subject area in the ninth data structure;performing step 1012 of forming the final data structure, in which the final data structure is formed, containing elements of the final data structure, which are the target constructions of the subject area, formed from the basic constructions of the subject area, contained in the ninth data structure, wherein the formation of the target constructions of the subject area is carried out on the basis of data on the formalized model of the target construction of the subject area, and also containing identification data of the target constructions of the subject area, which are the values of the target constructions of the subject area and the ordinal numbers of the target constructions of the subject area in the final data structure.
2. The method according to paragraph 1, characterized in that step 1003 is characterized by: performing step 10031 of identifying and generating elements of the second data structure, in which elements of the second data structure are identified and generated, representing integrated semantic parts of information objects of the digitized document, and representing grouped semantic parts contained in the first data structure with matching system features or grouped semantic parts contained in the first data structure parts with unique system features, as well as identification data of integrated semantic parts, which are non-repeating varieties of the said semantic parts with matching system features or the said semantic parts with unique system features, the values of the said semantic parts with matching system features or the said semantic parts with unique system features, and serial numbers in the digitalized document of the said semantic parts with matching system features or the said semantic parts with unique system features, wherein such said semantic parts with matching system features or such said semantic parts with unique system features constitute the said integrated semantic parts;performing step 10032 of forming a second data structure, in which the second data structure is formed from the identified and formed elements of the second data structure and their identification data.
3. The method according to item 1, characterized in that step 1004 is characterized by: performing step 10041 of identifying and generating elements of the third data structure, in which elements of the third data structure are identified and generated, which are linguistic constructions that represent integrated semantic parts of the information objects of the digitized document contained in the second data structure, wherein such integrated semantic parts have systemic features of text logical semantic parts, and also contain identification data of linguistic constructions that represent the values of linguistic constructions and ordinal numbers of linguistic constructions in the digitized document, wherein the linguistic constructions in the digitized document are one of the following linguistic constructions or a combination of the following linguistic constructions: ordinary linguistic constructions of the third data structure,which are linguistic sentences, special linguistic constructions of the third data structure, which are lists or enumeration headings, reconstructed linguistic constructions of the third data structure, which are tables containing at least two rows and two columns, wherein at least one row contains column headings and / or, respectively, at least one column contains row headings; by performing step 10042 of forming the third data structure, in which the third data structure is formed from the elements of the third data structure identified and formed in step 10041 and their identification data.
4. The method according to claim 1, characterized in that step 1005 is characterized by: performing step 10051 of identifying and generating first elements of the fourth data structure, in which the first elements of the fourth data structure are identified and generated, as well as identification data of the first elements of the fourth data structure, representing for each first element of the fourth data structure the value of the first element of the fourth data structure and the ordinal numbers of the first element of the fourth data structure in the fourth data structure; wherein the first elements of the fourth data structure are linguistic sentences formed from elements of the third data structure containing ordinary linguistic constructions, by identifying the linguistic sentences of the fourth data structure with ordinary linguistic constructions of the third data structure;performing step 10052 of identifying and generating second elements of the fourth data structure, in which the second elements of the fourth data structure are identified and generated, as well as identification data of the second elements of the fourth data structure, which represent, for each second element of the fourth data structure, the value of the second element of the fourth data structure and the serial numbers of the second element of the fourth data structure in the fourth data structure; wherein the second element of the fourth data structure are linguistic sentences formed from elements of the third data structure containing special linguistic constructions, by converting the special linguistic constructions into linguistic sentences of the fourth data structure;performing step 10053 of identifying and generating third elements of the fourth data structure, in which the third elements of the fourth data structure are identified and generated, as well as identification data of the third elements of the fourth data structure, which represent, for each third element of the fourth data structure, the value of the third element of the fourth data structure and the serial numbers of the third element of the fourth data structure in the fourth data structure; wherein the third element of the fourth data structure are linguistic sentences generated from elements of the third data structure containing reconstructed linguistic constructions, by recreating individual linguistic sentences of the fourth data structure from the data contained in the reconstructed linguistic constructions; performing step 10054 of forming a fourth data structure, in which the fourth data structure is formed from the first elements of the fourth data structure, the second elements of the fourth data structure, the third elements of the fourth data structure and their identification data.
5. The method according to paragraph 1, characterized in that step 1007 is characterized by: performing step 10071 of forming the first part of the linguistic-logical-subject features of the text elements of the linguistic sentences of the fourth data structure, in which, for the linguistic analysis of the text elements contained in the fifth data structure and classified as a word, identification data of such text elements are provided and linguistic characteristics of such text elements are obtained, as well as the values of the said linguistic characteristics;performing step 10072 of forming the second part of the linguistic-logical-subject features of the text elements of the linguistic sentences of the fourth data structure, in which, for the logical analysis of the text elements contained in the fifth data structure and classified as a word, the identification data of such text elements are provided, as well as the linguistic characteristics of such text elements together with the mentioned values of the linguistic characteristics, and the logical characteristics of such text elements of each linguistic sentence are obtained, as well as the values of the mentioned logical characteristics;by performing step 10073 of forming the third part of the linguistic-logical-subject features of the text elements of the linguistic sentences of the fourth data structure, in which, for the subject analysis of the text elements contained in the fifth data structure and classified as a word, the identification data of such text elements are provided, as well as the linguistic characteristics of such text elements together with the mentioned values of the linguistic characteristics, and the logical characteristics of such text elements are provided, as well as the values of the mentioned logical characteristics, and the characteristics of such text elements in the subject area, which are the subject characteristics, are obtained, as well as the values of the mentioned subject characteristics;performing step 10074 of forming a database of linguistic-logical-subject features, in which a database of linguistic-logical-subject features of text elements of linguistic sentences of the fourth data structure is formed, wherein the linguistic-logical-subject features of text elements of linguistic sentences of the fourth data structure are the said linguistic characteristics obtained for each text element during steps 10071, 10072 and 10073; the said logical characteristics and the said subject characteristics, which respectively have the said values of linguistic characteristics, the said values of logical characteristics and the said values of subject characteristics.
6. The method according to i. 1, characterized in that step 1008 is characterized by: performing step 10081, forming elements of the sixth data structure, in which, on the basis of information contained in the database of linguistic-logical-subject features, the fifth data structure, and also on the basis of data from the first user database containing data on current syntactic units and current logical objects, as well as on the current formalized model of the logical construction of the judgment, elements of the sixth data structure are identified and formed, which are components of a simple judgment of the corresponding linguistic sentences of the fourth data structure, as well as identification data of the said components of the simple judgment, representing for each said component of the simple judgment the type of such corresponding component of the simple judgment,the value of such corresponding simple proposition component and the ordinal number of such corresponding simple proposition component in the corresponding linguistic sentence of the fourth data structure; by performing step 10082 of forming the sixth data structure, in which the sixth data structure is formed from the mentioned simple proposition components and their identification data.
7. The method according to claim 1, characterized in that step 1009 is characterized by: performing step 10091 of forming elements of the seventh data structure, at which, on the basis of information contained in the database of linguistic-logical-subject features and the sixth data structure, from the formed components of simple judgments in accordance with the current formalized model of the logical construction of judgments, elements of the seventh data structure are formed, which are simple judgments, as well as identification data of simple judgments, representing the values of the corresponding simple judgments and the ordinal numbers of the corresponding simple judgments in the corresponding linguistic sentences of the fourth data structure; performing step 10092 of forming the seventh data structure, at which the seventh data structure is formed from the said simple judgments and their identification data.
8. The method according to claim 1, characterized in that step 1010 is characterized by: by performing step 10101 of forming elements of the eighth data structure, in which, on the basis of information contained in the database of linguistic-logical-subject features and in the seventh data structure, as well as in accordance with the current formalized model of the logical construction of judgments, elements of the eighth data structure are identified and formed, which are the final judgments of the corresponding linguistic sentences of the fourth data structure, as well as the identification data of the final judgments, which are the values of the final judgments and the ordinal numbers of the final judgments in the eighth data structure; by performing step 10102 of forming the eighth data structure, in which the eighth data structure is formed from the said final judgments and their identification data.
9. Method according to p.1, characterized in that step 1011 is characterized by: performing step 10111 of forming elements of the ninth data structure, at which, on the basis of information contained in the database of linguistic-logical-subject features, in the second user database, in the sixth data structure, and also in accordance with the current formalized model of the basic construction of the subject area, the elements of the ninth data structure are identified and, using the current formalized model of the logical construction of the judgment, the elements are formed which are the basic constructions of the subject area, as well as the identification data of the basic constructions of the subject area, representing the values of the basic constructions of the subject area and the ordinal numbers of the basic constructions of the subject area in the ninth data structure; performing step 10112 of forming the ninth data structure, at which the ninth data structure is formed from the aforementioned basic constructions of the subject area and their identification data.
10. The method according to claim 1, characterized in that step 1012 is characterized by: performing step 10121 of forming elements of the final data structure, in which, on the basis of information contained in the third user database and in the ninth data structure, as well as in accordance with the current formalized model of the target construction of the subject area, elements of the final data structure are identified and formed, which are the target constructions of the subject area, as well as identification data of the target constructions of the subject area, representing the values of the target constructions of the subject area, and the ordinal numbers of the target constructions of the subject area in the final data structure; performing step 10122 of forming the final data structure, in which the final data structure is formed from the mentioned target subject area constructs and their identification data.
Citation Information
Patent Citations
Method of transforming structured data array
RU2713568C1
Extracting Structure and Semantics from Tabular Data
US20190278853A1
Unstructured data clustering of information technology service delivery actions
US20200012728A1
Contrasting Document-Embedded Structured Data and Generating Summaries Thereof
US20210271654A1
Unified data classification techniques
US20230418859A1