Method for automatically searching for conceptual parameters of a text

A matrix-based multiple linear regression analysis method identifies and quantifies the substantive parameters of text, addressing the challenge of automatic text analysis by accurately determining the significance and influence of keywords and concepts.

WO2026049649A1PCT designated stage Publication Date: 2026-03-05LITUEV VIKTOR NIKOLAEVICH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/RU2025/050236
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-02
Filing Date
2025-08-12
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

The problem of automatically identifying and analyzing the substantive parameters of text, such as keywords and key concepts, has not been effectively addressed by existing methods.

Method used

A method involving the formation of a matrix from nouns and verbs, followed by multiple linear regression analysis, to determine the substantive parameters of text by calculating the dependence of each parameter on observations, thereby identifying the 'weight' and significance of keywords and concepts.

Benefits of technology

This method allows for accurate identification and quantification of the significance of keywords and concepts in texts, revealing hidden meanings and influences, particularly in legal and social contexts, by objectively determining the impact of independent parameters on dependent variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000003_0001
    Figure IMGF000003_0001
  • Figure IMGF000010_0001
    Figure IMGF000010_0001
  • Figure IMGF000015_0001
    Figure IMGF000015_0001
Patent Text Reader

Abstract

The invention relates to text processing, and more particularly to automatic searching for conceptual parameters of a text. The invention involves identifying nouns and generating an array of parameters therefrom; identifying in the text all verbs and verbal parts of speech that reflect an action, and generating an array of observations therefrom; generating a value matrix; processing the value matrix data using a multilinear regression system, alternately considering the parameter values to be dependent variables and the observation values to be independent variables, and calculating the dependency of each of the parameters on the observations; ranking the dependency values identified for the parameters, and considering the parameters having the greatest identified dependency value to be conceptual parameters of the text. The technical result is that of achieving the intended purpose by means of the invention.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]METHOD FOR AUTOMATIC SEARCH OF CONTENT PARAMETERS OF TEXT 5 Field of the invention relates. The invention relates to text processing, specifically to the automatic search for content parameters of text. The following terms are used in this description: Content parameters of text – words that reflect the author's understanding in the official text, embodied in the ideas of the text and expressed in nouns and their forms. Parameter array – the set of all nouns in the text. Observation array – the set of all verbs and verbal parts of speech in the text. Verbal parts of speech are parts of speech that are formed from verbs and are used to describe features or states occurring as a result of the verb's action, such as verbal adjectives and past participles. 20 State of the Art The problem of automatic text analysis in terms of identifying the content parameters of the text has not been solved.To solve this problem, the present method for automatic search of substantive parameters of text is proposed. 25 Disclosure of the invention. The present invention, mainly, is aimed at proposing a method for automatic search of substantive parameters of text. The technical result achieved is the implementation of the stated purpose by the invention. To achieve this goal: 30 a. nouns are selected, from which an array of parameters is formed; b. all verbs and verbal parts of speech that reflect an action and from which an array of observations are formed are selected from the text; c. a matrix of values ​​is formed, the rows of which are formed from the array of observations, and the columns of which are formed from the array of parameters, wherein in each cell of the matrix "1" is put if the parameter corresponding to this cell is found in the text next to the observation, and if not, then "0"; 5 d.The data of the specified matrix are processed using a multiple linear regression system, alternately considering the values ​​for the parameters as the dependent variable, and the values ​​for the observations as the independent variables, calculating the dependence of each of the parameters on the observations, e. ranking the found values ​​of the dependencies for the parameters, and 10 considering the parameters with the highest value of the found dependence as the substantive parameters of the text. Due to these advantageous characteristics, it becomes possible to automatically search for the substantive parameters of the text. 15 Carrying out the invention. The method for automatically searching for the substantive parameters of the text is carried out as follows. (An example is given that does not limit the application of the invention.) 20 Example 1. The National Anthem of the United States. Originally from the text, officially approved by the US Senate on March 3, 1931 (see website: ru.m.wikipedia.org) a list of parameters is created, based on the nouns of the lyrics of the anthem, and a list of observations, based on the verbs of the lyrics of the US national anthem. 25 Step 1. Select the nouns from which the array is formed. 1. You (US citizen) 2. We (US citizens) 30 3. Flag (it) 4. Our proud answer 5. Enemy (gang of murderers) 6. Home of the brave (brave) 7. Land of the free 35 8. War (flame) 9. Grave 10. Rebel 11. Land 12. Strength 5 Synonyms reflecting the essence of the concepts are given in brackets. Step 2. Select from the text all the verbs and verbal parts of speech that reflect the action and from which an array of observations is formed. 10 1. To see 2. To walk amid the battle 3. To appear again 4. To appear again as red and white fire 5. To cast light 15 6. To answer vile enemies 7. To be forever, where the free stronghold is 8. To rest 9. To be in the foggy silence 10. To sway in the wind 20 11. To give shine 12. The breeze unfolds 13. To be always 14. To swear boastfully 15. Confusion of the fallen in spirit 25 16. They will make us again 17. To give an answer in blood 18. There is no refuge for you, mercenary troops 19. It will be according to deeds 20. Decay awaits 30 21. Never die 22. To get up 23. To praise God 24. To give praise 25. To make us now 35 26. Keep by the people 27.Without fear of fate 28. To be with the motto "True to God" 29. To soar for now 30. To be the one in whom freedom is alive. 5 Step 3. A matrix of values ​​is formed, the rows of which are formed from an array of observations, and the columns of which are formed from an array of parameters, while in each cell of the matrix "1" is put if the parameter corresponding to this cell occurs in the text next to the observation, and "0" if not. 10 Thus, nouns are parameters, and verbs are observations. As a result, an asymmetric matrix can be formed using binary codes. Thus, during digital transformation of text, the nouns of the text are essentially parameters of the statistical population, and the verbs of the text represent the corresponding 15 observations. The matrix is ​​formed as follows: if (1). "You are (a US citizen)" the parameter occurs with observation (2). “See”, then at the intersection of columns and rows we put the number “1”, and if this observation is not there, then we put “0”, and so on.20 The result of digitalizing the text using binary codes using this method is the matrix below – the digital form of the US anthem. Table 1. Accordingly, the resulting digital matrix of the lyrics to the US anthem can be processed using any statistical methods. 5 Among classical statistical methods, the first method for processing matrix data can be the application of cluster and factor analysis methods and / or multiple regression methods. Representing the official lyrics to the US anthem in the form of a digital model, firstly, allows us to overcome the emotional attitude towards the official document. 10 Secondly, it allows us to see that the content of the US anthem contains at least four substructures. And everything is “hung” on parameter (5) “enemy – gang of murderers”. Moreover, the factor weight of this parameter has a noticeable significance and a minus sign. Thus, the lyrics to the anthem, approved on March 3, 1931, which have not changed for 93 15 years, structurally and implicitly create an evaluative negative aura among US citizens in relation to external communities. Step 4.The data from this matrix is ​​processed using a multiple linear regression system, alternately considering the 20 parameter values ​​as the dependent variable and the observation values ​​as the independent variables. The relationship between each parameter and the observations is calculated. Multiple regression is an effective method for studying official texts. The method's essence lies in the fact that the entire set of substantive elements of the text—nouns—are parameters, each of which, one at a time, can serve as the dependent variable, while the others serve as independent variables. This calculates the relationship between each dependent text parameter and the other set of five parameters specified in the general list of parameters for the US national anthem. After extensive calculations, the resulting data from processing the text parameters were compiled into the table "Ranked Rows of Parameters for the US National Anthem Lyrics." Table 2. 10 Note to Table 2: - the first column indicates the ranked numbers of the parameters; - the second column is the names of the parameters, and the parameter number is indicated in brackets; - the third column "ABS" - records the sum of the absolute values ​​of the independent parameters that collectively influence the dependent variable - 5 parameter; - the fourth column, nominated as "R2", is the determination coefficient that reliably explains the share of fluctuations in the dependent parameter; - the fifth column first indicates the total positive share by 10 relative value, which positively determines the level of influence on the dependent variable in %, and indicates the numbers of the independent parameters that influence the dependent parameter with the "+" sign; - the sixth column first indicates the total negative share in %, which negatively influences the dependent variable in %, and indicates the numbers of 15 independent parameters with the "-" sign on the dependent parameter. Step 5.The found values ​​of dependencies for the parameters are ranked and the parameters with the highest value of the found dependency are considered the substantive parameters of the text. 20 According to the data in the table above, it is revealed that the substantive parameters of the text of the US anthem are five institutional parameters of the text of the anthem: (5) Enemy (gang of murderers); (12) Strength; 25 (4) Our proud answer; (2) We (citizens of the USA); (5) The flag of the USA. These substantive parameters of the text of the US anthem fully reflect the core of the social psychology of the country's society. 30 Not only does the presented method of automatic search for substantive parameters of the text allow us to accurately calculate the "weight" and hidden meaning of keywords and key concepts, but we can also determine the influence of other words and concepts as parameters on keywords and key parameters. For example, let us pay attention to the first key parameter 5. "enemy, gang of murderers". Only one parameter "7" has a positive effect on it.free country", and negatively as many as nine parameters. (2). We, citizens of the United States, 5 (3). Flag, (4). Our proud answer, (6). Home of the brave (brave), (8). War, (9). Grave, 10 (10). Rebel, (11). Land, (12). Strength. A concrete historical interpretation can be that the higher in meaning and weight the "enemy, (gang of murderers)", the freer, more significant 15 and more significant the country should be in a linear sense. And in practice, the entire semantics of the text of the US anthem is configured negatively against the "enemy, (gang of murderers)" with the following independent parameters: (2). We, citizens of the United States, (3). Flag, 20 (4). Our proud answer, (6). Home of the brave (brave), (8). War, (9). Grave, (10). Rebel, 25 (11). Earth, (12). Power. This hidden information, which we have just demonstrated to you, cannot be obtained in any way other than by analyzing the text using the proposed method.30 As a result, using multiple regression, when any parameter (Y) as keywords and / or concepts as a dependent variable, is determined objectively as the sum of the values ​​of the interrelated independent variables (X) of other parameters. Furthermore, the proposed method for automatically searching for the substantive parameters of a text as a "weight" determination—parameters—of keywords and individual concepts assumes the influence of both positive and negative signs of the parameters of the US national anthem's lyrics as independent variables. This allows us to discover the basis for the semantics of a particular dependent parameter of an individual test, when we are able to see which independent parameters have a positive and which negative effect on the dependent variable. Example 2. Federal Law. Another area of ​​application of the proposed method 10 for automatically searching for the substantive parameters of a text is the study of legislative texts of documents of various types.Using the automatic search method for substantive text parameters, we will study Federal Law No. 84-FZ of April 22, 2024, "On Amendments to Article No. 155 of the Housing Code of the Russian Federation." Article 15 of the Law is devoted to regulating payments in the housing and utilities sector. As in the previous text analysis example, we will extract parameters and observations from the text of Law No. 84-FZ, as in Example No. 1 above. Step 1: Identify nouns from which an array is formed. 20 1. Federal Law No. 84 2. Date 3. Lessor of the residential premises 4. Agreement on assignment of rights 5. Provisions of this part 25 6. New lessor 7. Tenant 8. Specified owner 9. New tenant of the residential premises 10. President of the Russian Federation 30 Stage 2. Select from the text all verbs and verbal parts of speech that reflect the action and from which an array of observations is formed. 1. To be 2. To be signed on 22.04.24 3. To be published on 26.04.24 4. Adopted by the State Duma on 10.04.24 5. Approved by the Federation Council on 17.04.24 6. To amend… 7. To pay a fee 5 8. To be in accordance with the Housing Code 9. To pay a fee for non-residential premises 10. Not entitled to assign the right (claim) to repay overdue debt of individuals… to third parties, including credit institutions. 10 11. To conclude with… 12. To be null and void 13. Not applicable to overdue debt 14. Obliged to notify the owner of the premises in writing 15. Obliged to notify the tenant of the residential premises under the contract 15 16. To have overdue debt under the contract 17.Notify of assignment of claim 18. Have the right not to fulfill obligations to repay overdue debt to newly selected organizations 19. Not to do anything until the assignment of claim to the new lessor 20 20. Comes into effect 21. Be a newly selected organization for the new lessor. Step 3. Form a matrix of values, the rows of which are formed from the array of observations, and the columns of which are formed from the array of parameters, with each cell of the matrix assigned a "1" if the parameter corresponding to this cell appears in the text next to the observation, and "0" if not. This is performed similarly; it is not provided to avoid confusion. The result is a 10 / 21 matrix, i.e. ten parameters per twenty-one observations – a dimension that allows for the necessary calculations based on ten parameters. This makes it possible to objectively determine the "weight" and, accordingly, the significance for semantic analysis of keywords and concepts (parameters).This method of working with complex legal documents is virtually the only accurate tool for identifying the top 35 text parameters—keywords and key concepts—in legal practice. Let's look at the table of ranked parameters for the text of Federal Law No. 84 below. Step 4. The data in this matrix is ​​processed using a multiple linear regression system, alternately considering the dependent variable to be the values ​​for the parameters, and the independent variables to be the values ​​for the observations. The dependence of each parameter on the observations is calculated. Table 3. 10 Note to Table 3: - the first column denotes the ranked numbers of parameters; - the second column is the names of the parameters, and the number of the parameter in the general list is indicated in brackets; 15 - the third column "ABS" - records the sum of the absolute values ​​of the independent parameters that collectively influence the dependent variable - the text parameter; - the fourth column, nominated as "R2", is the determination coefficient that reliably explains the proportion of fluctuations in the dependent parameter; - the fifth column first indicates the total positive proportion by 5 relative value, which positively determines the level of influence on the dependent variable in %, and indicates the numbers of the independent parameters that influence the dependent parameter of the text with the "+" sign; - the sixth column first indicates the total negative proportion in %, which negatively influences the dependent variable - the parameter - in %, and 10 indicates the numbers of the independent parameters of the text of the law with the "-" sign on the dependent parameter.In general, this method of automatically searching for substantive text parameters allows us to solve a crucial problem: objectively identifying keywords and key concepts from the entire text. 15 Furthermore, it is important that keywords and concepts are not simply accurately identified, but that, for example, influential words and concepts from the text parameters are presented, which influence a keyword or key concept with a plus or minus sign. 20 Step 5. The found dependency values ​​for the parameters are ranked, and the parameters with the highest value of the found dependency are considered substantive text parameters. For example, the parameter (3) "lessor of residential premises" in the text of the law is positively and predominantly influenced by parameter (2) "date" and parameter (4) "assignment agreement", while parameter (1) "FZ-84" negatively influences (3) "lessor of residential premises".The negative impact was determined because (1) Federal Law 84 significantly changed the previous relationships regarding the rental of residential premises. Thus, the data in the table above, obtained after 30 processing of the digital text of (1) Federal Law 84, show that the most significant parameters of the text of the law are the following: (3) the lessor of the residential premises; (1) Federal Law 84; and (2) the date of entry into force of Federal Law 84. These most important parameters of Federal Law 84 emphasize that the provisions of this law have significantly changed the law enforcement practice from the corresponding date. 35 Moreover, the processed data clearly and accurately show that two parameters are distinguished that most frequently affect the known parameters. These are (4) the agreement on the assignment of the right of claim and (7) the tenant. As a result, the method of automatic search for the substantive parameters of the text of Federal Law 84 allows us to instrumentally and reliably solve three 5 interrelated problems: 1.obtain a comparable structure of an individual text that reflects the most important text parameters as keywords and concepts. 2. determine the most significant substantive parameters of keywords and concepts in the text of the law for legal texts. 10 3. since individual texts of legislative acts can be processed using a unified digital methodology, it becomes possible to compare the content of incomparable texts and identify hidden information that cannot be detected without obtaining quantitative data on the keywords and concepts of the text parameters. 15 In other words, the proposed method for automatically searching for substantive text parameters may well help the entire corporation of lawyers to adequately understand documents, interpret them correctly, and make fair decisions. 20 Example 3. Text of the prosecutor's office It is important to provide an example of processing the text of the Submission of a certain interdistrict prosecutor's office.The text of the Submission itself refers to violations of labor laws in a specific workforce. The crux of the matter is that the company is required to conduct medical examinations of employees before the start of the workday. Moreover, these examinations are, in fact, conducted regularly. Nevertheless, the junior justice adviser who signed this Submission demands punishment for those guilty of violating labor laws. Which culprits? For what? Because medical examinations are conducted regularly. To determine what the junior justice adviser is actually seeking, we will conduct a search for the substantive parameters of the text, which requires digitalizing the text and processing it using mathematical and statistical algorithms. Step 1: Identify nouns, from which an array is formed. 1. Interdistrict prosecutor's office 2. Employees of enterprises in general 5 3. Regulatory legal acts of the Russian Federation 4. Time 5. Procedure for implementing the order of the Ministry of Health of the Russian Federation 6. Order of the Ministry of Health of the Russian Federation dated 05 / 30 / 23 7. Healthcare worker 10 8. Organization (enterprises) 9. Medical organization 10. Medical activity 11. Medical work 12. Employees of a specific enterprise 15 13. A specific enterprise 14. Violation(s) 15. Presentation 16. Specific medical examinations 17. Head of the enterprise 20 18. Medical license 19. Junior Justice Advisor - Prosecutor Stage 2. Select from the text all the verbs and verbal parts of speech that reflect the action and from which an array of observations is formed. 25 1. conduct an inspection 2. establish 3. be employed in hazardous working conditions 4. undergo medical examinations 5. undergo unscheduled medical examinations 30 6. include the time of examinations in working hours 7. carry out work (examinations) 8. have a medical license 9. include in the staff 10.have medical education 35 11. hire 12. conclude a contract 13. confirm with judicial practice 14. no medical license 15. no contract with a licensed medical organization 5 16. have improperly organized medical work 17. take measures 18. consider 19. notify 20. report 10 21. eliminate Step 3. Form a matrix of values, the rows of which are formed from an array of observations, and the columns of which are formed from an array of parameters, while in each cell of the matrix they put "1" if the parameter corresponding to this cell occurs in the text next to the observation, and if not, then "0". 15 This is done similarly, but is not given to avoid complicating perception. Step 4. The data in the specified matrix are processed using a multiple linear regression system, alternately considering the values ​​​​for the parameters as the dependent variable, and the values ​​​​for 20 observations as the independent variables, and the dependence of each of the parameters on the observations is calculated.After converting the text of the prosecutor's Submission into digital form and processing it, the following results were obtained, concentrated in the table below. Table 4 5 Note to Table 4: - the first column means the ranked numbers of parameters; - the second column is the names of the parameters, and the number of the parameter in the general list is indicated in brackets; - the third column, nominated as "R2", is the coefficient of 10 determination, which reliably explains the share of fluctuations in the dependent parameter; - the fourth column "ABS" - records the sum of the absolute values ​​of the independent parameters that collectively influence the dependent variable - the text parameter; - the fifth column first indicates the cumulative positive share by 5 relative value, which positively determines the level of influence on the dependent variable in %, and indicates the numbers of the independent parameters that influence the dependent parameter of the text with the "+" sign; - the sixth column first indicates the cumulative negative share in %, which negatively influences the dependent variable - the parameter - in %, and 10 indicates the numbers of the independent parameters of the text of the law with the "-" sign on the dependent parameter. Step 5.The found dependency values ​​for the parameters are ranked, and the parameters with the highest value of the found dependency are considered meaningful text parameters. 15 Text Processing. The submission of the interdistrict prosecutor's office demonstrated that, according to prosecutorial oversight requirements, medical examinations of workers should be performed not only by medical specialists on staff at the enterprise, but also in specialized medical institutions. Thus, it is clear that accurately determining the "weight" of 20 meaningful concepts provides significant assistance in making qualified decisions, especially in the area of ​​legal issues. Industrial Applicability. The proposed method for automatically searching for meaningful 25 text parameters can be implemented by a specialist in practice and, when implemented, ensures the fulfillment of the stated purpose, which allows us to conclude that the invention meets the "industrial applicability" criterion.Overall, the proposed method for automatically searching for meaningful text parameters significantly improves the accuracy of identifying keywords and concepts in texts of any complexity. Accurately calculating the "weight" and priority positions of keywords and concepts in texts of any complexity is, in essence, a measure of understanding the meaning and reliability of the phenomena, events, and processes being described. This invention can be used to analyze individual texts and significantly improves the accuracy of identifying keywords and concepts in texts of any complexity. Accurately calculating the "weight" and priority positions of keywords and concepts in texts of any complexity is, in essence, a measure of understanding the meaning and reliability of the phenomena, events, and processes being described.

Claims

CLAIMS 1. A method for automatically searching for substantive parameters of a text, which comprises a. selecting nouns from which an array of parameters is formed; b. selecting from the text all verbs and verbal parts of speech that reflect an action and from which an array of observations is formed; c. forming a matrix of values, the rows of which are formed from the array of observations, and the columns of which are formed from the array of parameters, wherein in each cell of the matrix "1" is put if the parameter corresponding to this cell occurs in the text next to the observation, and "0" if not; d. the data of the said matrix are processed using a multiple linear regression system, alternately considering the values ​​​​for the parameters to be the dependent variable, and the values ​​​​for the observations to be the independent variables, calculating the dependence of each of the parameters on the observations, for which the parameters "R2" (the coefficient of determination) are calculated, and also determining the level of positive "%+" and negative "%-" influence e.The found values ​​of dependencies for the parameters are ranked and the parameters with the highest value of the found dependency are considered as meaningful parameters of the text.

Citation Information

Patent Citations

  • System and methodology of automatic language learning on basis of syntactic models frequency

    RU2632656C2

  • Machine learning-based relationship association and related discovery and search engines

    US20190354544A1

  • System and method for automating the generation of an ontology from unstructured documents

    US7987088B2

  • Rapid automatic keyword extraction for information retrieval and analysis

    US8131735B2

  • Reinforcement learning-based emotional image description method and system

    WO2023155460A1