A hierarchical relationship analysis method and device for influencing factors

By calculating the correlation and influence degree in the first document data, a triangular fuzzy number matrix is ​​generated, and the hierarchical relationship of influencing factors is automatically determined, which solves the problem of time-consuming and inaccurate manual analysis, and realizes efficient and accurate data analysis.

CN119378533BActive Publication Date: 2025-08-19TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411531653.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-08-19
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

In the prior art, it consumes a lot of time and is not very accurate by manually analyzing the hierarchical relationship of influencing factors, making it difficult to effectively parse data in complex systems.

Method used

By determining the first influencing factors in the first document data of the target keyword information, calculating the correlation and influencing degrees, generating a triangular fuzzy number matrix, clarifying and decomposing the fuzzy matrix, and automatically determining the hierarchical relationship of the second influencing factors.

Benefits of technology

Automatic data analysis is realized, reducing time costs and improving the accuracy of data analysis, and enabling more accurate determination of hierarchical relationships between influencing factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119378533B_ABST
    Figure CN119378533B_ABST
Patent Text Reader

Abstract

The present application provides a hierarchical relationship analysis method and device for influencing factors, including: determining the first influencing factors for target keyword information contained in the first document data corresponding to the target keyword information, determining the relevance of the first document data to each first influencing factor, and determining the influence of the first document data on each first influencing factor; based on the relevance and influence, determining the second influencing factors corresponding to the target keyword information from each first influencing factor; based on the correlation score between each second influencing factor, determining the first fuzzy matrix corresponding to the second influencing factor; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each second influencing factor, respectively; for the target keyword information, based on the first fuzzy matrix, determining the hierarchical relationship between each second influencing factor, can reduce the time cost of data analysis to a certain extent and improve the accuracy of data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data analysis technology, and in particular to a method and device for analyzing hierarchical relationships based on influencing factors. Background Art

[0002] Understanding the complex interactions between influencing factors is crucial for real-world decision-making in fields such as environmental science, social science, and public policy. Accurately analyzing these mechanisms, particularly in areas such as global climate change, public health, and resource management, can significantly assist scientists and policymakers in predicting future trends and developing effective strategies and policies to address complex systemic issues such as human behavior. However, with the increasing interdisciplinary nature of research and the surge in scientific data, challenges remain in understanding the complexities of human behavior and decision-making, as well as in extracting and analyzing data from the extensive academic literature.

[0003] In related technologies, since academic literature contains rich data and information, especially empirical literature often has real data as input and support, it can better reflect the combination of practice and theory. After proper processing and analysis, this information can provide a deeper analysis of complex problems. Therefore, some literature data extraction and analysis tools can be used to obtain corresponding academic literature, and manually analyze the influencing factors from the academic literature to obtain the hierarchical relationship between the various influencing factors.

[0004] However, in the above methods, manual data analysis consumes a lot of time and cost, and the hierarchical relationship of influencing factors obtained through manual analysis often has errors, and the accuracy of manual analysis is not high. Summary of the Invention

[0005] In view of the above problems, embodiments of the present application provide a hierarchical relationship analysis method, device, electronic device and readable storage medium for influencing factors, so as to overcome the above problems or at least partially solve the above problems.

[0006] In a first aspect, an embodiment of the present application provides a hierarchical relationship analysis method for influencing factors, the method comprising:

[0007] Determining a first influencing factor for the target keyword information contained in the first document data corresponding to the target keyword information;

[0008] Determining the relevance of the first document data to each of the first influencing factors, and determining the influence of the first document data on each of the first influencing factors; wherein the relevance represents the degree of association between the first document data and each of the first influencing factors; and the influence represents the degree of influence of each of the first influencing factors on the target keyword information;

[0009] Based on the relevance and the influence, determining a second influencing factor corresponding to the target keyword information from each of the first influencing factors;

[0010] Determining a first fuzzy matrix corresponding to each of the second influencing factors based on the correlation scores between the second influencing factors; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each of the second influencing factors;

[0011] For the target keyword information, based on the first fuzzy matrix, a hierarchical relationship between each of the second influencing factors is determined.

[0012] Optionally, determining the first fuzzy matrix corresponding to the second influencing factors based on the correlation scores between the second influencing factors includes:

[0013] Determine the triangular fuzzy numbers corresponding to the correlation scores between the respective second influencing factors;

[0014] Based on the triangular fuzzy number, a first fuzzy matrix corresponding to the second influencing factor is generated.

[0015] Optionally, determining the hierarchical relationship between the second influencing factors based on the first fuzzy matrix for the target keyword information includes:

[0016] Based on the relationship between the element value of each element in the first fuzzy matrix and the first preset threshold, the first fuzzy matrix is clarified to obtain a first reachable matrix;

[0017] Decomposing the first reachable matrix to obtain a first reachable set and a first predecessor set;

[0018] For the target keyword information, based on the first reachable set and the first precedent set, a hierarchical relationship between each of the second influencing factors is determined.

[0019] Optionally, the elements in the first fuzzy matrix are triangular fuzzy numbers, and the method further includes:

[0020] The element value of each element in the first fuzzy matrix is determined based on the lower limit value, the most likely value and the upper limit value respectively corresponding to each triangular fuzzy number in the first fuzzy matrix.

[0021] Optionally, determining the relevance of the first document data to each of the first influencing factors includes:

[0022] Performing semantic recognition on each first document in the first document data to obtain first keyword information corresponding to each first document;

[0023] Determining the similarity between the first keyword information and each of the first influencing factors;

[0024] Filtering the first documents to obtain second documents containing the first keyword information corresponding to a similarity greater than or equal to a second preset threshold value;

[0025] Based on the first number of the second documents associated with each of the first influencing factors, the relevance of the first document data to each of the first influencing factors is determined.

[0026] Optionally, determining the influence of the first document data on each of the first influencing factors includes:

[0027] performing semantic recognition on each of the second documents to obtain first context information corresponding to each of the second documents;

[0028] For each of the first influencing factors, determining from the second documents a third document corresponding to the first context information containing a preset sensitive word;

[0029] Based on the second number of the third documents, the influence of the first document data on each of the first influencing factors is determined.

[0030] Optionally, for each of the first influencing factors, determining from the second documents a third document corresponding to the first context information containing a preset sensitive word includes:

[0031] For each of the first influencing factors, determining, from the first context information, second context information within a preset range surrounding the first keyword information in the second document;

[0032] In a case where the second context information includes a preset sensitive word, the second document corresponding to the second context information is determined to be the third document corresponding to the first context information.

[0033] Optionally, determining a second influencing factor corresponding to the target keyword information from each of the first influencing factors based on the relevance and the influence includes:

[0034] Determine, from each of the first influencing factors, a third influencing factor whose influence is greater than or equal to a third preset threshold;

[0035] Determine, from each of the first influencing factors, a fourth influencing factor whose influence is less than the third preset threshold and whose correlation is greater than or equal to a fourth preset threshold;

[0036] The third influencing factor and the fourth influencing factor are integrated to obtain a second influencing factor corresponding to the target keyword information.

[0037] Optionally, the method further includes:

[0038] Based on the hierarchical relationship, a directed graph of hierarchical relationships between each of the second influencing factors and the target keyword information is generated.

[0039] In a second aspect, an embodiment of the present application provides a hierarchical relationship analysis device for influencing factors, the device comprising:

[0040] A first determining module, configured to determine a first influencing factor for the target keyword information contained in the first document data corresponding to the target keyword information;

[0041] a second determining module, configured to determine a relevance of the first document data to each of the first influencing factors, and determine an influence of the first document data on each of the first influencing factors; wherein the relevance indicates a degree of association between the first document data and each of the first influencing factors; and the influence indicates a degree of influence of each of the first influencing factors on the target keyword information;

[0042] a third determining module, configured to determine, based on the relevance and the influence, a second influencing factor corresponding to the target keyword information from each of the first influencing factors;

[0043] a fourth determining module, configured to determine a first fuzzy matrix corresponding to each of the second influencing factors based on the correlation scores between the second influencing factors; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each of the second influencing factors;

[0044] A fifth determining module is configured to determine, for the target keyword information, a hierarchical relationship between each of the second influencing factors based on the first fuzzy matrix.

[0045] Optionally, the fourth determining module includes:

[0046] A first determining submodule is configured to determine the triangular fuzzy numbers corresponding to the correlation scores between the respective second influencing factors;

[0047] A generating submodule is used to generate a first fuzzy matrix corresponding to the second influencing factor based on the triangular fuzzy number.

[0048] Optionally, the fifth determining module includes:

[0049] a clarifying submodule, configured to clarify the first fuzzy matrix based on a magnitude relationship between an element value of each element in the first fuzzy matrix and a first preset threshold value, to obtain a first reachable matrix;

[0050] a decomposition submodule, configured to decompose the first reachable matrix to obtain a first reachable set and a first predecessor set;

[0051] The second determining submodule is configured to determine, for the target keyword information, a hierarchical relationship between each of the second influencing factors based on the first reachable set and the first preceding set.

[0052] Optionally, the elements in the first fuzzy matrix are triangular fuzzy numbers, and the apparatus further includes:

[0053] The sixth determining module is configured to determine an element value of each element in the first fuzzy matrix based on a lower limit value, a most probable value, and an upper limit value respectively corresponding to each triangular fuzzy number in the first fuzzy matrix.

[0054] Optionally, the second determining module includes:

[0055] A first identification submodule is configured to perform semantic recognition on each first document in the first document data to obtain first keyword information corresponding to each first document;

[0056] a third determining submodule, configured to determine the similarity between the first keyword information and each of the first influencing factors;

[0057] A fourth determining submodule is configured to filter out from the first document a second document containing the first keyword information corresponding to a similarity greater than or equal to a second preset threshold;

[0058] The fifth determining submodule is configured to determine the relevance of the first document data to each of the first influencing factors based on the first quantity of the second documents associated with each of the first influencing factors.

[0059] Optionally, the second determining module includes:

[0060] A second recognition submodule is configured to perform semantic recognition on each of the second documents to obtain first context information corresponding to each of the second documents;

[0061] a sixth determining submodule, configured to determine, for each of the first influencing factors, from the second documents a third document corresponding to the first context information containing a preset sensitive word;

[0062] The seventh determining submodule is configured to determine the influence of the first document data on each of the first influencing factors based on the second number of the third documents.

[0063] Optionally, the sixth determining submodule includes:

[0064] a second determining unit configured to determine, from the first context information, second context information within a preset range surrounding the first keyword information in the second document for each of the first influencing factors;

[0065] The third determining unit is configured to determine, when the second context information includes a preset sensitive word, the second document corresponding to the second context information as the third document corresponding to the first context information.

[0066] Optionally, the third determining module includes:

[0067] an eighth determining submodule, configured to determine, from each of the first influencing factors, a third influencing factor whose influence is greater than or equal to a third preset threshold;

[0068] a ninth determining submodule, configured to determine, from among the first influencing factors, a fourth influencing factor whose influence is less than the third preset threshold and whose correlation is greater than or equal to a fourth preset threshold;

[0069] The integration submodule is configured to integrate the third influencing factor and the fourth influencing factor to obtain a second influencing factor corresponding to the target keyword information.

[0070] Optionally, the device further comprises:

[0071] A generating module is configured to generate a directed graph of hierarchical relationships between each of the second influencing factors and the target keyword information based on the hierarchical relationship.

[0072] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the hierarchical relationship analysis method for influencing factors as described in any one of the above items.

[0073] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the hierarchical relationship analysis method for influencing factors as described in any one of the above items is implemented.

[0074] The specific beneficial effects are:

[0075] The embodiment of the present application determines the first influencing factors for the target keyword information contained in the first document data corresponding to the target keyword information, determines the relevance of the first document data to each first influencing factor, and determines the influence of the first document data on each first influencing factor. Based on the relevance and influence, the second influencing factors corresponding to the target keyword information are determined from each first influencing factor. Based on the correlation scores between each second influencing factor, the first fuzzy matrix corresponding to the second influencing factor is determined; wherein, the rows and columns of the first fuzzy matrix correspond one-to-one to each second influencing factor respectively. For the target keyword information, based on the first fuzzy matrix, the hierarchical relationship between each second influencing factor is determined. The first influencing factor corresponding to the target keyword can be found from the first document data corresponding to the target keyword information, and the first influencing factor can be screened by the relevance and influence to obtain the second influencing factor. Finally, the hierarchical relationship between each second influencing factor is obtained through the fuzzy matrix of the second influencing factor. The document data corresponding to the target keyword information can be automatically analyzed, thereby reducing the time cost of data analysis to a certain extent and improving the accuracy of data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0077] Figure 1 This is a flow chart of a hierarchical relationship analysis method for influencing factors provided in an embodiment of the present application;

[0078] Figure 2 1 is a flow chart of another hierarchical relationship analysis method for influencing factors provided in an embodiment of the present application;

[0079] Figure 3 is a schematic diagram of a reachability matrix provided in an embodiment of the present application;

[0080] Figure 4 is a schematic diagram of a hierarchical directed graph provided in an embodiment of the present application;

[0081] Figure 5 This is a logic block diagram of a hierarchical relationship analysis device for influencing factors provided by an embodiment of the present application;

[0082] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0083] The exemplary embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0084] Reference Figure 1 , Figure 1 A flowchart of a hierarchical relationship analysis method for influencing factors provided in an embodiment of the present application is provided, wherein the method includes:

[0085] Step 101: Determine a first influencing factor for the target keyword information contained in first document data corresponding to the target keyword information.

[0086] In an embodiment of the present application, the target keyword information may be the central vocabulary around which the content of the document revolves. For example, the target keyword information may be research content in various fields such as water-saving behavior, water-using behavior, electricity-using behavior, travel frequency, and speech generation. The first document data corresponding to the target keyword information may refer to the paper document data containing the target keyword information in the subject or abstract. The first document data may be segmented to obtain other vocabulary information in addition to the target keyword information in the first document data. The above-mentioned other vocabulary information is the first influencing factor for the target keyword information contained in the first document data.

[0087] For example, the target keyword information may be water-saving behavior, and the first document data may be relevant research papers on water-saving behavior. The first influencing factors contained in the first document data may include but are not limited to the degree of part-time employment, planting scale, technical perception, technical training, land endowment, incentive policy, cultural level, water resource endowment, family income, age, family population, etc.

[0088] Step 102: Determine the relevance of the first document data to each of the first influencing factors, and determine the influence of the first document data on each of the first influencing factors; wherein the relevance represents the degree of association between the first document data and each of the first influencing factors; and the influence represents the degree of influence of each of the first influencing factors on the target keyword information.

[0089] In the embodiments of the present application, the relevance may refer to the degree of association between the first document data and each first influencing factor, which may be calculated by the frequency of occurrence of the first influencing factor in the first document data. For example, the frequency of occurrence of the first influencing factor may be directly used as the relevance of the first document data to each first influencing factor, or the frequency of occurrence of the first influencing factor may be normalized using a normalization function, and the normalized value may be used as the relevance of the first document data to each first influencing factor. The influence is an evaluation index that measures the comprehensive influence of a certain factor on other factors in the system. In the embodiments of the present application, the influence may represent the influence of each first influencing factor on the target keyword information. Generally, the corresponding influence may be calculated based on the survey data or experimental data of each paper document contained in the first document data. For example, in a control experiment, the relative rate of change between the experimental data of the experimental group and the experimental data of the control group may be used as the influence of the variable condition on the control experiment; for each first influencing factor, the relative difference between the survey data or experimental data containing the first influencing factor and the survey data or experimental data not containing the first influencing factor may be used as the influence of the first document data on the first influencing factor.

[0090] Step 103: Determine a second influencing factor corresponding to the target keyword information from each of the first influencing factors based on the relevance and the influence.

[0091] In an embodiment of the present application, after obtaining the relevance and influence of the first document data to each of the first influencing factors, the second influencing factor corresponding to the target keyword information can be screened from each of the first influencing factors based on the above relevance and influence. Among them, the relevance and influence can be used as conditions for screening independently, or they can be combined to form a joint condition for screening the first influencing factor. For example, the first influencing factor with a relevance greater than or equal to the first threshold can be used as the second influencing factor, and the first influencing factor with an influence greater than or equal to the second threshold can be used as the second influencing factor; the first influencing factor with a relevance greater than or equal to the first threshold and an influence greater than or equal to the second threshold can also be used as the second influencing factor.

[0092] Step 104: Determine a first fuzzy matrix corresponding to each of the second influencing factors based on the correlation scores between the second influencing factors; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each of the second influencing factors.

[0093] In an embodiment of the present application, the relevance score can be an evaluation index of the relevance between each second influencing factor, and can be obtained by scoring by a group of experts in the field described by the target keyword information. The relevance score can have a preset score set or score interval. A fuzzy matrix is a matrix used to represent fuzzy relationships. If set X has m elements and set Y has n elements, the fuzzy relationship from set X to set Y can be represented by a matrix. The elements of the fuzzy matrix have values between [0, 1], where 0 represents complete irrelevance and 1 represents complete correlation. The value of each element in the fuzzy matrix can be determined based on the mapping between the score set or score interval of the relevance score and the element value interval of the fuzzy matrix.

[0094] For example, if the score set of the relevance score is {1, 2, 3, 4, 5}, then the element value set corresponding to the score set can be {0.2, 0.4, 0.6, 0.8, 1}, where a relevance score of 1 can correspond to an element value of 0.2, a relevance score of 2 can correspond to an element value of 0.4, a relevance score of 3 can correspond to an element value of 0.6, a relevance score of 4 can correspond to an element value of 0.8, and a relevance score of 5 can correspond to an element value of 1.

[0095] Step 105 : For the target keyword information, based on the first fuzzy matrix, determine the hierarchical relationship between each of the second influencing factors.

[0096] In an embodiment of the present application, the hierarchical relationship can represent the level of influence of each second influencing factor on the target keyword information, and the hierarchical relationship between each second influencing factor can be determined based on the first fuzzy matrix. The hierarchical relationship and the target keyword information can correspond to each other. For example, the level of the second influencing factor corresponding to each column can be determined based on the sum of the element values in each column of the first fuzzy matrix. The larger the sum of the element values, the smaller the level of the second influencing factor, indicating that the second influencing factor has a stronger relevance to the target keyword information.

[0097] In an embodiment of the present application, by determining the first influencing factors for the target keyword information contained in the first document data corresponding to the target keyword information, determining the relevance of the first document data to each first influencing factor, and determining the influence of the first document data on each first influencing factor, based on the relevance and influence, determining the second influencing factors corresponding to the target keyword information from each first influencing factor, and based on the correlation scores between each second influencing factor, determining the first fuzzy matrix corresponding to the second influencing factor; wherein, the rows and columns of the first fuzzy matrix respectively correspond to each second influencing factor one-to-one, for the target keyword information, based on the first fuzzy matrix, the hierarchical relationship between each second influencing factor is determined, the first influencing factor corresponding to the target keyword can be found from the first document data corresponding to the target keyword information, and the first influencing factor can be screened by relevance and influence to obtain the second influencing factor, and finally the hierarchical relationship between each second influencing factor is obtained through the fuzzy matrix of the second influencing factor, and the document data corresponding to the target keyword information can be automatically analyzed, thereby reducing the time cost of data analysis to a certain extent and improving the accuracy of data analysis.

[0098] Reference Figure 2 , Figure 2 A flowchart of a method for analyzing and determining hierarchical relationships of influencing factors provided in an embodiment of the present application may include:

[0099] Step 201: Determine a first influencing factor for the target keyword information contained in first document data corresponding to the target keyword information.

[0100] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 101 and will not be repeated here.

[0101] Step 202: Perform semantic recognition on each first document in the first document data to obtain first keyword information corresponding to each first document.

[0102] In an embodiment of the present application, the first document data may include multiple first documents, and a word segmenter or a common semantic recognition model (such as a BERT model) may be used to perform semantic recognition on each first document, thereby obtaining the first keyword information corresponding to each first document.

[0103] Step 203: Determine the similarity between the first keyword information and each of the first influencing factors.

[0104] In an embodiment of the present application, the similarity between the first keyword information and each first influencing factor can be calculated. For example, the Pearson Correlation Coefficient between the first keyword information and each first influencing factor can be used as the similarity, or the Euclidean distance can be used as the similarity. The vector used in calculating the similarity can be obtained by vectorizing the first keyword information and each first influencing factor using embedding.

[0105] Step 204 : Filter the first documents to obtain second documents containing the first keyword information corresponding to a similarity greater than or equal to a second preset threshold.

[0106] In an embodiment of the present application, a second document containing first keyword information corresponding to a similarity greater than or equal to a second preset threshold value can be screened out from the first document, that is, the first document containing first keyword information corresponding to a similarity greater than or equal to the second preset threshold value is determined, and this part of the first document is used as the second document.

[0107] Step 205 : Determine the relevance of the first document data to each of the first influencing factors based on the first number of the second documents associated with each of the first influencing factors.

[0108] In the embodiments of the present application, as can be seen from the aforementioned embodiments, each first influencing factor can have an association relationship with the first document. Therefore, a second document can also have an association relationship with each first influencing factor. Therefore, the relevance of the first document data to each first influencing factor can be determined based on the first number of second documents associated with each first influencing factor. The specific details of the relevance can be found in the embodiments of step 102 and will not be further described here.

[0109] Step 206: Perform semantic recognition on each of the second documents to obtain first context information corresponding to each of the second documents.

[0110] In an embodiment of the present application, semantic recognition can be performed on each second document to obtain the first context information corresponding to each second document. For example, the first context information of each second document can be obtained using a common improved Transformer-based model (such as a BERT model) or a common large language model.

[0111] Step 207 : For each of the first influencing factors, determine from the second documents a third document corresponding to the first context information containing a preset sensitive word.

[0112] In an embodiment of the present application, for each first influencing factor, a third document corresponding to the first context information containing a preset sensitive word can be screened from the second document to determine. Obviously, for each first influencing factor, there can be a corresponding third document. Among them, the preset sensitive word can be a word indicating that the first influencing factor has a strong correlation with the target keyword information, such as "significantly positive", "significantly negative", "relatively strong", "extremely strong", etc.

[0113] Optionally, step 207 may include the following sub-steps:

[0114] Sub-step 2071 : For each of the first influencing factors, determine second context information within a preset range surrounding the first keyword information in the second document from the first context information.

[0115] In an embodiment of the present application, for each first influencing factor, second context information within a preset range surrounding the first keyword information in the second document can be determined from the first context information. The preset range surrounding the first keyword information in the second document can refer to a query range with the first keyword information as the query center and a preset number of bytes as the query radius. Thus, the process of determining the second context information can be the query result obtained by querying the first context information based on the aforementioned query range.

[0116] Sub-step 2072: When the second context information contains a preset sensitive word, the second document corresponding to the second context information is determined as the third document corresponding to the first context information.

[0117] In an embodiment of the present application, when the second context information includes a preset sensitive word, the second document corresponding to the second context information may be determined as the third document corresponding to the first context information.

[0118] Step 208: Determine the influence of the first document data on each of the first influencing factors based on the second number of the third documents.

[0119] In the embodiment of the present application, the influence of the first document data on each first influencing factor can be determined based on the second number of the third document. For details on the influence, please refer to the embodiment of step 102 and will not be repeated here.

[0120] Step 209 : determining a second influencing factor corresponding to the target keyword information from each of the first influencing factors based on the relevance and the influence.

[0121] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 103 and will not be repeated here.

[0122] Optionally, step 209 may include the following sub-steps:

[0123] Sub-step 2091: Determine, from each of the first influencing factors, a third influencing factor whose influence is greater than or equal to a third preset threshold.

[0124] In the embodiment of the present application, third influencing factors having an influence greater than or equal to a third preset threshold may be screened and determined from among the first influencing factors according to the magnitude of the influence.

[0125] Sub-step 2092: Determine, from each of the first influencing factors, a fourth influencing factor whose influence is less than the third preset threshold and whose correlation is greater than or equal to a fourth preset threshold.

[0126] In an embodiment of the present application, third influencing factors having an influence less than the third preset threshold and a correlation greater than or equal to the fourth preset threshold can be screened and determined from the various first influencing factors based on the influence and the correlation.

[0127] Sub-step 2093: integrating the third influencing factor and the fourth influencing factor to obtain a second influencing factor corresponding to the target keyword information.

[0128] In an embodiment of the present application, the third influencing factor and the fourth influencing factor may be integrated, so that the result of the integration may be used as the second influencing factor corresponding to the target keyword information.

[0129] Step 210: Determine a first fuzzy matrix corresponding to each of the second influencing factors based on the correlation scores between the second influencing factors; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each of the second influencing factors.

[0130] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 104 and will not be repeated here.

[0131] Optionally, step 210 may include the following sub-steps:

[0132] Sub-step 2101: determining the triangular fuzzy numbers corresponding to the correlation scores between the respective second influencing factors.

[0133] In the embodiments of the present application, triangular fuzzy numbers are a mathematical tool for representing uncertain information. Triangular fuzzy numbers are composed of three parameters: a lower limit value (l), a most likely value (m), and an upper limit value (r). These three values together define a triangular membership function, which describes the membership of a variable within a given range, that is, the degree to which the variable belongs to a certain set. The correlation scores between the various second influencing factors can be determined based on the pre-established correspondence between the correlation scores and the triangular fuzzy numbers. A correspondence between a correlation score and a triangular fuzzy number can be shown in Table 1 below:

[0134] Table 1

[0135] Relevance score Corresponding triangular fuzzy number TFN(l,m,r) 1 (0.0,0.1,0.3) 2 (0.1,0.3,0.5) 3 (0.3,0.5,0.7) 4 (0.5,0.7,0.9) 5 (0.7,0.9,1.0)

[0136] Sub-step 2102: generating a first fuzzy matrix corresponding to the second influencing factor based on the triangular fuzzy number.

[0137] In an embodiment of the present application, a first fuzzy matrix corresponding to the second influencing factors can be generated based on the triangular fuzzy numbers. The triangular fuzzy numbers can be directly used as element values of the first fuzzy matrix, or the most likely values in the triangular fuzzy numbers can be used as element values of the first fuzzy matrix.

[0138] Step 211 : For the target keyword information, based on the first fuzzy matrix, determine the hierarchical relationship between each of the second influencing factors.

[0139] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 105 and will not be repeated here.

[0140] Optionally, step 211 may include the following sub-steps:

[0141] Sub-step 2111 : determining the element value of each element in the first fuzzy matrix based on the lower limit value, the most likely value, and the upper limit value corresponding to each triangular fuzzy number in the first fuzzy matrix.

[0142] In the embodiment of the present application, the element value of each element in the first fuzzy matrix can be calculated and determined according to the lower limit value, most likely value and upper limit value corresponding to each triangular fuzzy number in the first fuzzy matrix. The calculation method is as follows:

[0143] As shown in formula 3:

[0144]

[0145]

[0146] In the above formulas 1 to 3, n ij represents the element value of row i and column j in the first fuzzy matrix, represents the upper limit value of the triangular fuzzy number in row i and column j, represents the most likely value among the triangular fuzzy numbers in row i and column j, represents the lower limit of the triangular fuzzy number in row i and column j. The remaining variables are intermediate parameters and are not explained here.

[0147] Sub-step 2112: Based on the magnitude relationship between the element value of each element in the first fuzzy matrix and a first preset threshold, the first fuzzy matrix is clarified to obtain a first reachable matrix.

[0148] In an embodiment of the present application, a reachable matrix is a matrix form used to describe the degree to which nodes in a directed connection graph can be reached after a path of a certain length, and the element values of the reachable matrix can be 0 or 1. In a directed graph, if there is a path from vertex Si to vertex Sj, Si is said to be reachable to Sj, and accordingly, the value of element Sij is 1. The first fuzzy matrix can be clarified by the relationship between the element values of each element in the first fuzzy matrix and the first preset threshold, thereby obtaining a first reachable matrix. If the element value is greater than or equal to the first preset threshold, the element corresponding to the element value is set to 1; if the element value is less than the first preset threshold, the element corresponding to the element value is set to 0.

[0149] For example, refer to Figure 3 , Figure 3 A schematic diagram of a reachable matrix provided in an embodiment of the present application. In the figure, S0 represents the target keyword information, S1 to S10 represent the second influencing factors, a total of 10. The elements on the diagonal of the matrix are all defaulted to 1 and have no practical meaning. For element S ij For example, if S ij =1, it means that there is a connected path from Si to Sj between the influencing factor Si and the influencing factor Sj. ij =0, it means that there is no connected path between the influencing factor Si and the influencing factor Sj.

[0150] Sub-step 2113: Decompose the first reachable matrix to obtain a first reachable set and a first predecessor set.

[0151] In the embodiment of the present application, the first reachable matrix can be decomposed to obtain the first reachable set and the first predecessor set. The reachable set contains all other influencing factors that can be reached from the influencing factor i, which is expressed as R(i)={S j |S ij =1}, the predecessor set includes all other influencing factors that can reach influencing factor i, expressed as A(i)={S j |S ji=1}. According to the properties of the reachable set, the first reachable matrix is decomposed row by row to obtain the first reachable set corresponding to each second influencing factor; according to the properties of the predecessor set, the first reachable matrix is decomposed column by column to obtain the first predecessor set corresponding to each second influencing factor.

[0152] Continuing with the above example, for influencing factor S5, its reachable set is {S1, S3, S6, S10}, and its predecessor set is {S4, S8}.

[0153] Sub-step 2114 , for the target keyword information, based on the first reachable set and the first predecessor set, determining the hierarchical relationship between each of the second influencing factors.

[0154] In an embodiment of the present application, with respect to target keyword information, the hierarchical relationship between the second influencing factors may be determined based on the first reachable set and the first precedent set corresponding to each second influencing factor.

[0155] Continuing with the above example, for influencing factor S5, S5 can point to S1, S3, S6, and S10, while S4 and S8 can also point to S5. This gives us the hierarchical relationship between influencing factor S5 and the other influencing factors. By traversing all influencing factors, we can obtain the hierarchical relationship between each secondary influencing factor.

[0156] Step 212: Based on the hierarchical relationship, generate a directed graph of the hierarchical relationship between each of the second influencing factors and the target keyword information.

[0157] In an embodiment of the present application, a directed hierarchical relationship graph between each second influencing factor and target keyword information may be generated based on the hierarchical relationship between each second influencing factor. When generating the directed hierarchical relationship graph, the cross-influence between each second influencing factor may be considered.

[0158] Using the above example, refer to Figure 4 , Figure 4 A schematic diagram of a hierarchical directed graph provided in an embodiment of the present application.

[0159] In the figure, if S0 is water-saving behavior, S1 is the degree of concurrent employment, S2 is the scale of planting, S3 is land endowment, S4 is age, S5 is educational level, S6 is technology perception, S7 is incentive policy, S8 is family income, S9 is water resource endowment, and S10 is technical training, then the generated hierarchical relationship directed graph is Figure 4As shown in the form. For the second influencing factor S5, there is one path pointing to S5, namely the path of "age → family income → education level". Corresponding to the influencing factors S4 and S8, there are four paths pointing to other influencing factors S1, S3, S6 and S10, namely the path of "education level → land endowment → degree of part-time employment", the path of "education level → land endowment → planting scale", the path of "education level → land endowment → technology perception" and the path of "education level → land endowment → technology training". In the path of "age → family income → education level", since S4 has directed connected paths to all other influencing factors, and S8 has directed connected paths to all other influencing factors except S4, it can be determined that the directional relationship is "S4 points to S8". For other influencing factors, the above analysis method for S5 can be referred to for analysis, and will not be repeated here. For the target keyword S0, there is only a path pointing to S0, and there is no path pointed to by S0. Therefore, S0 is at the end of the directional path of the hierarchical relationship directed graph.

[0160] Reference Figure 5 , Figure 5 This is a logic block diagram of a hierarchical relationship analysis device for influencing factors provided in an embodiment of the present application. The hierarchical relationship analysis device 500 for influencing factors may include:

[0161] A first determining module 501 is configured to determine a first influencing factor for the target keyword information contained in the first document data corresponding to the target keyword information;

[0162] A second determining module 502 is configured to determine the relevance of the first document data to each of the first influencing factors, and determine the influence of the first document data on each of the first influencing factors;

[0163] A third determining module 503 is configured to determine a second influencing factor corresponding to the target keyword information from each of the first influencing factors based on the relevance and the influence;

[0164] A fourth determining module 504 is configured to determine a first fuzzy matrix corresponding to each of the second influencing factors based on the correlation scores between the second influencing factors; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each of the second influencing factors;

[0165] The fifth determining module 505 is configured to determine, for the target keyword information, a hierarchical relationship between each of the second influencing factors based on the first fuzzy matrix.

[0166] Optionally, the fourth determining module 504 includes:

[0167] A first determining submodule is configured to determine the triangular fuzzy numbers corresponding to the correlation scores between the respective second influencing factors;

[0168] A generating submodule is used to generate a first fuzzy matrix corresponding to the second influencing factor based on the triangular fuzzy number.

[0169] Optionally, the fifth determining module 505 includes:

[0170] a clarifying submodule, configured to clarify the first fuzzy matrix based on a magnitude relationship between an element value of each element in the first fuzzy matrix and a first preset threshold value, to obtain a first reachable matrix;

[0171] a decomposition submodule, configured to decompose the first reachable matrix to obtain a first reachable set and a first predecessor set;

[0172] The second determining submodule is configured to determine, for the target keyword information, a hierarchical relationship between each of the second influencing factors based on the first reachable set and the first preceding set.

[0173] Optionally, the elements in the first fuzzy matrix are triangular fuzzy numbers, and the hierarchical relationship analysis device 500 for influencing factors further includes:

[0174] The sixth determining module is configured to determine an element value of each element in the first fuzzy matrix based on a lower limit value, a most probable value, and an upper limit value respectively corresponding to each triangular fuzzy number in the first fuzzy matrix.

[0175] Optionally, the second determining module 502 includes:

[0176] A first identification submodule is configured to perform semantic recognition on each first document in the first document data to obtain first keyword information corresponding to each first document;

[0177] a third determining submodule, configured to determine the similarity between the first keyword information and each of the first influencing factors;

[0178] A fourth determining submodule is configured to filter out from the first document a second document containing the first keyword information corresponding to a similarity greater than or equal to a second preset threshold;

[0179] The fifth determining submodule is configured to determine the relevance of the first document data to each of the first influencing factors based on the first quantity of the second documents associated with each of the first influencing factors.

[0180] Optionally, the second determining module 502 includes:

[0181] A second recognition submodule is configured to perform semantic recognition on each of the second documents to obtain first context information corresponding to each of the second documents;

[0182] a sixth determining submodule, configured to determine, for each of the first influencing factors, from the second documents a third document corresponding to the first context information containing a preset sensitive word;

[0183] The seventh determining submodule is configured to determine the influence of the first document data on each of the first influencing factors based on the second number of the third documents.

[0184] Optionally, the sixth determining submodule includes:

[0185] a second determining unit configured to determine, from the first context information, second context information within a preset range surrounding the first keyword information in the second document for each of the first influencing factors;

[0186] The third determining unit is configured to determine, when the second context information includes a preset sensitive word, the second document corresponding to the second context information as the third document corresponding to the first context information.

[0187] Optionally, the third determining module 503 includes:

[0188] an eighth determining submodule, configured to determine, from each of the first influencing factors, a third influencing factor whose influence is greater than or equal to a third preset threshold;

[0189] a ninth determining submodule, configured to determine, from among the first influencing factors, a fourth influencing factor whose influence is less than the third preset threshold and whose correlation is greater than or equal to a fourth preset threshold;

[0190] The integration submodule is configured to integrate the third influencing factor and the fourth influencing factor to obtain a second influencing factor corresponding to the target keyword information.

[0191] Optionally, the hierarchical relationship analysis device 500 for influencing factors further includes:

[0192] A generating module is configured to generate a directed graph of hierarchical relationships between each of the second influencing factors and the target keyword information based on the hierarchical relationship.

[0193] The hierarchical relationship analysis device for influencing factors in the embodiment of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or a device other than a terminal. Exemplarily, the electronic device can be a GPUBOX, a mobile phone, a tablet computer, a laptop computer, a PDA, a car-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., and the embodiment of the present application is not specifically limited.

[0194] The hierarchical relationship analysis device for influencing factors in the embodiment of the present application can be a device having an operating system. The operating system can be an Android operating system, a Linux operating system, a Windows operating system, etc., or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0195] The hierarchical relationship analysis device for influencing factors provided by the embodiment of the present application can achieve Figures 1 to 3 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0196] The present application provides an electronic device. Figure 6 The electronic device 60 includes: a processor 601, a memory 602, and a computer program 6021 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the program, the hierarchical relationship analysis method for influencing factors of the aforementioned embodiment is implemented.

[0197] An embodiment of the present application also provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps in the hierarchical relationship analysis method for influencing factors disclosed in the embodiment of the present application are implemented.

[0198] The embodiment of the present application further provides a computer program product, which, when executed on an electronic device, enables a processor to implement the steps of the hierarchical relationship analysis method for influencing factors disclosed in the embodiment of the present application.

[0199] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0200] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0201] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0202] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0203] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0204] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0205] The above is a detailed introduction to the hierarchical relationship analysis method for influencing factors provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.

Claims

1. A hierarchical relationship analysis method for influencing factors, characterized by: The method comprises: Determining a first influencing factor for the target keyword information contained in the first document data corresponding to the target keyword information; Determining the relevance of the first document data to each of the first influencing factors, and determining the influence of the first document data on each of the first influencing factors; wherein the relevance represents the degree of association between the first document data and each of the first influencing factors; and the influence represents the degree of influence of each of the first influencing factors on the target keyword information; Based on the relevance and the influence, determining a second influencing factor corresponding to the target keyword information from each of the first influencing factors; Determining a first fuzzy matrix corresponding to each of the second influencing factors based on the correlation scores between the second influencing factors; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each of the second influencing factors; For the target keyword information, determining the hierarchical relationship between each of the second influencing factors based on the first fuzzy matrix; The step of determining the first fuzzy matrix corresponding to the second influencing factors based on the correlation scores between the second influencing factors includes: Determine the triangular fuzzy numbers corresponding to the correlation scores between the respective second influencing factors; Based on the triangular fuzzy number, a first fuzzy matrix corresponding to the second influencing factor is generated.

2. The method according to claim 1, characterized in that The step of determining the hierarchical relationship between the second influencing factors based on the first fuzzy matrix for the target keyword information includes: Based on the relationship between the element value of each element in the first fuzzy matrix and the first preset threshold, the first fuzzy matrix is clarified to obtain a first reachable matrix; Decomposing the first reachable matrix to obtain a first reachable set and a first predecessor set; For the target keyword information, based on the first reachable set and the first precedent set, a hierarchical relationship between each of the second influencing factors is determined.

3. The method according to claim 2, characterized in that The elements in the first fuzzy matrix are triangular fuzzy numbers, and the method further includes: The element value of each element in the first fuzzy matrix is determined based on the lower limit value, the most likely value and the upper limit value respectively corresponding to each triangular fuzzy number in the first fuzzy matrix.

4. The method according to claim 1, wherein Determining the relevance of the first document data to each of the first influencing factors includes: Performing semantic recognition on each first document in the first document data to obtain first keyword information corresponding to each first document; Determining the similarity between the first keyword information and each of the first influencing factors; Filtering the first documents to obtain second documents containing the first keyword information corresponding to a similarity greater than or equal to a second preset threshold value; Based on the first number of the second documents associated with each of the first influencing factors, the relevance of the first document data to each of the first influencing factors is determined.

5. The method according to claim 4, characterized in that Determining the influence of the first document data on each of the first influencing factors includes: performing semantic recognition on each of the second documents to obtain first context information corresponding to each of the second documents; For each of the first influencing factors, determining from the second documents a third document corresponding to the first context information containing a preset sensitive word; Based on the second number of the third documents, the influence of the first document data on each of the first influencing factors is determined.

6. The method according to claim 5, characterized in that The determining, for each of the first influencing factors, from the second documents, a third document corresponding to the first context information containing a preset sensitive word includes: For each of the first influencing factors, determining, from the first context information, second context information within a preset range surrounding the first keyword information in the second document; In a case where the second context information includes a preset sensitive word, the second document corresponding to the second context information is determined to be the third document corresponding to the first context information.

7. The method according to claim 1, characterized in that The determining, based on the relevance and the influence, a second influencing factor corresponding to the target keyword information from each of the first influencing factors includes: Determine, from each of the first influencing factors, a third influencing factor whose influence is greater than or equal to a third preset threshold; Determine, from each of the first influencing factors, a fourth influencing factor whose influence is less than the third preset threshold and whose correlation is greater than or equal to a fourth preset threshold; The third influencing factor and the fourth influencing factor are integrated to obtain a second influencing factor corresponding to the target keyword information.

8. The method according to claim 1, characterized in that The method further comprises: Based on the hierarchical relationship, a directed graph of hierarchical relationships between each of the second influencing factors and the target keyword information is generated.

9. A hierarchical relationship analysis device for influencing factors, characterized in that: The device comprises: A first determining module, configured to determine a first influencing factor for the target keyword information contained in the first document data corresponding to the target keyword information; a second determining module, configured to determine a relevance of the first document data to each of the first influencing factors, and determine an influence of the first document data on each of the first influencing factors; wherein the relevance indicates a degree of association between the first document data and each of the first influencing factors; and the influence indicates a degree of influence of each of the first influencing factors on the target keyword information; a third determining module, configured to determine, based on the relevance and the influence, a second influencing factor corresponding to the target keyword information from each of the first influencing factors; a fourth determining module, configured to determine a first fuzzy matrix corresponding to each of the second influencing factors based on the correlation scores between the second influencing factors; wherein the rows and columns of the first fuzzy matrix correspond one-to-one to each of the second influencing factors; a fifth determining module, configured to determine, for the target keyword information, a hierarchical relationship between each of the second influencing factors based on the first fuzzy matrix; Wherein, the fourth determining module includes: A first determining submodule is configured to determine the triangular fuzzy numbers corresponding to the correlation scores between the respective second influencing factors; A generating submodule is used to generate a first fuzzy matrix corresponding to the second influencing factor based on the triangular fuzzy number.

Citation Information

Patent Citations

  • Keyword extraction method and system and storage medium

    CN110598209A

  • Literature information pushing method, device and system and storage medium

    CN117527888A