A method, device, computer equipment and medium for determining watershed business relations

By obtaining the text and data set of river basin business data, using word frequency and reverse file frequency methods and clustering analysis for business identification, building a causal graph and traversing the causal graph using a preset particle diffusion model, the problems of business relationship discovery and new business opportunities mining in river basin management are solved, and efficient resource optimization and management improvement are achieved.

CN119477098BActive Publication Date: 2025-05-23THREE GORGES GROUP IND DEVELOPMENT (BEIJING) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510039741.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-23
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing methods and tools cannot meet the needs of business relationship discovery and new business opportunity mining in watershed management, rely on expert experience and difficult to make full use of internal enterprise data.

Method used

By obtaining the text and data set of river basin business data, using word frequency and reverse file frequency methods and clustering analysis for business identification, constructing a causal graph and traversing the causal graph using a preset particle diffusion model to determine the river basin business relationship.

Benefits of technology

It has realized the systematic sorting and in-depth exploration of the causal relationship between complex factors in the basin information, discovered hidden business connections, optimized resource allocation, improved management efficiency, and improved economic, social and ecological benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477098B_ABST
    Figure CN119477098B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and discloses a method, device, computer equipment and medium for determining watershed business relationships. The present invention uses word frequency and reverse file frequency methods and cluster analysis methods to perform business identification, and can intelligently and quickly extract a valuable list of watershed business activities from a large amount of text data, avoiding the inefficiency and subjectivity of traditional manual identification methods. Furthermore, a causal graph is constructed based on the data set and the activity list, and a preset particle diffusion model is used to traverse the causal graph to obtain a causal list of watershed business, thereby achieving a systematic combing and in-depth mining of the causal relationship between complex factors in watershed information, and can discover business connections hidden behind the data. Finally, the watershed business relationship is determined in combination with the actual watershed business process, so that the discovered business relationship can be closely integrated with the actual business, thereby facilitating the targeted discovery of new business opportunities in watershed management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, computer equipment and medium for determining a watershed business relationship. Background Art

[0002] In the field of modern enterprises and management, discovering new business opportunities is crucial for sustainable development. Although traditional new business opportunity discovery methods such as SWOT, PEST, 5W1H, etc. can provide a macro analysis framework, these methods rely heavily on expert experience for subjective judgment. This approach of relying on expert experience has obvious limitations. On the one hand, the decision-making speed is slow because experts need to spend a lot of time collecting information, analyzing and discussing; on the other hand, it is difficult to make full use of the large amount of data and knowledge accumulated within the enterprise, resulting in these valuable resources not being effectively mined and applied in the business discovery process.

[0003] At the same time, some existing business analysis intelligent tools, such as Google Analytics, mainly focus on understanding customer behavior analysis, and their design and functional positioning are oriented to the business market and customer relationship management. In the specific and highly professional field of watershed information management, these tools cannot meet the needs of business relationship discovery and new business opportunity mining in watershed management due to their lack of targeted processing capabilities for watershed characteristics, business processes and related data. Summary of the invention

[0004] In view of this, the present invention provides a method, apparatus, computer equipment and medium for determining watershed business relationships to solve the problem that existing methods and tools cannot meet the needs of business relationship discovery and new business opportunity mining in watershed management.

[0005] In a first aspect, the present invention provides a method for determining a watershed service relationship, the method comprising:

[0006] Acquire the watershed business information text, watershed business activity data set and actual watershed business process; based on the watershed business information text, use the word frequency and inverse file frequency method and cluster analysis method to identify the business and obtain the watershed business activity list; based on the watershed business activity data set and the watershed business activity list, construct a cause-effect graph; based on the watershed business activity list and the watershed business activity data set, use the preset particle diffusion model to traverse the cause-effect graph and obtain the watershed business cause-effect linked list; based on the watershed business cause-effect linked list and the actual watershed business process, determine the watershed business relationship.

[0007] The method for determining watershed business relationships provided by the present invention uses word frequency and reverse file frequency methods and cluster analysis methods to identify business, and can intelligently and quickly extract a valuable list of watershed business activities from a large amount of text data, avoiding the inefficiency and subjectivity of traditional manual identification methods. Furthermore, a causal graph is constructed based on the data set and the activity list, and a preset particle diffusion model is used to traverse the causal graph to obtain a causal list of watershed business, thereby achieving a systematic combing and in-depth mining of the causal relationship between complex factors in watershed information, and can discover business connections hidden behind the data. Finally, the watershed business relationship is determined in combination with the actual watershed business process, so that the discovered business relationship can be closely integrated with the actual business, thereby helping to optimize resource allocation, improve management efficiency, and discover new business opportunities in watershed management in a targeted manner, thereby improving the economic, social and ecological benefits of watershed management.

[0008] In an optional implementation, based on the watershed business data text, word frequency and reverse document frequency methods and cluster analysis methods are used to perform business identification to obtain a list of watershed business activities, including:

[0009] The text of the basin business information is segmented to obtain multiple initial entries; based on the preset entry list, multiple entry vectors of the multiple initial entries are constructed; based on the multiple entry vectors, multiple target entries are obtained through word frequency and reverse file frequency methods; cluster analysis is performed on the multiple target entries and a list of basin business activities is determined.

[0010] The method for determining watershed business relationships provided by the present invention performs word segmentation processing on the text of watershed business information, and can convert continuous text into analyzable entry units. Furthermore, an entry vector is constructed based on a preset entry list, so that each entry can be represented in a unified vector space. Furthermore, by processing using word frequency and inverse file frequency methods, the importance of each entry in the text collection can be effectively measured, and key target entries that are closely related to watershed business can be screened out, reducing the interference of data noise and redundant information. Finally, by performing cluster analysis on multiple target entries, similar entries can be grouped together, thereby clearly dividing different areas of watershed business activities and improving the accuracy and efficiency of business identification.

[0011] In an optional implementation, a cause-effect diagram is constructed based on a watershed business activity data set and a watershed business activity list, including:

[0012] Based on the list of watershed business activities, the random forest algorithm is used to perform feature selection on the watershed business activity data set to obtain multiple key data of watershed business; a causal diagram is constructed based on multiple key data of watershed business.

[0013] The method for determining the relationship between river basin business provided by the present invention performs feature selection on the river basin business activity data set through the random forest algorithm, and can screen out key data with high relevance to river basin business activities from a large amount of data, which not only reduces the dimension of the data and the computational complexity of subsequent analysis, but also can highlight the information that has an important impact on the river basin business activities, avoids the interference of irrelevant or redundant information on the analysis results, and improves the efficiency and accuracy of data processing. Furthermore, by using multiple key data of river basin business to construct a causal graph, the causal graph can be made more concise and clear, and more accurately reflect the causal relationship in the river basin business activities.

[0014] In an optional implementation, a cause-effect graph is constructed based on multiple key data of watershed services, including:

[0015] Based on multiple key data of watershed business, multiple correlation values ​​are obtained through processing with Kendall rank correlation coefficient and partial correlation analysis method; based on multiple key data of watershed business and multiple correlation values, a cause-effect diagram is constructed.

[0016] The method for determining the relationship between river basin business provided by the present invention is based on multiple river basin business key data, and multiple correlation values ​​are obtained through processing by the Kendall rank correlation coefficient and partial correlation analysis method, which can accurately quantify the linear and nonlinear correlations between the key data. Furthermore, a causal graph is constructed by combining multiple river basin business key data and the obtained multiple correlation values, and the existence and direction of the causal relationship can be determined based on the real correlation between the data, so that the constructed causal graph is more scientific and reasonable and in line with the actual situation.

[0017] In an optional implementation, based on the watershed business activity list and the watershed business activity data set, a preset particle diffusion model is used to traverse the causal graph to obtain a watershed business causal linked list, including:

[0018] Based on the watershed business activity list and the watershed business activity data set, a mapping relationship table is determined; the causal graph is traversed using a preset particle diffusion model to obtain a node causal relationship table; based on the node causal relationship table and the mapping relationship table, a watershed business causal chain table is determined.

[0019] The method for determining the watershed business relationship provided by the present invention determines a mapping relationship table based on the watershed business activity list and the watershed business activity data set, establishes a connection bridge between business activities and data, and can effectively map the causal relationship at the data level with the activities at the business level. Further, the node causal relationship table is obtained by traversing the causal graph using a preset particle diffusion model, which reveals in detail the causal relationship details between each node in the causal graph, providing a basis for in-depth understanding of the microstructure of watershed business relationships. Finally, the watershed business causal chain table is determined based on the node causal relationship table and the mapping relationship table, and the causal relationship at the data level is integrated with the business activities to form a complete business causal relationship chain, which clearly demonstrates the causal logic between watershed business activities, and provides comprehensive and accurate information support for watershed management decisions such as discovering new business opportunities, optimizing business processes, and improving resource utilization efficiency.

[0020] In an optional implementation, determining the watershed service relationship based on the watershed service causal chain table and the actual watershed service process includes:

[0021] Perform time series analysis on the causal chain list of watershed business to obtain the first impact set; decompose the actual watershed business process to obtain the second impact set; determine the watershed business relationship based on the first impact set and the second impact set.

[0022] The method for determining the watershed business relationship provided by the present invention can capture the changing rules and influencing factors of the watershed business relationship in the time dimension by performing time series analysis on the watershed business causal chain table, and reveals the dynamic characteristics of the business relationship. At the same time, by decomposing the actual watershed business process, the inherent structure and influencing factors in the existing business process are deeply analyzed, and the operating mechanism and key links of the current business are clarified. Finally, by comparing and integrating the results of the time series analysis with the results of the actual watershed business process decomposition, it is possible to discover business relationships and potential new business opportunities that are not fully recognized in the existing business process, which provides direction for the innovation and optimization of watershed management, helps to improve the economic, social and ecological benefits of watershed management, and achieve the sustainable utilization of watershed resources and the continuous improvement of management level.

[0023] In a second aspect, the present invention provides a device for determining a watershed service relationship, the device comprising:

[0024] An acquisition module is used to acquire the text of watershed business information, the data set of watershed business activities and the actual watershed business process; an identification module is used to identify the business based on the watershed business information text using the word frequency and reverse file frequency method and the cluster analysis method to obtain a list of watershed business activities; a construction module is used to construct a causal graph based on the watershed business activity data set and the watershed business activity list; a traversal module is used to traverse the causal graph based on the watershed business activity list and the watershed business activity data set using a preset particle diffusion model to obtain a causal linked list of watershed business; a determination module is used to determine the watershed business relationship based on the watershed business causal linked list and the actual watershed business process.

[0025] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the method for determining the watershed business relationship of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0026] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for determining watershed business relationships of the first aspect or any corresponding embodiment thereof.

[0027] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the method for determining watershed business relationships according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0029] Figure 1 is a flow chart of a method for determining a watershed service relationship according to an embodiment of the present invention;

[0030] Figure 2 is a flow chart of another method for determining a watershed service relationship according to an embodiment of the present invention;

[0031] Figure 3 is a flow chart of another method for determining a flow domain service relationship according to an embodiment of the present invention;

[0032] Figure 4 is a structural block diagram of a device for determining a watershed service relationship according to an embodiment of the present invention;

[0033] Figure 5 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0035] An embodiment of the present invention provides a method for determining watershed business relationships, which identifies businesses through word frequency and inverse file frequency methods and clustering analysis methods, and uses a particle diffusion model to traverse a constructed causal graph to achieve a systematic combing and in-depth mining of the causal relationships between complex factors in watershed information, and to discover the business connections hidden behind the data.

[0036] According to an embodiment of the present invention, an embodiment of a method for determining a watershed business relationship is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] In this embodiment, a method for determining a watershed service relationship is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 1 is a flow chart of a method for determining a watershed service relationship according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0038] Step S101, obtaining watershed business data text, watershed business activity data set and actual watershed business process.

[0039] Among them, the text of watershed business information covers textual information on various aspects of watershed management, operation and related business development. It may include a basic description of the watershed (such as watershed area, flow area, topographical features, etc.), construction status of various water conservancy projects (project name, construction time, scale, function, etc.), water resources monitoring data records and analysis (water level, flow, water quality parameters, etc. changes over time and corresponding interpretations), the development of watershed ecological protection work (changes in animal and plant species, implementation of ecological restoration projects, etc.), flood control and drought relief measures and historical event records in the watershed, as well as administrative management regulations related to watershed business, business process descriptions, and other content.

[0040] Furthermore, the watershed business activity dataset represents a collection of data related to the actual development of watershed business, which can be collected through existing systems within the organization such as ERP, supply chain management system, financial system, contract system and other business systems as well as the organization's internal database.

[0041] Furthermore, the actual watershed business process represents the process description of the order, rules and methods in which various business activities are actually operated and circulated in the real scenario of watershed management, which can be obtained through UML modeling.

[0042] Step S102, based on the watershed business data text, business identification is performed using word frequency and reverse document frequency methods and cluster analysis methods to obtain a list of watershed business activities.

[0043] Among them, term frequency–inverse document frequency (TF-IDF) represents a commonly used weighting technique for information retrieval and data mining.

[0044] Term Frequency (TF) is used to measure how frequently a word appears in a specific text; Inverse Document Frequency (IDF) is used to measure the rarity or uniqueness of a word in the entire text collection.

[0045] Furthermore, the clustering analysis method may be a K-Means clustering algorithm, a hierarchical clustering algorithm, or the like.

[0046] Specifically, the watershed business data text can be segmented, and the important words in the watershed business data text that are more valuable and representative for distinguishing different business contents can be screened out through word frequency and reverse file frequency methods.

[0047] Furthermore, through cluster analysis, a corresponding list of basin business activities can be generated.

[0048] Step S103: construct a cause-effect diagram based on the watershed business activity data set and the watershed business activity list.

[0049] Among them, the causal graph and particle diffusion model represent a graphical model used to describe the causal relationship between variables, which can show how the various factors in the system influence and interact with each other in an intuitive graphical way.

[0050] Furthermore, in watershed business activities, causal diagrams can use nodes to represent various variables in watershed business activities, such as natural factors (rainfall, water level, etc.), human activity factors (water resource exploitation, sewage treatment, etc.), and use directed edges to represent the causal relationship between variables, that is, how changes in one variable cause changes in another variable.

[0051] Step S104, based on the watershed business activity list and the watershed business activity data set, the causal graph is traversed using a preset particle diffusion model to obtain a watershed business causal linked list.

[0052] Among them, the preset particle diffusion model represents a dynamic model based on probability and state transition, which can abstract objects or variables in the system into particles, and describe the dynamic changes and evolution laws of the system by simulating the random diffusion process of particles between different states.

[0053] Furthermore, in watershed business activities, the preset particle diffusion model can regard the different values ​​or states of various variables in watershed business activities as different states of particles. By defining the transition probability and rules of particles between these states, it can simulate and analyze the changing trends and mutual influence relationships of variables in watershed business activities over time.

[0054] The watershed business causal linked list represents a data structure used to present the causal relationship between various factors in the watershed business activities. The variables or factors involved in different business activities in the watershed are arranged in order according to the causal relationship. Each element (node) in the linked list represents a business-related factor, and clearly reflects the sequence of "cause" and "effect" between the factors and the mutual influence relationship. Through this linked list, it is possible to clearly sort out which factors' changes will trigger corresponding changes in other factors in the watershed business activity scenario, thereby helping managers, researchers, etc. to deeply understand the internal operating mechanism of the watershed business system and provide a strong basis for decision-making, business optimization, etc.

[0055] Specifically, the preset particle diffusion model is used to traverse the causal graph to obtain the basin business causal chain list, which realizes the systematic sorting and in-depth mining of the causal relationship between complex factors in the basin information, and can discover the business connections hidden behind the data.

[0056] Step S105, determining the watershed business relationship based on the watershed business causal chain table and the actual watershed business process.

[0057] Specifically, the watershed business relationships are determined in combination with the actual watershed business processes, so that the discovered business relationships can be closely integrated with the actual business, which will help to optimize resource allocation, improve management efficiency, and discover new business opportunities in watershed management, thereby enhancing the economic, social and ecological benefits of watershed management.

[0058] The method for determining the watershed business relationship provided in this embodiment uses the word frequency and reverse file frequency method and the cluster analysis method for business identification, which can intelligently and quickly extract a valuable list of watershed business activities from a large amount of text data, avoiding the inefficiency and subjectivity of traditional manual identification methods. Furthermore, a causal graph and a particle diffusion model are constructed based on the data set and the activity list, and the particle diffusion model is used to traverse the causal graph to obtain a causal list of watershed business, thereby achieving a systematic combing and in-depth mining of the causal relationship between complex factors in watershed information, and discovering the business connections hidden behind the data. Finally, the watershed business relationship is determined in combination with the actual watershed business process, so that the discovered business relationship can be closely integrated with the actual business, thereby helping to optimize resource allocation, improve management efficiency, and discover new business opportunities in watershed management in a targeted manner, thereby improving the economic, social and ecological benefits of watershed management.

[0059] In this embodiment, a method for determining a watershed service relationship is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 2 is a flow chart of a method for determining a watershed service relationship according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0060] Step S201, obtain the watershed business data text, watershed business activity data set and actual watershed business process. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0061] Step S202, based on the watershed business data text, business identification is performed using word frequency and reverse document frequency methods and cluster analysis methods to obtain a list of watershed business activities.

[0062] Specifically, the above step S202 includes:

[0063] Step S2021, segment the watershed business data text to obtain multiple initial entries.

[0064] Specifically, you can choose a suitable word segmentation tool based on the language characteristics of the basin business data text (such as Chinese or English, etc.). For Chinese text, you can use a professional Chinese word segmentation tool such as Jieba, which can split continuous text into independent word units based on Chinese vocabulary, grammar and other rules; if it is English text, you can use spaces and some common punctuation marks (such as commas, periods, etc.) as separators to split the text into word-shaped entries.

[0065] Furthermore, the selected word segmentation tool is used to segment the watershed business data text and obtain multiple initial entries. For example, for the Chinese text "the flood control and drought relief work in the watershed requires the coordination and cooperation of multiple parties, and the construction of water conservancy projects must also be carried out in an orderly manner", after word segmentation, initial entries such as "watershed", "within", "of", "flood control and drought relief", "work", "need", "multiple parties", "cooperation", "cooperation", "water conservancy project", "construction", "also", "need", "orderly", and "promote" will be obtained.

[0066] Step S2022: construct multiple term vectors of multiple initial terms based on a preset term list.

[0067] Among them, the preset entry list represents a pre-set collection of various possible words related to watershed business, which can be determined by sorting out professional knowledge in the field of watershed business, summarizing high-frequency words in previous related research or projects, etc.

[0068] Specifically, for each initial term, a corresponding term vector is constructed. Furthermore, the dimension of the term vector is the same as the dimension of the preset term list, that is, the length of the term vector is equal to the number of terms in the preset term list.

[0069] Furthermore, the preset term list can be traversed to check the number of times each term in the list appears in the text where the initial term whose vector is to be constructed is located, and the number of occurrences is used as the value of the corresponding position of the term vector.

[0070] For example, the preset term list has 5 terms, namely "flood control", "water resources", "ecology", "engineering", and "management". For the initial term "flood control and drought relief work", after statistics, "flood control" appears 1 time in the text, "water resources" appears 0 times, "ecology" appears 0 times, "engineering" appears 0 times, and "management" appears 0 times, so the constructed term vector is [1,0,0,0,0].

[0071] Step S2023, based on the multiple term vectors, multiple target terms are obtained by processing using the term frequency and inverse document frequency methods.

[0072] Among them, the target entries represent important words in the text of watershed business materials that are more valuable and representative for distinguishing different business contents, and are more closely related to the key activities of the watershed business.

[0073] Specifically, for the initial term corresponding to each term vector, its word frequency in all basin business data texts is counted, and then the word frequency value of the initial term can be calculated (word frequency value = the sum of the number of times the initial term appears in each text divided by the total number of words in all texts).

[0074] Furthermore, the number of texts containing each initial term is counted, and then the IDF value of each initial term can be calculated using the Inverse Document Frequency (IDF) method.

[0075] Furthermore, the term frequency (TF) value and the inverse document frequency (IDF) value of each initial term are multiplied to obtain the TF-IDF value.

[0076] Furthermore, the initial terms may be sorted in descending order of TF-IDF values ​​and initial terms with higher TF-IDF values ​​may be selected as target terms.

[0077] Step S2024, cluster analysis is performed on multiple target terms and a list of basin business activities is determined.

[0078] Specifically, you can choose a suitable clustering algorithm based on data characteristics and business needs. For example, if you have a rough expectation of the number of clusters in the clustering results and hope that the algorithm is relatively simple and efficient, you can consider using the K-Means clustering algorithm; if you want to present the hierarchical structure of clusters more intuitively, you can choose the hierarchical clustering algorithm.

[0079] Furthermore, the target terms are used as input data for cluster analysis and the selected clustering algorithm is run.

[0080] Taking the K-Means algorithm as an example, first you need to determine the number of clusters K (you can determine the appropriate K value by testing different K values ​​multiple times, combining business understanding and clustering effect evaluation indicators such as the silhouette coefficient), then randomly initialize K cluster centers, and then calculate the distance from each target term to these cluster centers (the distance measurement method can be Euclidean distance, etc.), assign the target term to the cluster with the nearest cluster center, and then recalculate the cluster center of each cluster. Repeat the above process of assigning and updating cluster centers until the cluster center no longer changes significantly or reaches the preset number of iterations.

[0081] Through cluster analysis, similar target terms are classified into the same category, and target terms in different categories are quite different.

[0082] Furthermore, by analyzing and summarizing each type of target term after clustering, and combining the actual background and professional knowledge of the watershed business, a business activity name is given to each type of target term, so that a corresponding list of watershed business activities can be formed.

[0083] For example, if one category of target terms includes "flood prevention", "flood control", "flood fighting", etc., the corresponding business activity name of this category can be defined as "flood prevention and disaster reduction business"; if another category includes "water resources monitoring", "water quality testing", "water volume statistics" and other terms, the corresponding business activity name can be "water resources comprehensive management business" and so on.

[0084] In this way, we can eventually obtain a list covering multiple business activities, clearly presenting the different main business activity areas in the basin. At the same time, we can further determine the secondary classification, i.e. the main business activities, within the basin based on the clustering results (such as the size of the K value), thus providing a basis for more detailed sorting out of basin business relationships.

[0085] Furthermore, after obtaining the list of business activities, you can also use the Unified Modeling Language (UML) for modeling to map the basic activity domain of watershed information in a more intuitive and standardized way. For example, use case diagrams can be used to show the relationship between different business activities and user roles, and class diagrams can be used to present the class relationship between business activities. This can further sort out and present the overall architecture of the watershed business and the relationship between the business activities, assisting in the in-depth understanding of the watershed business and subsequent analysis operations.

[0086] Step S203: construct a cause-effect diagram based on the watershed business activity data set and the watershed business activity list.

[0087] Specifically, the above step S203 includes:

[0088] Step S2031, based on the watershed business activity list, use the random forest algorithm to perform feature selection on the watershed business activity data set to obtain multiple watershed business key data.

[0089] Among them, the Random Forest Algorithm represents an integrated learning algorithm.

[0090] Specifically, feature selection can be performed through the following steps:

[0091] (1) Data preprocessing: This may include cleaning data, handling missing values, encoding categorical variables, such as one-hot encoding, and standardizing or normalizing data.

[0092] (2) Training the Random Forest Model

[0093] First, the parameters of the random forest can be reasonably determined based on factors such as the scale of the watershed business activity dataset, the number of features, and the characteristics of the business problem.

[0094] Secondly, when constructing each decision tree, a feature subset is randomly selected from all preprocessed features according to the set parameters.

[0095] Then, for each tree, use a randomly selected feature subset to recursively split the data based on the feature values ​​of the data and the corresponding target values ​​(in supervised learning scenarios, if it is unsupervised clustering, it is based on rules such as data similarity). For example, for regression problems, select the feature value that minimizes the mean square error as the split node; for classification problems, select the feature value that reduces the node's impurity (such as the Gini coefficient) the most for splitting, and continue this process until the stopping condition is met. The stopping condition can be reaching the pre-set maximum depth of the tree (such as the value set previously), or the number of samples contained in the node is less than a certain minimum number of samples (such as 5 samples), etc., and through continuous segmentation, complete decision trees are constructed.

[0096] Finally, if it is a classification problem, such as predicting the risk level (high, medium, or low risk) of flooding in a certain area within a river basin, after all decision trees are constructed and predictions are made on new data samples, the final category is determined by majority voting, that is, whichever risk level is predicted by the majority of decision trees is used as the final classification result; if it is a regression problem, such as predicting the water level change value within a river basin, then the average of the prediction results of all decision trees is calculated, and this average is used as the final water level prediction value. Through this aggregation method, a comprehensive and relatively stable prediction result is obtained.

[0097] (3) Feature selection

[0098] First, after the random forest model is trained, an importance score can be provided for each feature based on the set evaluation criteria. For classification problems, the Gini impurity reduction is often used as a measurement standard, that is, the importance of each feature is evaluated by calculating the contribution of each feature to reducing the Gini impurity of the node in all decision trees; for regression problems, the mean square error reduction is used to measure, that is, to see how much each feature contributes to reducing the mean square error of the prediction result.

[0099] Secondly, all features involved in the training can be sorted in descending order of importance based on the feature importance scores provided by the random forest model.

[0100] Finally, the top N most important features can be selected based on the importance score. The value of N can be determined based on actual business needs and experience. It can be a fixed number, such as selecting the top 5 or top 10 features; or it can be dynamically selected based on some criteria, such as selecting features with importance scores exceeding a certain threshold (such as 0.2).

[0101] (4) Building a new model: You can use the selected feature subset to retrain a new random forest model according to the previous random forest model training steps. Furthermore, the new random forest model is learned and trained only based on the selected key features. The purpose is to further evaluate the actual impact of these features on model performance and see whether it is possible to streamline and optimize the features while maintaining or improving model performance, so that the model can focus more on the key factors that have a core influence on watershed business activities.

[0102] (5) Evaluate model performance

[0103] Specifically, the performance of the new model can be evaluated through cross-validation (such as the commonly used K-fold cross-validation, which divides the dataset into K parts, uses K-1 parts as training sets each time, and 1 part as a test set, repeats K times and takes the average performance indicator) or an independent test set (reserving a part of the original data, which does not participate in model training and is specifically used for the final test of model performance).

[0104] Furthermore, the performance of the new model is compared with that of the full-feature model (i.e., the random forest model trained using all the original features). If these performance indicators of the new model do not decline significantly, or even improve in some cases, then it can be considered that the feature selection is successful, and the selected key features can indeed effectively represent the core information in the basin's business activities and help the model to better predict and analyze; conversely, if the performance declines significantly, it may be necessary to reconsider the feature selection strategy, such as appropriately increasing the number of selected features, or trying other feature selection methods, and then performing the above feature selection and model building evaluation process again to continuously optimize the feature subset and model performance.

[0105] Among them, the performance indicators can be accuracy, recall, F1 value, etc. in classification problems, or mean square error, mean absolute error, etc. in regression problems.

[0106] Finally, through the above process, key data of multiple watershed businesses can be obtained.

[0107] Step S2032, constructing a cause-effect diagram based on key business data of multiple watersheds.

[0108] Specifically, a corresponding cause-effect diagram can be constructed by combining the key business data of multiple watersheds.

[0109] In some optional implementations, the above step S2032 includes:

[0110] Step a1, based on multiple key watershed business data, multiple correlation values ​​are obtained through processing using the Kendall rank correlation coefficient and partial correlation analysis method.

[0111] Among them, the Kendall rank correlation coefficient represents a statistic used to measure the correlation between two ordered variables; the partial correlation analysis method represents a statistical method for studying the net correlation between two variables while controlling the influence of other variables, which can more accurately understand the direct relationship between the two variables and eliminate the interference of other variables.

[0112] Specifically, for each pair of variables in the basin business activity dataset (for example, variable A is water level change, variable B is water resource allocation), the Kendall rank correlation coefficient method can be used to measure the correlation between them:

[0113] First, the data of each variable are sorted separately to obtain their respective rank sequences, and then the number of synergistic pairs (logarithms of the two variables with the same rank change direction) and non-synergistic pairs (logarithms of the rank change directions) are counted. Finally, the Kendall rank correlation coefficient between each pair of variables is calculated according to the calculation formula of the Kendall rank correlation coefficient (which is more complicated and involves the calculation of parameters such as the number of synergistic pairs and non-synergistic pairs).

[0114] For example, for variable A and variable B, the Kendall rank correlation coefficient between them is calculated to be 0.6 (just an example value), which indicates that there is a certain degree of positive correlation between the two (the closer the coefficient is to 1, the stronger the positive correlation is, the closer it is to -1, the stronger the negative correlation is, and the closer it is to 0, the weaker the correlation is).

[0115] Furthermore, by combining all variables in the watershed business activity dataset in pairs and repeating the above calculation process, we can obtain the Kendall rank correlation coefficient values ​​between multiple pairs of variables and form a set of correlation results.

[0116] Furthermore, since the correlation between two variables may be affected by other variables, partial correlation analysis was performed based on the calculation of Kendall's rank correlation coefficient to determine the net correlation between the two variables.

[0117] For example, when studying the partial correlation between variables C (precipitation) and variable D (water level changes), other related variables that may affect their relationship (such as upstream water volume, regulation of surrounding water conservancy projects, etc.) are selected as control variables. Through a specific partial correlation analysis algorithm (such as calculation based on the covariance matrix, etc.), the true correlation between variables C and D after excluding the influence of these control variables is calculated, and the partial correlation coefficient value is obtained.

[0118] Furthermore, multiple rounds of partial correlation analysis can be performed on each variable in the watershed business activity dataset, considering different combinations of control variables to obtain more partial correlation coefficient values ​​and further enrich the results of the correlation analysis.

[0119] Finally, all the correlation coefficient values ​​obtained from the Kendall rank correlation coefficient and partial correlation analysis can be sorted and comprehensively considered to obtain multiple correlation values ​​that can accurately reflect the correlation between different variables in the watershed business activity dataset.

[0120] Step a2: construct a cause-effect diagram based on multiple watershed business key data and multiple correlation values.

[0121] First, each data element in the basin business activity data set is used as a node to construct a cause-and-effect diagram. For example, specific data elements such as water level, rainfall, water resource extraction, sewage treatment, and flow control parameters of various water conservancy projects correspond to a node in the cause-and-effect diagram, clearly representing each key variable in the actual business activities in the diagram.

[0122] Secondly, the causal relationship between each data element (node) can be defined based on the multiple correlation values ​​obtained. If the correlation value between two variables shows that the change of one variable can cause the regular change of another variable, and it conforms to the causal logic in business activities, then it is determined that there is a causal relationship between them.

[0123] For example, after analysis, it was found that there is a strong positive correlation between the increase in rainfall and the rise in water level. From a business logic perspective, the increase in rainfall will indeed lead to an increase in the water level in the basin. Therefore, it can be determined that there is a causal relationship between the rainfall node and the water level node, with rainfall being the "cause" and water level being the "effect."

[0124] Finally, directed edges (arrows) are used to indicate the direction of causal relationships. Arrows point from the identified cause nodes to the result nodes, visually showing the causal flow between variables in the diagram.

[0125] Specifically, during the drawing process, special attention should be paid to avoid loops, that is, a node should not return to itself through a series of directed edges. If a loop occurs, it is necessary to return to the previous correlation analysis steps, recalculate the correlation between the elements carefully, check whether there are analysis errors or missing influencing factors, and make adjustments and corrections.

[0126] At the same time, for the elements with correlation values ​​between (-0.4, 0.4), no causal relationship is constructed according to the set rules (the correlation is considered to be too weak to construct a causal relationship), that is, no directed edges are drawn between the nodes corresponding to these elements. This simplifies the causal diagram, highlights the key and reliable causal relationships, and finally obtains a clear and reasonable causal diagram to display the causal relationship network between key variables in watershed business activities.

[0127] Step S204: Based on the watershed business activity list and the watershed business activity data set, the causal graph is traversed using the preset particle diffusion model to obtain a watershed business causal chain table. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0128] Step S205: Determine the watershed business relationship based on the watershed business causal chain table and the actual watershed business process. Figure 1 Step S105 of the illustrated embodiment will not be described in detail here.

[0129] The method for determining the relationship between watershed business provided in this embodiment performs word segmentation processing on the text of watershed business data, and can convert continuous text into analyzable entry units. Further, an entry vector is constructed based on a preset entry list so that each entry can be represented in a unified vector space. Further, by using the word frequency and inverse file frequency method for processing, the importance of each entry in the text collection can be effectively measured, and the key target entries closely related to the watershed business can be screened out, reducing the interference of data noise and redundant information. Finally, by clustering multiple target entries, similar entries can be classified into one category, thereby clearly dividing different watershed business activity areas, and improving the accuracy and efficiency of business identification. Further, by performing feature selection on the watershed business activity data set through the random forest algorithm, key data with high relevance to the watershed business activities can be screened out from a large amount of data. Further, based on multiple key data of watershed business, multiple correlation values ​​are obtained by processing the Kendall rank correlation coefficient and partial correlation analysis method, which can accurately quantify the linear and nonlinear correlations between key data. Furthermore, by combining multiple key data of watershed business and multiple correlation values ​​to construct a causal graph, the existence and direction of causal relationship can be determined based on the real correlation between the data, making the constructed causal graph more scientific and reasonable and in line with the actual situation.

[0130] In this embodiment, a method for determining a watershed service relationship is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 3 is a flow chart of a method for determining a watershed service relationship according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0131] Step S301, obtain the watershed business data text, watershed business activity data set and actual watershed business process. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0132] Step S302: Based on the watershed business data text, business identification is performed using word frequency and reverse document frequency methods and cluster analysis methods to obtain a list of watershed business activities. Figure 2Step S202 of the illustrated embodiment will not be described in detail here.

[0133] Step S303: construct a cause-effect diagram based on the watershed business activity dataset and the watershed business activity list. Figure 2 Step S203 of the illustrated embodiment will not be described in detail here.

[0134] Step S304, based on the watershed business activity list and the watershed business activity data set, the causal graph is traversed using a preset particle diffusion model to obtain a watershed business causal linked list.

[0135] Specifically, the above step S304 includes:

[0136] Step S3041, determining a mapping relationship table based on the watershed business activity list and the watershed business activity data set.

[0137] Specifically, professional knowledge and business logic can be used to match each business activity in the business activity list with the relevant data in the data set to clarify which data supports and reflects the development of each business activity.

[0138] Furthermore, the structure of the mapping relationship table can be determined, which can usually be designed in a table format, including two or more columns (depending on specific needs), one of which is used to record the name or identifier of the business activity in the watershed business activity list, and the other column (or columns) is used to record the key information such as the data field name or identifier in the corresponding watershed business activity data set. For example, the first column of the table is "business activity name" and the second column is "related data field".

[0139] Furthermore, the business activities and corresponding data information can be filled into the mapping relationship table row by row according to the determined correspondence. For example, in one row, the "Business Activity Name" column is filled with "Flood Prevention and Disaster Reduction Business", and the "Related Data Field" column is filled with "Rainfall, River Water Level, Reservoir Water Storage" and other related data field names, and so on, to complete the filling of the entire mapping relationship table, clearly presenting the mapping relationship between business activities and specific data.

[0140] Step S3042: traverse the causal graph using a preset particle diffusion model to obtain a node causal relationship table.

[0141] Specifically, one or more starting nodes in the causal graph can be selected (usually those nodes that have a significant impact on the basin's business activities and are at the core of the causal relationship are given priority, such as the "rainfall" node, because it is often the starting factor for changes in many other business-related variables) and used as the initial state input of the particle diffusion model.

[0142] Furthermore, according to the rules set by the particle diffusion model, the "particles" representing the key variables of each business activity can begin to diffuse and move in the causal diagram to simulate the change propagation of these variables in actual business activities.

[0143] Furthermore, in each process of the particle diffusion model traversing the causal graph, when the change of one node triggers a corresponding change in another node (that is, a state transfer occurs along a directed edge), the causal relationship between the two nodes is recorded in detail to clarify which is the "cause" node and which is the "effect" node, as well as the relevant attributes of this causal relationship (such as the possible impact intensity, probability, etc., if there are relevant settings).

[0144] Finally, the recorded causal relationships between these nodes can be organized in a certain format. For example, a node causal relationship table in tabular form can be constructed, which contains columns such as "cause node name", "effect node name", "causal relationship strength (optional)", and "causal relationship trigger probability (optional)". The causal relationship information recorded in each traversal is filled in line by line to form a complete node causal relationship table, which intuitively presents the specific causal relationship between the nodes in the causal graph.

[0145] Step S3043, based on the node causal relationship table and the mapping relationship table, determine the watershed business causal chain table.

[0146] Specifically, the node causal relationship table and the mapping relationship table are analyzed, and the "cause node name" and "effect node name" in the node causal relationship table are used as a bridge. By searching the mapping relationship table, the specific business activities in the watershed business activity list corresponding to these node names and their related data information in the watershed business activity data set are found, so that the causal relationship based on the node level can be associated with the actual business activities and data.

[0147] Furthermore, based on the above association results, the related business activities and the causal relationships between them can be linked together in sequence according to the causal logical order between the business activities to form a watershed business causal chain list.

[0148] Among them, in the process of constructing the causal chain table of watershed business, we can determine the many-to-many relationship between different business activities (reflecting the detailed causal relationship) and the many-to-one relationship between business activities and more macro business fields (such as the different business field classifications represented by L), and clarify the multiple causal chain tables B contained in each business field. i .

[0149] Step S305, determining the watershed business relationship based on the watershed business causal chain table and the actual watershed business process.

[0150] Specifically, the above step S305 includes:

[0151] Step S3051, perform time series analysis on the basin business causal chain table to obtain the first impact set.

[0152] Among them, the time series analysis method can be a moving average method, an exponential smoothing method, an autoregressive moving average model (ARMA), an autoregressive integrated moving average model (ARIMA), etc.

[0153] The first impact set can fully reflect the mutual influence relationship between various business activities in the basin business causal chain table based on time series.

[0154] Specifically, the time series data corresponding to each business activity in the basin business causal chain table can be used as the object and analyzed using the corresponding time series analysis method.

[0155] For example, for the time series data of "river water level", if the ARIMA model is used, the data must first be tested for stationarity (such as using the unit root test method to determine whether the data is stationary, and if it is not stationary, then a difference operation is performed to make it stationary), and then the order of the model is determined (including autoregressive order p, difference order d, moving average order q). Tools such as the autocorrelation function (ACF) graph and the partial autocorrelation function (PACF) graph can be used to assist in determining the appropriate order, and then the model is fitted and the parameters are estimated to obtain a suitable ARIMA model.

[0156] Furthermore, through analysis, we can reveal the law of changes of variables related to various business activities over time. At the same time, we can analyze how the change of a business activity variable affects other subsequent business activity variables in the time dimension, and quantify the degree of such impact and time delay and other factors.

[0157] Furthermore, the specific impact relationships between the business activities in the time dimension obtained through the time series analysis are sorted and summarized to form a first impact set.

[0158] Among them, the first impact set records in detail the impact of each business activity on other business activities at different time stages, which can fully reflect the mutual influence relationship between business activities based on time series in the basin business causal chain table.

[0159] Step S3052: Decompose the actual watershed business process to obtain a second impact set.

[0160] Among them, the second impact can reflect the inherent impact mechanism between various parts in the actual watershed business process.

[0161] Specifically, according to the description of step S101, the actual watershed business process can be obtained through UML modeling.

[0162] Furthermore, the Archimate enterprise architecture modeling language and methods can be used to further decompose the existing business processes presented through UML modeling.

[0163] Among them, Archimate can break down business processes into multiple components from different levels such as business, application, and technology, analyze the relationship between each element and their role and impact in the entire business architecture. For example, a business activity in a business process can be further decomposed into specific operating steps, information system support involved, technical resource requirements and other sub-elements to explore deeper structures and relationships in the business process.

[0164] Furthermore, by decomposing the business process with Archimate, the mutual influence relationship between various business activities and their sub-elements under the existing business process framework can be determined, thereby forming a corresponding second influence set.

[0165] Step S3053: Determine the watershed service relationship based on the first impact set and the second impact set.

[0166] Specifically, carefully compare the first impact set and the second impact set and look for differences between the two.

[0167] Furthermore, by performing a difference operation on the first impact set and the second impact set, we can find out the impact relationship between the business activities that exist in the first impact set but are not reflected in the second impact set, that is, the newly discovered business relationship. For example, it is found that under a certain time series change, a business activity that was originally considered to have little to do with flood control (such as a specific operation in water resources allocation) has a significant positive impact on flood control and disaster reduction effects in a specific time period, but it is not considered as a key correlation factor in the existing actual basin business process. Then this is a newly discovered business relationship, which provides valuable clues and directions for optimizing basin business processes, exploring new business opportunities, and improving basin management benefits.

[0168] Furthermore, the newly discovered business relationships can be integrated with the business relationships in the actual watershed business processes concentrated in the original second impact to form a complete watershed business relationship.

[0169] Furthermore, the obtained watershed business relationships cover the existing business activity connections from the actual business operation level, as well as the potential and new business activity connections discovered through time series analysis. They comprehensively and accurately present the mutual influence, mutual restriction and collaborative cooperation among various business activities in the watershed, providing a solid basis for subsequent watershed management decision-making, business optimization, resource allocation and other work, and helping to improve the overall efficiency and benefits of watershed management.

[0170] The method for determining the watershed business relationship provided in this embodiment determines the mapping relationship table based on the watershed business activity list and the watershed business activity data set, establishes a bridge between business activities and data, and can effectively map the causal relationship at the data level with the activities at the business level. Further, the node causal relationship table is obtained by traversing the causal graph using the preset particle diffusion model, which reveals the causal relationship details between each node in the causal graph in detail, and provides a basis for in-depth understanding of the microstructure of the watershed business relationship. Finally, the watershed business causal chain table is determined based on the node causal relationship table and the mapping relationship table, and the causal relationship at the data level is integrated with the business activities to form a complete business causal chain, which clearly shows the causal logic between the watershed business activities. Further, by performing time series analysis on the watershed business causal chain table, the changing laws and influencing factors of the watershed business relationship in the time dimension can be captured, revealing the dynamic characteristics of the business relationship. At the same time, by decomposing the actual watershed business process, the internal structure and influencing factors in the existing business process are deeply analyzed, and the operation mechanism and key links of the current business are clarified. Finally, by comparing and integrating the results of time series analysis with the decomposition results of actual watershed business processes, we can discover business relationships and potential new business opportunities that are not fully recognized in existing business processes, which provides direction for innovation and optimization of watershed management, helps to improve the economic, social and ecological benefits of watershed management, and achieve sustainable utilization of watershed resources and continuous improvement of management levels.

[0171] In one example, a smart watershed business relationship discovery decision support method is provided, comprising:

[0172] 1. Identification of watershed business. Segment the data of watershed business, and construct a vector for each segmented word. The dimension of the vector is the same as the dimension of the entry list, and the value of the vector is the number of times each entry in the entry list appears in the text. Use the TF (term frequency)-IDF (inverse document frequency) method to determine the similarity of business fields and find important entries. Perform cluster analysis on the results of important entries, and use UML to model the results of cluster analysis to map the basic activity domain of watershed information (for example, water conservancy project construction, flood control and drought relief, hydropower generation, ecological protection, water administration, shipping and transportation, immigration management, etc.), and determine the secondary classification within the watershed, that is, the main business activities, based on the clustering results (K value), and obtain the business activity list L.

[0173] 2. Collection of watershed information data: Collect watershed business activity-related data D through existing internal systems such as ERP, supply chain management system, financial system, contract system and other business systems as well as internal databases of the organization. Repeat the method of obtaining L, map D to L, and obtain a many-to-one mapping relationship table Y between D and L.

[0174] 3. Selection of key variables of watershed information: According to the results of watershed information data collection, based on L, a random forest is constructed for all data, and feature selection is performed to obtain the key variables K of watershed information; combined with the Kendall rank correlation coefficient and partial correlation analysis, the correlation R between different data is obtained for all data.

[0175] 4. Combine the watershed information characteristics to construct an event-driven causal graph G: Construct an event-driven causal graph and use the causal graph to visualize the relationship between key variables and identify possible causal paths and interactions.

[0176] (1) Determine the nodes: Determine the nodes in the graph, each node is a data element in D;

[0177] (2) Define causal relationship: Based on R, define the causal relationship between D;

[0178] (3) Draw directed edges: Use directed edges (arrows) to indicate the direction of causal relationships, with arrows pointing from cause to effect. Avoid loops when drawing. If a loop occurs, return to step 3 to recalculate the correlation between elements. For elements with R values ​​between (-0.4, 0.4), do not build causal relationships, and obtain graph G;

[0179] (4) Construct a particle diffusion model and use it to traverse G to obtain all associated nodes D' in graph G for each node D, and record the causal relationship in the causal list C;

[0180] (5) According to the node relationship, map table C and table Y so that table Y is associated with table C, and obtain the basin business causal chain table B, obtain the many-to-many relationship of D to D (reflecting the detailed causal relationship) and the many-to-one relationship of D to L, reflecting the multiple causal chain tables B under L i .

[0181] 5. Time series analysis: Due to the obvious time variation characteristics of basin business, a time series analysis is conducted on B to obtain the specific impact I of D on L.

[0182] 6. Obtain the existing business process P through UML modeling, use Archimate to decompose P to obtain the impact subset I', and then subtract I-I' to obtain the newly discovered business relationship.

[0183] The intelligent watershed business relationship discovery decision support method provided in this example can help watershed management optimize resource utilization, improve water quality, enhance flood control capabilities, and improve the economic, social, and ecological benefits of watershed management by intelligently discovering the business relationships of watershed information and discovering new business opportunities in the field of watershed management.

[0184] In this embodiment, a device for determining a watershed business relationship is also provided, and the device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0185] This embodiment provides a device for determining a watershed service relationship. Figure 4 As shown, the device comprises:

[0186] The acquisition module 401 is used to acquire the watershed business data text, watershed business activity data set and actual watershed business process;

[0187] The identification module 402 is used to identify the business based on the watershed business data text, using the word frequency and reverse document frequency method and cluster analysis method to obtain a list of watershed business activities;

[0188] A construction module 403 is used to construct a cause-effect diagram based on the watershed business activity data set and the watershed business activity list;

[0189] A traversal module 404 is used to traverse the cause-effect graph based on the watershed business activity list and the watershed business activity data set using a preset particle diffusion model to obtain a watershed business cause-effect linked list;

[0190] The determination module 405 is used to determine the watershed business relationship based on the watershed business causal chain table and the actual watershed business process.

[0191] In some optional implementations, the identification module 402 includes:

[0192] The first processing submodule is used to perform word segmentation processing on the watershed business data text to obtain multiple initial entries.

[0193] The first construction submodule is used to construct multiple term vectors of multiple initial terms based on a preset term list.

[0194] The second processing submodule is used to obtain multiple target terms based on multiple term vectors through word frequency and inverse document frequency methods.

[0195] The analysis and determination submodule is used to perform cluster analysis on multiple target terms and determine the list of basin business activities.

[0196] In some optional implementations, the construction module 403 includes:

[0197] The feature selection submodule is used to select features of the watershed business activity data set based on the watershed business activity list using the random forest algorithm to obtain multiple key watershed business data.

[0198] The second construction submodule is used to construct a cause-effect diagram based on key business data of multiple watersheds.

[0199] In some optional embodiments, the second building block includes:

[0200] The processing unit is used to obtain multiple correlation values ​​based on multiple key watershed business data through Kendall rank correlation coefficient and partial correlation analysis method.

[0201] A construction unit is used to construct a cause-effect diagram based on multiple watershed business key data and multiple correlation values.

[0202] In some optional implementations, the traversal module 404 includes:

[0203] The first determination submodule is used to determine a mapping relationship table based on the watershed business activity list and the watershed business activity data set.

[0204] The traversal submodule is used to traverse the causal graph using a preset particle diffusion model to obtain a node causal relationship table.

[0205] The second determination submodule is used to determine the watershed business causal chain table based on the node causal relationship table and the mapping relationship table.

[0206] In some optional implementations, the determination module 405 includes:

[0207] The analysis submodule is used to perform time series analysis on the basin business causal chain list to obtain the first impact set.

[0208] The decomposition submodule is used to decompose the actual watershed business process to obtain the second impact set.

[0209] The third determination submodule is used to determine the watershed service relationship based on the first impact set and the second impact set.

[0210] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0211] The watershed business relationship determination device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0212] The embodiment of the present invention also provides a computer device having the above Figure 4 The watershed business relationship determination device shown.

[0213] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device comprises: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other, and can be installed on a common mainboard or installed in other ways as required. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0214] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0215] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0216] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0217] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0218] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0219] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0220] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0221] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for determining a watershed service relationship, characterized in that: The method comprises: Obtain basin business data text, basin business activity data set and actual basin business process; Based on the watershed business data text, business identification is performed using word frequency and reverse document frequency methods and cluster analysis methods to obtain a list of watershed business activities; constructing a cause-and-effect diagram based on the watershed business activity dataset and the watershed business activity list; Based on the watershed business activity list and the watershed business activity data set, the causal graph is traversed using a preset particle diffusion model to obtain a watershed business causal chain list, wherein the preset particle diffusion model is used to regard different values ​​or states of various variables in watershed business activities as different states of particles, and to simulate and analyze the changing trends and mutual influence relationships of variables in watershed business activities over time by defining the transition probabilities and rules of particles between these states; Determine the watershed business relationship based on the watershed business causal chain table and the actual watershed business process; Wherein, based on the watershed business activity list and the watershed business activity data set, the causal graph is traversed using a preset particle diffusion model to obtain a watershed business causal linked list, including: Determine a mapping relationship table based on the watershed business activity list and the watershed business activity data set; Using the preset particle diffusion model to traverse the causal graph, a node causal relationship table is obtained, wherein the node causal relationship table is used to associate the causal relationship based on the node level with the actual business activities and data; Based on the node causal relationship table and the mapping relationship table, the watershed business causal chain table is determined, and the watershed business causal chain table includes the many-to-many relationship between different business activities and the many-to-one relationship between business activities and more macro business areas.

2. The method according to claim 1, characterized in that: Based on the text of the watershed business data, the word frequency and reverse document frequency method and cluster analysis method are used to identify the business, and a list of watershed business activities is obtained, including: Performing word segmentation processing on the watershed business data text to obtain a plurality of initial entries; Based on a preset term list, construct a plurality of term vectors of the plurality of initial terms; Based on the multiple term vectors, multiple target terms are obtained by processing the term frequency and inverse document frequency method; Cluster analysis is performed on the multiple target terms and a list of business activities in the watershed is determined.

3. The method according to claim 1, characterized in that Based on the watershed business activity data set and the watershed business activity list, a cause-effect diagram is constructed, including: Based on the watershed business activity list, a random forest algorithm is used to perform feature selection on the watershed business activity data set to obtain a plurality of key watershed business data; The cause-effect diagram is constructed based on the multiple watershed business key data.

4. The method according to claim 3, characterized in that: Constructing the cause-effect diagram based on the multiple key business data of the watersheds includes: Based on the multiple key watershed business data, multiple correlation values ​​are obtained through processing using Kendall rank correlation coefficient and partial correlation analysis methods; The cause-effect graph is constructed based on the multiple watershed business-critical data and the multiple correlation values.

5. The method according to claim 1, characterized in that Based on the watershed business causal chain table and the actual watershed business process, determining the watershed business relationship includes: Performing time series analysis on the basin business causal chain table to obtain a first impact set; Decomposing the actual watershed business process to obtain a second impact set; The flow domain service relationship is determined based on the first impact set and the second impact set.

6. A device for determining a watershed service relationship, characterized in that: The device is used to execute the method for determining a watershed service relationship according to any one of claims 1 to 5, comprising: The acquisition module is used to obtain the watershed business data text, watershed business activity data set and actual watershed business process; An identification module, used for identifying business based on the watershed business data text by using word frequency and reverse document frequency method and cluster analysis method to obtain a list of watershed business activities; A construction module, configured to construct a cause-effect diagram based on the watershed business activity data set and the watershed business activity list; A traversal module, used to traverse the causal graph based on the watershed business activity list and the watershed business activity data set using a preset particle diffusion model to obtain a watershed business causal chain list; A determination module is used to determine the watershed business relationship based on the watershed business causal chain table and the actual watershed business process.

7. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for determining the watershed business relationship according to any one of claims 1 to 5 by executing the computer instructions.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for determining the watershed service relationship according to any one of claims 1 to 5.

9. A computer program product, characterized in that It comprises computer instructions, and the computer instructions are used to enable a computer to execute the method for determining the watershed business relationship according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for establishing information chain of emergent pollution accident in river basin

    CN107705001A

  • Text data classification method and device, electronic equipment and computer readable medium

    CN111241273A

  • Internet insurance user feature screening method and device based on diffusion model

    CN117972381A