Industrial chain graph association rule mining method and device and storage medium

By constructing a knowledge graph and using frequent pattern growth algorithms to mine association rules, the problem of difficulty in capturing changes in rural industrial chains in the existing technology is solved, and the efficiency and accuracy of mining of industrial chain association relationship information is improved.

CN119962652AActive Publication Date: 2025-05-09BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411928999.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-09
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The existing industrial chain map construction method is difficult to capture the rapid changes in the structure and association of rural industrial chains in real time, resulting in a decrease in the timeliness and accuracy of the map.

Method used

By obtaining the data set of rural characteristic industrial chains, building a knowledge graph, performing feature vector transformation, dimensionality reduction and discretization of service resource information, combining frequent mode growth algorithms to mine association rules, and obtaining target association rules between multiple service resource information.

Benefits of technology

It improves the efficiency and accuracy of information mining of related relationships in the industrial chain, and can more accurately locate and analyze the service entities, production entities and service processes in the industrial chain, enhancing the timeliness and accuracy of the map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962652A_ABST
    Figure CN119962652A_ABST
Patent Text Reader

Abstract

The invention provides an industrial chain graph association rule mining method and device and a storage medium, and the method comprises the steps: carrying out the construction of a knowledge graph based on a service main body, a production main body and a service process, and obtaining an industrial chain graph of a target industrial chain; performing feature vector conversion on the plurality of pieces of service resource information in the industrial chain atlas to obtain semantic feature vectors and attribute feature vectors; performing feature dimension reduction and discretization processing on the semantic feature vector to obtain semantic feature representation; discretizing the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set; performing association rule mining based on the service main body resource item set corresponding to each piece of service resource information through a frequent pattern growth algorithm to obtain a target association rule among the multiple pieces of service resource information; according to the invention, the efficiency and accuracy of industrial chain association relationship information mining can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graph technology, and in particular to a method, device and storage medium for mining association rules of industrial chain graphs. Background Art

[0002] The industrial chain map is a visualization tool that describes the various links in an industry and their interrelationships. It integrates data such as enterprises, transactions, industries, information, and popularity to form a rich knowledge map, and presents the various links in the industrial chain and their relationships, or their development status and trends in a graphical way.

[0003] Rural industries are affected by various factors such as the market, policies, and environment, and their industrial chain structures and relationships are constantly changing. Existing industrial chain map construction methods often find it difficult to capture these changes in real time, resulting in a decrease in the timeliness and accuracy of the map. Summary of the invention

[0004] The present invention provides a method, device and storage medium for mining association rules of an industrial chain graph, so as to solve the defect in the prior art that the industrial chain graph is difficult to adapt to the rapidly changing rural industrial environment, and to improve the efficiency and accuracy of mining industrial chain association relationship information.

[0005] The present invention provides an industrial chain graph association rule mining method, comprising the following steps. Obtain a data set of a target industrial chain, wherein the data set includes: a service subject, a production subject and a service process, and the target industrial chain is a rural characteristic industrial chain; construct a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; perform feature vector conversion on multiple service resource information in the industrial chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; perform feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; perform discretization processing on the attribute feature vector to obtain an attribute feature representation; combine the semantic feature representation with the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; perform association rule mining based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain a target association rule between the multiple service resource information.

[0006] According to a method for mining association rules of an industrial chain graph provided by the present invention, a knowledge graph is constructed based on the service subject, the production subject and the service process to obtain the industrial chain graph of the target industrial chain, including: extracting knowledge based on the service subject, the production subject and the service process to obtain multiple entities and entity relationships between the multiple entities; integrating knowledge of the multiple entities and the relationships between the multiple entities based on a predefined ontology to obtain the industrial chain graph of the target industrial chain, wherein the predefined ontology is used to define entity categories, entity attributes and entity relationships.

[0007] According to a method for mining association rules in an industrial chain graph provided by the present invention, the service resource information includes text information, numerical information, type information and time information, and the feature vector conversion of multiple service resource information in the industrial chain graph to obtain semantic feature vectors and attribute feature vectors corresponding to each service resource information includes: extracting semantic features from the text information to obtain semantic feature vectors; normalizing the numerical information to obtain numerical features; encoding the type information to obtain type features; performing time series processing on the time information to obtain time features; and using the numerical features, the type features and the time features as the attribute feature vectors of the service resource information.

[0008] According to a method for mining association rules of an industrial chain graph provided by the present invention, after combining the semantic feature representation with the attribute feature representation to obtain the service subject resource item set corresponding to each service resource information, the method further includes: determining the ratio between the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set and the number of all service subject resource item sets, as the support between the first service subject resource item set and the second service subject resource item set; determining the ratio between the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set and the number of the first service subject resource item set, as the confidence between the first service subject resource item set and the second service subject resource item set.

[0009] According to a method for mining association rules of an industrial chain graph provided by the present invention, the method also includes: taking the ratio of the confidence to the support of the second service subject resource item set as the improvement between the first service subject resource item set and the second service subject resource item set.

[0010] According to a method for mining association rules in an industrial chain graph provided by the present invention, association rules are mined based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain target association rules between the multiple service resource information, including: traversing all service subject resource item sets to obtain the frequency of each service subject resource item set; determining the support of each service subject resource item set based on the frequency of each service subject resource item set; deleting the service subject resource item set whose support is less than a support threshold to obtain an item header table; sorting the service subject resource item sets in the item header table in reverse order according to support and inputting them into an FP tree; generating multiple frequent item sets based on the item header table and the FP tree; determining the confidence of the association rule corresponding to each frequent item set in the multiple frequent item sets, and taking the association rule whose confidence is greater than the confidence threshold as a strong association rule; determining the lift of each of the strong association rules, and taking the strong association rule whose lift is greater than the lift threshold as the target association rule.

[0011] The present invention also provides an industrial chain graph association rule mining device, comprising the following modules: an acquisition module, used to acquire a data set of a target industrial chain, wherein the data set includes: a service subject, a production subject and a service process, and the target industrial chain is a rural characteristic industrial chain; a construction module, used to construct a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; a conversion module, used to perform feature vector conversion on multiple service resource information in the industrial chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; a processing module, used to perform feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; the processing module is also used to discretize the attribute feature vector to obtain an attribute feature representation; a combination module, used to combine the semantic feature representation with the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; a mining module, used to perform association rule mining based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain a target association rule between the multiple service resource information.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for mining association rules of an industrial chain graph as described above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for mining association rules of industrial chain graphs as described in any one of the above is implemented.

[0014] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned industrial chain graph association rule mining methods.

[0015] The industrial chain graph association rule mining method, device and storage medium provided by the present invention can accurately locate and analyze the service subjects, production subjects and service processes in the target industrial chain by acquiring the data set of the rural characteristic industrial chain; construct a knowledge graph based on the acquired data set, which can intuitively display the association relationship and service process between the subjects in the industrial chain; convert the service resource information in the industrial chain graph into feature vectors, and perform dimensionality reduction and discretization processing on it, which can simplify the complexity of the data, ensure that all attributes can be analyzed in a unified feature space, and improve the efficiency of subsequent association rule mining; use the frequent pattern growth algorithm to mine association rules based on the service subject resource item set, and obtain the target association rules between multiple service resource information, revealing the potential association and dependency relationship between multiple service resource information; thereby improving the efficiency and accuracy of mining industrial chain association relationship information. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 It is a flow chart of the industrial chain graph association rule mining method provided by the present invention.

[0018] Figure 2 It is a flow chart of the rural characteristic industry information crawling process provided by the present invention.

[0019] Figure 3 It is a flow chart for constructing a knowledge graph provided by the present invention.

[0020] Figure 4 It is a schematic diagram of the framework of the industrial chain map for constructing the target industrial chain provided by the present invention.

[0021] Figure 5 It is a structural schematic diagram of the industrial chain graph association rule mining device provided by the present invention.

[0022] Figure 6 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] As a visualization tool, the industrial chain map can clearly display the links of the industrial chain and their interrelationships, providing strong support for mining association relationships. The present invention relates to the technical field of industrial chain analysis and data mining, and aims to deeply mine the association relationships between various elements in rural characteristic industries by constructing and analyzing industrial chain maps.

[0025] Rural industries involve many sub-sectors and complex economic environments, with scattered data and uneven quality. Effectively collecting, cleaning and integrating these data has become a challenge in building an accurate industrial chain map. The lack of unified data standards and sharing platforms has led to serious data silos, limiting the comprehensiveness and accuracy of data.

[0026] Secondly, rural industries are affected by various factors such as the market, policies, and environment, and their industrial chain structures and relationships are constantly changing. Existing industrial chain map construction methods often have difficulty capturing these changes in real time, resulting in reduced timeliness and accuracy of the map. The lack of a dynamic update mechanism makes it difficult for the map to adapt to the rapidly changing rural industrial environment.

[0027] Finally, the existing methods of constructing industrial chain maps focus on displaying the basic structure and main links of the industrial chain, but are still insufficient in exploring the deep and complex relationships between multiple factors within and between industries. This limits the potential of the map in revealing the laws of industrial development and predicting industrial trends.

[0028] The purpose of this invention is to build a complete data collection, integration and verification mechanism to ensure that data from different channels and in different formats can be efficiently and accurately integrated into the industrial chain map. By introducing advanced data cleaning and preprocessing technologies, the accuracy and availability of data can be improved, laying a solid foundation for subsequent association mining. Secondly, using the FP-Growth algorithm, this method will deeply explore the deep and complex associations between multiple factors within rural characteristic industries and between industries. Finally, by constructing a multi-factor association mining device and storage device for rural characteristic industries, a scientific and comprehensive reference basis is provided for decision makers.

[0029] Optionally, the industrial chain graph association rule mining method of the embodiment of the present application can be executed by a server, or by a terminal device, or jointly by a server and a terminal device, taking the execution of the industrial chain graph association rule mining method of the embodiment by a server as an example.

[0030] Figure 1 is a flow chart of the method for mining association rules of industrial chain graph provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 101, obtaining a data set of a target industrial chain, wherein the data set includes: a service subject, a production subject, and a service process, and the target industrial chain is a rural characteristic industrial chain.

[0031] In the process of constructing the map of the rural characteristic industrial chain, the latest industry dynamics, trend analysis and market reports were obtained through relevant rural characteristic information release websites. Through the research and analysis of the industrial chain, the classification of the industry was clarified, which was divided into agriculture, forestry, animal husbandry, fishery, etc. The industrial chain was divided into production, processing, storage, transportation, sales and other links.

[0032] refer to Figure 2 , Figure 2 It is a flow chart of the process of crawling rural characteristic industry information provided by the present invention, which includes sending a request (GET); server response; parsing data, using regular expressions to extract social service information; and maintaining information and storing it in a database.

[0033] Using web crawler technology, we extract third-party data sources and social service-related websites to collect data and information related to the industry chain, including text data, geographic data, image data, etc. Figure 2 shown.

[0034] In an embodiment of the present invention, the service subject data set includes entities or institutions that provide various services for the rural characteristic industrial chain. For example, service agencies, agricultural cooperatives, agricultural enterprises, scientific research institutions, financial institutions, etc. The service subject data set may include: Service subject name, used to accurately record the full name of the service subject. Service type, used to describe in detail the type of service provided by the service subject, such as technical training, market information, financial services, logistics support, policy guidance, etc. Service scope, used to illustrate the geographical scope, industry scope or customer group covered by the service subject. Service effect, used to evaluate the contribution of the service subject to the rural characteristic industrial chain, such as improving production efficiency, increasing output, improving product quality, promoting sales, etc.

[0035] The production entity dataset mainly covers the actual producers in the rural characteristic industrial chain, including small farmers, family farms, agricultural enterprises, etc. The production entity dataset includes: the name of the production entity, which is used to accurately record the full name or abbreviation of the production entity. The production type is used to describe the type of agricultural production engaged in by the production entity, such as planting, breeding, agricultural product processing, etc. The production scale is used to illustrate the production scale of the production entity, such as planting area, breeding number, processing capacity, etc. Product types are used to list the main types of agricultural products produced by the production entity. Production and sales are used to record the actual production and sales of the production entity, as well as sales channels and market distribution.

[0036] The service process data set mainly includes various service processes provided by service entities to production entities in the rural characteristic industrial chain, including multiple links such as prenatal, intrapartum, and postpartum. The service process data set may include: service links, which are used to clearly describe the links to which the service process belongs, such as prenatal preparation, planting / breeding process, product processing, quality inspection, packaging and transportation, and market sales. Service content, which is used to describe in detail the specific service content provided by the service entity in each service link, such as the specific content of technical training, the way of providing market information, the types and conditions of financial services, etc. Service time, which is used to record the specific time or time range of service provision. Service effect, which is used to evaluate the degree of help of the service process to the production entity, such as improving production efficiency, reducing costs, improving product quality, and promoting sales.

[0037] Step 102, construct a knowledge graph based on the service subject, production subject and service process to obtain the industrial chain graph of the target industrial chain.

[0038] In the embodiment of the present invention, the collected data sets are cleaned, deduplicated, formatted, and other operations are performed to ensure the accuracy and consistency of the data.

[0039] Use natural language processing techniques, such as named entity recognition (NER), to automatically identify entities with specific meanings from text data, such as service entity names, production entity names, service names, etc.

[0040] Extract the relationship between entities from the text, such as the service relationship between the service subject and the production subject, the production relationship between the production subject and the product, etc. Relationship extraction can adopt rule-based methods, deep learning-based methods, etc.

[0041] Collect attribute information of specific entities, such as service type and service scope of service subject; production scale and technical level of production subject; product type and output, etc. Attribute extraction helps to enrich the content of knowledge graph and make it more complete and accurate.

[0042] Integrate the extracted entities, relationships, and attribute information to eliminate contradictions and ambiguities. Through technologies such as entity linking and knowledge merging, information from different sources is integrated into a unified knowledge graph.

[0043] Use graph databases (such as Neo4j) or RDF standard storage formats to store the constructed knowledge graph. Graph databases have significantly improved the efficiency of associated queries compared to traditional relational data storage methods, so they are more suitable for storing complex knowledge graphs.

[0044] Rural characteristic industries focus on fields, subjects and services, among which fields are divided into characteristic agricultural products, regional characteristics, local resources, characteristic agricultural product processing, leisure agriculture, rural tourism, etc. Subjects are divided into subject name and type (supply and marketing cooperatives, farmers' professional cooperatives, service-oriented associations, rural collective economic organizations, leading enterprises, professional service companies, family farms, individual business households, individual demonstration households, and others), business status, region, detailed address, contact name, telephone number, service subject introduction, service items and other information. At the same time, the service subject is also associated with agricultural materials, agricultural technology, products, markets, credit and other data. Service data refers to service process data, mainly including agricultural production services, agricultural material supply, agricultural technology services, post-production processing services, agricultural machinery leasing, and agricultural financial services. Service data creates a connection between service subjects and production subjects. Based on the relationship between service subjects, production subjects and service processes, a knowledge graph framework for rural socialized service resources is constructed. According to an industrial chain graph association rule mining method provided by the present invention, a knowledge graph is constructed based on a service subject, a production subject and a service process to obtain an industrial chain graph of a target industrial chain, including: Based on the service subject, production subject and service process, knowledge is extracted to obtain multiple entities and entity relationships between multiple entities; Based on a predefined ontology, knowledge integration is performed on multiple entities and the relationships between multiple entities to obtain an industrial chain map of the target industrial chain, wherein the predefined ontology is used to define entity categories, entity attributes and entity relationships.

[0045] Here, knowledge extraction includes using natural language processing (NLP) technology, such as named entity recognition (NER), to automatically identify and extract entities such as service entities and production entities from text (datasets). Identify and extract attribute information related to these entities, such as company name, address, size, main business, etc. Analyze the association between entities in the text and extract the service relationship, production relationship, etc. between them. Relationships can include supply chain relationships (such as supplier-manufacturer relationships), cooperative relationships (such as manufacturer-distributor relationships), service relationships (such as service provider-customer relationships), etc.

[0046] Knowledge integration includes pre-defining an ontology to define entity categories, entity attributes, and entity relationships. The ontology should cover the entity categories and relationship types of all targets in the industrial chain to ensure the completeness and accuracy of knowledge integration. Map the extracted entities and relationships with the entity categories and relationship types in the ontology. For entities or relationships that cannot be directly mapped, new categories or types can be created and added to the ontology. Integrate the mapped entities and relationships into a unified knowledge graph. Use graph databases or storage formats such as RDF to store knowledge graphs for efficient query and analysis.

[0047] refer to Figure 3 , Figure 3 It is a knowledge graph construction flow chart provided by the present invention, which includes: entities, relationships, attributes; ontology, relationships, attributes; knowledge extraction, knowledge integration, knowledge storage, update, control; structured data, unstructured data; structured triples and knowledge base.

[0048] In an embodiment of the present invention, the collected data is used to draw an industrial chain map to clarify the connection between each link. First, the ontology is clearly defined, and the entity categories, attributes and relationships are defined. It mainly includes four parts: knowledge extraction, knowledge integration, knowledge storage, knowledge update maintenance and quality control. Among them, knowledge extraction requires natural language processing technology and other technologies to identify entities in the text and the relationship between entities. Knowledge integration, to integrate the extracted knowledge into a unified framework in an appropriate manner, requires first defining an ontology to determine the categories, attributes and relationships in the knowledge graph. Knowledge storage, to effectively store and index knowledge for subsequent query and analysis, can be stored based on an open source graph database. The knowledge graph construction flow chart is as follows: Figure 3 shown.

[0049] Through the embodiments of the present invention, through knowledge extraction, key information related to service subjects, production subjects and service processes can be efficiently extracted from a large amount of text data; through pre-defined ontologies, multiple entities and their relationships can be integrated into a unified knowledge system; ontology-based knowledge integration supports semantic reasoning, and can derive new relationships implicit in the knowledge graph, further enriching and improving the industrial chain graph.

[0050] Step 103, performing feature vector conversion on multiple service resource information in the industrial chain map to obtain semantic feature vectors and attribute feature vectors corresponding to each service resource information.

[0051] In the embodiment of the present invention, the socialized service resource information includes long text information such as service crops, service items, and service subject introductions. Through the preset semantic feature extraction representation model, each item will be converted into a semantic-level feature vector. For attribute information such as service crop type, location, and time, feature engineering processing is performed. This includes appropriate processing of different types of attributes such as numerical, categorical, and time types, numerical normalization, category coding, and time series processing to ensure that all attributes can be analyzed in a unified feature space.

[0052] According to an industrial chain graph association rule mining method provided by the present invention, service resource information includes text information, numerical information, type information and time information, and multiple service resource information in the industrial chain graph are converted into feature vectors to obtain semantic feature vectors and attribute feature vectors corresponding to each service resource information, including: Extract semantic features from text information to obtain semantic feature vectors; Perform numerical normalization on the numerical information to obtain numerical features; Encode the type information to obtain the type characteristics; Perform time series processing on time information to obtain time features; The numerical features, type features and time features are used as attribute feature vectors of service resource information.

[0053] In an embodiment of the present invention, semantic features of text information are extracted through natural language processing technology (NLP). For example, a word segmentation tool (such as jieba) is used to segment the text into words or phrases; a pre-trained word vector model (such as Word2Vec, BERT) is used to convert each word into a high-dimensional vector; the vectors of all words are aggregated (such as average, weighted average, TF-IDF weighting, etc.) to obtain a semantic feature vector of the entire text.

[0054] Numerical information may have different dimensions and ranges. In order to unify the processing, it needs to be normalized; for example, through minimum-maximum normalization, the values ​​are scaled to a specified range (such as 0 to 1); through Z-score normalization, the data is scaled according to the mean and standard deviation of the data to make it conform to the standard normal distribution.

[0055] Type information is usually represented as discrete category labels, which can be converted into numerical features through encoding. For example, through One-Hot Encoding, a separate binary column is created for each category, with a value of 1 if the data belongs to that category and 0 otherwise. Through Label Encoding, each category is mapped to a unique integer. Through Target Encoding, the target variable mean of the category label is encoded, which is often used to process categorical data with target information.

[0056] Time information usually includes date, timestamp, etc., and features can be extracted through time series analysis; for example, decomposing time into components such as years, months, days, hours, and minutes; calculating the difference between the current time and a reference time (such as the start time of an event); extracting periodic features of time, such as day of the week, quarter, etc.; using time series analysis techniques (such as autoregressive models, moving averages, etc.) to extract features.

[0057] Through the embodiments of the present invention, text information is converted into high-dimensional vectors through semantic feature extraction, and these vectors can capture the semantic relationship and context information between words, thereby enhancing the machine's ability to understand the text content; numerical normalization can eliminate the dimensional differences between different numerical features; type encoding converts discrete category information into numerical features, so that the category information can participate in subsequent mathematical operations or model training; time series processing can capture trends and periodic changes in time information, providing an important basis for subsequent prediction and analysis.

[0058] Step 104, performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation.

[0059] Feature dimensionality reduction is to reduce the dimension of data while retaining the important information of the original data as much as possible. For semantic feature vectors, dimensionality reduction helps to reduce redundant information and improve computational efficiency.

[0060] For example, according to the principal component analysis method, a set of variables that may be correlated is converted into a set of linearly uncorrelated variables, namely, principal components, through orthogonal transformation. Applying the principal component analysis method on the semantic feature vector can remove redundant features and retain the most representative components.

[0061] Use pre-trained word embedding models (such as Word2Vec, GloVe, etc.) to map words to low-dimensional vector space, and use these vectors to represent the semantic information of words. In the semantic feature vector, the word embedding model can be used to convert the high-dimensional one-hot encoding into a low-dimensional dense vector, thereby achieving dimensionality reduction.

[0062] Discretization is the process of converting continuous data into discrete data representations, usually to simplify data representation, improve computational efficiency, or meet the requirements of specific algorithms. For semantic feature vectors, discretization may involve converting continuous vector values ​​into discrete category labels or binary values.

[0063] For example, through the threshold method, on the semantic feature vector, the threshold can be set according to the value range of each feature to convert the continuous feature value into a discrete category label; through the clustering method, the semantic feature vector is clustered using algorithms such as K-means and hierarchical clustering to convert the continuous feature value into a discrete cluster label.

[0064] After feature dimensionality reduction and discretization, a simplified semantic feature representation can be obtained. Semantic feature representation not only reduces the dimension and complexity of the data, but also retains important information in the original data, which is helpful for subsequent analysis and modeling tasks.

[0065] Step 105, discretize the attribute feature vector to obtain the attribute feature representation.

[0066] In the embodiment of the present invention, the attribute feature vector is discretized to convert the continuous numerical features into discrete category labels or binary values, thereby simplifying data representation and improving calculation efficiency.

[0067] According to the pre-set discretization method and determined parameters, the attribute feature vector is discretized to convert the continuous numerical features into discrete category labels or binary values; after discretization, the attribute feature vector will be converted into a vector composed of discrete category labels or binary values, that is, the attribute feature representation.

[0068] Step 106: Combine the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information.

[0069] In the embodiment of the present invention, feature dimension reduction and discretization processing are performed on the semantic feature vector, and discretization processing is performed on other attribute features, and the two are combined to form a service subject resource item set. After completing the feature engineering, the FP-Growth algorithm is used to mine association rules on the processed data. The FP-growth algorithm is an extension of the Apriori algorithm. Since the algorithm only scans the data set twice, it is very efficient in the discovery process of frequent item sets.

[0070] In some embodiments, attributes such as "service organization name", "service items", "service location", "service time", "service crop type", "service subject type", "service subject introduction", "annual turnover", and "establishment time (year)" are selected to form an attribute set. Among them, annual turnover and establishment time (year) are in numerical form, and the numerical values ​​are discretized according to the interval according to the probability distribution. Categorical attributes form item sets according to each value of each type of attribute, and are represented by category coding. For example, production service types include pruning, fertilizing, spraying, weeding, harvesting, full trusteeship, and others.

[0071] After feature engineering, the resource item set of a single service entity is in the form of: { "Service item semantic vector": [0.1, 0.3, 0.5, ..., 0.2], "Service subject introduction semantic vector": [0.2, 0.4, 0.1, ..., 0.3], "Service Industry Type": 1, "Service Location": 2, "Service Time_Year": 2024, "Service Time_Month": 8, "Service time_day": 16, "Annual turnover": "100,000-600,000", "Duration": "1-5 years" } According to an industrial chain graph association rule mining method provided by the present invention, after combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information, the method further includes: Determine the ratio between the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set and the number of all service subject resource item sets as the support between the first service subject resource item set and the second service subject resource item set; A ratio between the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set and the number of the first service subject resource item set is determined as a confidence level between the first service subject resource item set and the second service subject resource item set.

[0072] The support is calculated by the ratio between the number of times two service subject resource item sets (the first service subject resource item set and the second service subject resource item set) appear at the same time and the total number of times all service subject resource item sets appear. This ratio reflects the frequency of the two item sets appearing together in the total data set.

[0073] The higher the support, the more times the two service subject resource item sets appear together in the dataset, and the stronger the correlation between them may be.

[0074] The confidence calculation is the probability that the second service subject resource item set will also appear when the first service subject resource item set appears. This probability reflects the possibility of the second service subject resource item set appearing under certain conditions (i.e. the first service subject resource item set appears).

[0075] The higher the confidence, the greater the possibility that the second service subject resource item set will also appear when the first service subject resource item set appears, that is, the higher the reliability of the association rule.

[0076] In the embodiment of the present invention, all user data sets (i.e., all service subject resource item sets) are I, and the support of the association rule A⇒B is the probability of A and B appearing at the same time, i.e., the ratio of the number of data where A and B appear at the same time to the total number of records, as shown in formula (1): (1) Where A represents the first service subject resource item set, B represents the second service subject resource item set, Indicates the support between the first service subject resource item set and the second service subject resource item set; Indicates the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set. Indicates the number of all service principal resource item sets.

[0077] The confidence level is the probability of B appearing when A appears, that is, the ratio of the number of data where A and B appear at the same time to the number of records where A appears, as shown in formula (2): (2) in, represents the confidence between the first service subject resource item set and the second service subject resource item set, A represents the first service subject resource item set, B represents the second service subject resource item set, Indicates the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set. Indicates the number of resource item sets of the first service principal.

[0078] Through the embodiments of the present invention, by calculating the support and confidence, frequent item sets and strong association rules can be quickly screened out, reducing the amount of calculation and time cost in the data mining process; by calculating the support and confidence, the correlation between the service subject resource item sets can be more intuitively displayed, enhancing the interpretability and readability of the data.

[0079] According to a method for mining association rules of an industrial chain graph provided by the present invention, the method further includes: The ratio of the confidence level to the support level of the second service subject resource item set is used as the lift between the first service subject resource item set and the second service subject resource item set.

[0080] In the embodiment of the present invention, the lift indicates whether the appearance of A has a positive or negative effect on the appearance of B, that is, The ratio of the confidence of A to the support of B is shown in formula (3). In the application, a strong association rule with a lift greater than 3 is considered to be a valid association rule.

[0081] (3) in, Indicates the degree of improvement between the first service subject resource item set and the second service subject resource item set. Indicates the confidence between the first service subject resource item set and the second service subject resource item set, Indicates the support of the second service subject resource item set (that is, the occurrence frequency of the second service subject resource item set).

[0082] Here, for the first service subject resource item set and the second service subject resource item set, support indicates the frequency of the two item sets appearing at the same time. Confidence indicates the probability of the other item set appearing when one item set appears. Lift indicates the ratio of confidence to support, which is used to measure the degree of influence of the appearance of one item set on the appearance of another item set.

[0083] Lift is used to indicate whether the impact of the appearance of A (the first service subject resource item set) on the appearance of B (the second service subject resource item set) is independent, positively correlated, or negatively correlated. If the lift is greater than 1, it means that the appearance of A has a positive impact on the appearance of B (that is, A and B are positively correlated); if the lift is equal to 1, it means that A and B are independent; if the lift is less than 1, it means that the appearance of A has a negative impact on the appearance of B (that is, A and B are negatively correlated).

[0084] According to the embodiment of the present invention, lift is an important indicator to measure whether the appearance of one item set (or rule) is independent of another item set. By calculating lift, it can be determined whether the association between two item sets is random, independent, or has some kind of dependency.

[0085] Step 107 , using a frequent pattern growth algorithm, association rule mining is performed based on the service subject resource item set corresponding to each service resource information to obtain target association rules between multiple service resource information.

[0086] By using the Frequent Pattern Growth (FPGrowth) algorithm to mine association rules, the target association rules between multiple service resource information can be efficiently identified based on the service subject resource item set corresponding to each service resource information.

[0087] Convert the service principal resource item set into a form suitable for processing by the FPGrowth algorithm, such as a transaction database.

[0088] Traverse the transaction database and calculate the frequency of occurrence of each item set (or item in the service subject resource item set). Filter out frequent items according to the set minimum support threshold. Based on frequent items, build a frequent item set tree (FP-Tree). FP-Tree is a compressed data structure used to store frequent item sets and their associations.

[0089] Starting from the root node of the FP-Tree, all frequent patterns (i.e., frequent item sets) are mined recursively. For each frequent item, its corresponding conditional FP-Tree is constructed to further mine the frequent patterns containing the item.

[0090] For each frequent pattern, calculate the confidence of the corresponding association rule. According to the set minimum confidence threshold, filter out the association rules that meet the conditions. Sort and optimize the filtered association rules (i.e., target association rules) to make them easier to understand and apply.

[0091] According to an industrial chain graph association rule mining method provided by the present invention, association rule mining is performed based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain target association rules between multiple service resource information, including: Traverse all service subject resource item sets to obtain the frequency of each service subject resource item set; Based on the frequency of each service subject resource item set, determine the support of each service subject resource item set; The service subject resource item set whose support is less than the support threshold is deleted to obtain the item header table; Sort the service subject resource item sets in the item header table in descending order of support and input them into the FP tree; Generate multiple frequent item sets based on the item header table and FP tree; Determine the confidence of the association rules corresponding to each frequent item set in multiple frequent item sets, and take the association rules whose confidence is greater than the confidence threshold as strong association rules; The lift of each strong association rule is determined, and the strong association rule with a lift greater than a lift threshold is used as a target association rule.

[0092] In the embodiment of the present invention, the frequent pattern growth (FP-growth) algorithm performs the following steps: Step 1: traverse the data set and calculate the frequency based on each possible value of each attribute.

[0093] In this step, the algorithm goes through the entire data set and counts each value of each attribute to calculate the frequency of each item (that is, the number of times the item appears in the data set).

[0094] Step 2: Calculate the minimum data item threshold according to the support, delete the data items that are less than the threshold, put them into the item header table, and sort them in descending order.

[0095] Set a support threshold (usually determined based on experience or business needs), and then remove items whose frequency is less than the threshold because these items are considered infrequent.

[0096] The remaining frequent items are put into the Header Table and sorted in descending order according to their frequency (support). The Header Table is used to quickly find and access frequent items.

[0097] Step 3, traverse the data set and delete the itemsets deleted in step 2. Insert the data into the FP tree in order of support.

[0098] Traverse the dataset again, but this time consider only those frequent items in the item header table. Insert the transaction data into a frequent pattern tree (FP-Tree) based on the frequency (support) of the items and the order in which they appear in the transaction.

[0099] FP-Tree is a special prefix tree used to efficiently store and find frequent itemsets.

[0100] Step 4: Traverse the item header table in reverse order to find the corresponding conditional pattern base to obtain frequent item sets.

[0101] Traverse each frequent item from the item header table in order of frequency (support) from high to low. For each frequent item, find all its prefix paths in the FP-Tree (these paths constitute the conditional pattern base of the item). Use these conditional pattern bases to build the conditional FP-Tree and recursively mine all frequent item sets that contain the current frequent item.

[0102] Step 5: For each non-empty subset of the frequent itemsets, calculate the confidence and obtain the strong association rule (i.e., the target association rule) according to the confidence threshold.

[0103] For each frequent item set mined, generate all possible non-empty subsets (these subsets constitute the antecedents of the association rules). Calculate the confidence of each association rule (i.e., the probability of the consequent appearing when the antecedent appears). According to the set confidence threshold, filter out those association rules whose confidence is higher than the threshold. These rules are considered strong association rules.

[0104] Step 6: For each strong association rule, calculate the lift and output the association rule if the threshold is met.

[0105] In association rule mining, lift is an important indicator to measure the effectiveness of rules, which reflects the independence between the antecedent and consequent of the rule.

[0106] For the antecedent and consequent of the strong association rule, their support in the data set is calculated respectively. At the same time, the support of the antecedent and consequent appearing at the same time is calculated, which is usually calculated when mining frequent itemsets.

[0107] Confidence is another measure of rule strength, which indicates the probability that the consequent will also occur when the antecedent occurs. For example, confidence = (support of antecedent & consequent) / support of antecedent.

[0108] Lift is the ratio of confidence to support of the consequent, which reflects the impact of the antecedent on the probability of occurrence of the consequent. Lift = confidence / support of the consequent.

[0109] According to the set lift threshold, the association rules with lifts higher than the threshold are filtered out.

[0110] If the lift of a rule is greater than 1, it means that there is a positive correlation between the antecedent and the consequent, that is, the appearance of the antecedent increases the probability of the consequent appearing.

[0111] If the lift of a rule is equal to 1, it means that the antecedent and consequent are independent, that is, the occurrence of the antecedent has no effect on the probability of the consequent occurring.

[0112] If the lift of a rule is less than 1, it means that there is a negative correlation between the antecedent and the consequent, that is, the appearance of the antecedent reduces the probability of the consequent appearing.

[0113] For association rules that meet the lift threshold, they are output as final valid rules.

[0114] Through the embodiments of the present invention, by calculating the lifting degree and screening the association rules, the mined strong association rules can be further refined, and only the target association rules that can reflect the potential association in the data set are retained.

[0115] This method can reveal the potential associations between different service resources on the rural social service platform. For example, it can be found that certain service items are more likely to be used together in specific locations and time periods.

[0116] refer to Figure 4 , Figure 4 It is a framework schematic diagram of the industrial chain map for constructing the target industrial chain provided by the present invention, which includes: an information acquisition module, an association relationship establishment module, a core industry field identification module, a basic association relationship establishment module, and an association relationship generation module.

[0117] The device for mining multi-factor association relationships of characteristic rural industries mainly includes an information acquisition module, which obtains the industry fields of interest to the subject and service content data based on the fields where the service subject and the production subject are located. The association establishment module is used to establish the basic relationship and other association relationships between the service subject and the generating subject based on the field, subject, service content and other data. The core industry field identification module is used to identify the core field in the field based on the basic association relationship. The basic association establishment module is used to establish the basic association relationship between the field and the subject and service based on the basic association relationship based on the core field. The field subject association generation module is used to generate enterprise-customer association relationships by superimposing the other association relationships on the basis of the basic association relationship. Figure 4 .

[0118] At the same time, the method provided by this invention can be implemented in a terminal environment, which mainly includes a processor, a memory, and a display screen. The memory can be used to store one or more instructions, and the corresponding instructions can be loaded and executed by the processor. The display screen is used to display the user interface of each application.

[0119] In an embodiment of the present invention, a rural characteristic industry map is constructed, and relevant data is collected using web crawler technology to draw an industrial chain map and clarify the connection between each link. The FP-Growth algorithm is used to mine association rules on the processed data, and frequent item sets and potential association rules are quickly identified. By processing the feature vectors, different types of attributes (numerical, categorical, and time) are ensured to be analyzed in a unified feature space. A multi-factor association relationship mining device and storage device for rural characteristic industries are constructed, and the device includes an information acquisition module, an association relationship establishment module, a core industry field identification module, a basic association relationship establishment module, and an association relationship generation module. These modules cooperate with each other, and then through the data storage device, jointly complete the mining of the multi-factor association relationship of the rural characteristic industry map.

[0120] The present invention is based on advanced technologies such as knowledge graphs and big data processing. It takes the characteristic industry field as the center, integrates internal and external related data, deeply analyzes various relationships such as industry fields, subjects, and service contents, and digs out various public and implicit associations related to characteristic industries. It also uses the construction of knowledge graphs to enhance the relevance and comprehensiveness of feature representation, and significantly improves the efficiency and accuracy of data analysis. It effectively solves the challenges brought by the diversity, complexity and personalized needs of rural social service resources, and provides a systematic and intelligent solution. Through the comprehensive analysis of semantic features and associations, it not only optimizes resource allocation, but also improves resource utilization efficiency.

[0121] The industrial chain graph association rule mining device provided by the present invention is described below. The industrial chain graph association rule mining device described below and the industrial chain graph association rule mining method described above can be referenced to each other.

[0122] refer to Figure 5 , Figure 5 It is a structural schematic diagram of the industrial chain graph association rule mining device provided by the present invention.

[0123] The acquisition module 501 is used to acquire a data set of a target industrial chain, wherein the data set includes: a service subject, a production subject, and a service process, and the target industrial chain is a rural characteristic industrial chain; A construction module 502 is used to construct a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; The conversion module 503 is used to perform feature vector conversion on multiple service resource information in the industrial chain map to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; A processing module 504 is used to perform feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; The processing module 504 is further used to discretize the attribute feature vector to obtain an attribute feature representation; A combination module 505, used to combine the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; The mining module 506 is used to mine association rules based on the service subject resource item set corresponding to each service resource information by using a frequent pattern growth algorithm to obtain target association rules between the multiple service resource information.

[0124] Specifically, the above-mentioned industrial chain graph association rule mining device provided by the present invention can implement all the method steps implemented by the above-mentioned industrial chain graph association rule mining method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as the method embodiment in this embodiment will not be described in detail.

[0125] Figure 6 is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as Figure 6 As shown, the electronic device may include: a processor (processor) 610 , a communication interface (Communications Interface) 620 , a memory (memory) 630 and a communication bus 640 , wherein the processor 610 , the communication interface 620 , and the memory 630 communicate with each other through the communication bus 640 . The processor 610 can call the logic instructions in the memory 630 to execute the industrial chain graph association rule mining method, which includes: obtaining a data set of the target industrial chain, wherein the data set includes: service subjects, production subjects and service processes, and the target industrial chain is a rural characteristic industrial chain; constructing a knowledge graph based on the service subjects, production subjects and service processes to obtain an industrial chain graph of the target industrial chain; performing feature vector conversion on multiple service resource information in the industrial chain graph to obtain semantic feature vectors and attribute feature vectors corresponding to each service resource information; performing feature dimension reduction and discretization on the semantic feature vector to obtain a semantic feature representation; performing discretization on the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation with the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; performing association rule mining based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain a target association rule between multiple service resource information.

[0126] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0127] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the industrial chain graph association rule mining method provided by the above methods, and the method includes: obtaining a data set of a target industrial chain, wherein the data set includes: a service subject, a production subject and a service process, and the target industrial chain is a rural characteristic industrial chain; constructing a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; performing feature vector conversion on multiple service resource information in the industrial chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; performing feature dimensionality reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; performing discretization processing on the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; performing association rule mining based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain a target association rule between multiple service resource information.

[0128] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the industrial chain graph association rule mining method provided by the above-mentioned methods, the method comprising: obtaining a data set of a target industrial chain, wherein the data set comprises: a service subject, a production subject and a service process, and the target industrial chain is a rural characteristic industrial chain; constructing a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; performing feature vector conversion on multiple service resource information in the industrial chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; performing discretization processing on the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation with the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; performing association rule mining based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain a target association rule between multiple service resource information.

[0129] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0130] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for mining association rules of industrial chain graphs, characterized in that: include: Acquire a data set of a target industrial chain, wherein the data set includes: a service subject, a production subject, and a service process, and the target industrial chain is a rural characteristic industrial chain; A knowledge graph is constructed based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; Performing feature vector conversion on multiple service resource information in the industrial chain map to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; Performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; Discretizing the attribute feature vector to obtain an attribute feature representation; Combining the semantic feature representation with the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; By using a frequent pattern growth algorithm, association rule mining is performed based on the service subject resource item set corresponding to each service resource information to obtain target association rules between the multiple service resource information.

2. The method for mining association rules of industrial chain graph according to claim 1 is characterized in that: The knowledge graph is constructed based on the service subject, the production subject and the service process to obtain the industrial chain graph of the target industrial chain, including: Extracting knowledge based on the service subject, the production subject and the service process to obtain multiple entities and entity relationships between the multiple entities; Based on a predefined ontology, knowledge integration is performed on the multiple entities and the relationships between the multiple entities to obtain an industrial chain map of the target industrial chain, wherein the predefined ontology is used to define entity categories, entity attributes and entity relationships.

3. The method for mining association rules of industrial chain graph according to claim 1 is characterized in that: The service resource information includes text information, numerical information, type information and time information. The feature vector conversion of the multiple service resource information in the industrial chain map to obtain the semantic feature vector and attribute feature vector corresponding to each service resource information includes: Extracting semantic features from the text information to obtain a semantic feature vector; Normalizing the numerical information to obtain numerical features; Performing type encoding on the type information to obtain type features; Performing time series processing on the time information to obtain time features; The numerical feature, the type feature and the time feature are used as an attribute feature vector of the service resource information.

4. The method for mining association rules of industrial chain graph according to claim 1 is characterized in that: After combining the semantic feature representation with the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information, the method further includes: Determine the ratio between the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set and the number of all service subject resource item sets as the support between the first service subject resource item set and the second service subject resource item set; A ratio between the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set and the number of the first service subject resource item set is determined as a confidence level between the first service subject resource item set and the second service subject resource item set.

5. The method for mining association rules of industrial chain graph according to claim 4 is characterized in that: The method further comprises: The ratio of the confidence level to the support level of the second service subject resource item set is used as the lift between the first service subject resource item set and the second service subject resource item set.

6. The method for mining association rules of industrial chain graph according to claim 1 is characterized in that: The frequent pattern growth algorithm is used to perform association rule mining based on the service subject resource item set corresponding to each service resource information to obtain the target association rule between the multiple service resource information, including: Traverse all service subject resource item sets to obtain the frequency of each service subject resource item set; Determining the support of each service subject resource item set based on the frequency of each service subject resource item set; The service subject resource item set whose support is less than the support threshold is deleted to obtain an item header table; Sort the service subject resource item sets in the item header table in descending order of support and input them into the FP tree; Generate multiple frequent item sets based on the item header table and the FP tree; Determine the confidence of the association rule corresponding to each frequent item set in the multiple frequent item sets, and take the association rule whose confidence is greater than the confidence threshold as a strong association rule; The lifting degree of each of the strong association rules is determined, and the strong association rules whose lifting degrees are greater than a lifting degree threshold are taken as target association rules.

7. An industrial chain graph association rule mining device, characterized in that: include: An acquisition module is used to acquire a data set of a target industrial chain, wherein the data set includes: a service subject, a production subject, and a service process, and the target industrial chain is a rural characteristic industrial chain; A construction module, used to construct a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; A conversion module, used to perform feature vector conversion on multiple service resource information in the industrial chain map to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; A processing module, used for performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; The processing module is further used to discretize the attribute feature vector to obtain an attribute feature representation; A combination module, used to combine the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; The mining module is used to mine association rules based on the service subject resource item set corresponding to each service resource information through a frequent pattern growth algorithm to obtain target association rules between the multiple service resource information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for mining association rules of the industrial chain graph as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for mining association rules of industrial chain graphs as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for mining association rules of industrial chain graphs as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Government affair service recommendation method, device, equipment and computer readable storage medium

    CN113722611A

  • Multi-domain knowledge fusion method based on semantic tree

    CN116542332A

  • Tax analysis service system based on big data

    CN116795923A

  • Large language model knowledge enhancement method and system

    CN117474013A