Industrial chain map association rule mining method and device and storage medium
By constructing a knowledge graph and utilizing a frequent pattern growth algorithm, the timeliness and accuracy of rural industrial chain graphs in existing technologies are addressed, enabling efficient and accurate mining of relationships and providing scientific decision support.
Patent Information
- Application Number
- CN202411928999.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing methods for constructing industrial chain maps are unable to capture the rapidly changing structure and relationships of rural industrial chains in real time, resulting in reduced timeliness and accuracy, and are insufficient in uncovering deeper and more complex relationships.
By acquiring a dataset of rural characteristic industrial chains, data cleaning and feature vector transformation are performed to construct a knowledge graph. The frequent pattern growth algorithm is used to mine association rules, obtaining association rules between service subjects, production subjects, and service processes. This includes feature dimensionality reduction and discretization processing, and the FP-Growth algorithm is used for efficient mining.
It improves the efficiency and accuracy of mining information on industrial chain relationships, enabling precise analysis of potential connections and dependencies between service providers and production entities, and providing scientific decision-making references.
Smart Images

Figure CN119962652B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of knowledge graph, and in particular to an industry chain graph correlation rule mining method and device and a storage medium. BACKGROUND
[0002] An industry chain graph is a visualization tool that describes various links in an industry and their mutual relationships. It integrates enterprise, transaction, industry, information, and heat data to form a rich knowledge graph, presenting various links and their relationships on the industry chain, or their development status, trends, etc. in a graphical manner.
[0003] Rural industries are influenced by market, policy, environment, and other factors, and their industry chain structure and associated relationships are constantly changing. Existing industry chain graph construction methods often fail to capture these changes in real time, resulting in decreased timeliness and accuracy of the graph. SUMMARY
[0004] The present application provides an industry chain graph correlation rule mining method and device and a storage medium to address the shortcomings of existing industry chain graphs that fail to adapt to rapidly changing rural industry environments, thereby improving the efficiency and accuracy of industry chain association relationship information mining.
[0005] The present application provides an industry chain graph correlation rule mining method, including the following steps. A data set of a target industry chain is obtained, wherein the data set includes service subjects, production subjects, and service processes, and the target industry chain is a rural characteristic industry chain. A knowledge graph is constructed based on the service subjects, production subjects, and service processes to obtain an industry chain graph of the target industry chain. Feature vector conversion is performed on multiple service resource information in the industry chain graph to obtain semantic feature vectors and attribute feature vectors corresponding to each service resource information. Feature dimension reduction and discretization processing are performed on the semantic feature vectors to obtain semantic feature representations. Discretization processing is performed on the attribute feature vectors to obtain attribute feature representations. The semantic feature representations and attribute feature representations are combined to obtain service subject resource item sets corresponding to each service resource information. A frequent pattern growth algorithm is used to mine correlation rules based on the service subject resource item sets corresponding to each service resource information to obtain target correlation rules between the multiple service resource information.
[0006] According to the industrial chain graph association rule mining method provided by the application, the knowledge graph is constructed based on the service subject, the production subject and the service process, and the industrial chain graph of the target industrial chain is obtained, which comprises the following steps: knowledge extraction is performed based on the service subject, the production subject and the service process, and a plurality of entities and entity relationships between the plurality of entities are obtained; the plurality of entities and the relationships between the plurality of entities are integrated based on a predefined ontology, and the industrial chain graph of the target industrial chain is obtained, wherein the predefined ontology is used to define entity categories, entity attributes and entity relationships.
[0007] According to the industrial chain graph association rule mining method provided by the application, the service resource information comprises text information, numerical information, type information and time information, the feature vector conversion is performed on the plurality of service resource information in the industrial chain graph, and the semantic feature vector and the attribute feature vector corresponding to each service resource information are obtained, which comprises the following steps: the semantic feature extraction is performed on the text information, and the semantic feature vector is obtained; the numerical information is subjected to numerical normalization, and the numerical feature is obtained; the type information is subjected to type coding, and the type feature is obtained; the time information is subjected to time sequence processing, and the time feature is obtained; and the numerical feature, the type feature and the time feature are taken as the attribute feature vector of the service resource information.
[0008] According to the industrial chain graph association rule mining method provided by the application, after the combination of the semantic feature representation and the attribute feature representation is performed, the method further comprises the following steps: the ratio between the number of simultaneous occurrence of the first service subject resource item set and the second service subject resource item set and the number of all service subject resource item sets is taken as the support degree between the first service subject resource item set and the second service subject resource item set; and the ratio between the number of simultaneous occurrence of the first service subject resource item set and the second service subject resource item set and the number of the first service subject resource item set is taken as the confidence degree between the first service subject resource item set and the second service subject resource item set.
[0009] According to the industrial chain graph association rule mining method provided by the application, the method further comprises the following step: the ratio between the confidence degree and the support degree of the second service subject resource item set is taken as the lift degree between the first service subject resource item set and the second service subject resource item set.
[0010] According to the industrial chain graph association rule mining method provided by the application, the frequent pattern growth algorithm is used to perform association rule mining on the basis of the service subject resource item set corresponding to each service resource information, so that the target association rule between the plurality of service resource information is obtained.
[0011] The application further provides an industrial chain graph association rule mining device, comprising the following modules: an acquisition module, configured to acquire a data set of a target industrial chain, wherein the data set comprises a service subject, a production subject and a service process, and the target industrial chain is a rural characteristic industrial chain; a construction module, configured to perform knowledge graph construction on the basis of the service subject, the production subject and the service process, so as to obtain an industrial chain graph of the target industrial chain; a conversion module, configured to perform feature vector conversion on a plurality of service resource information in the industrial chain graph, so as to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; a processing module, configured to perform feature dimension reduction and discretization processing on the semantic feature vector, so as to obtain a semantic feature representation; the processing module is further configured to perform discretization processing on the attribute feature vector, so as to obtain an attribute feature representation; a combination module, configured to combine the semantic feature representation and the attribute feature representation, so as to obtain a service subject resource item set corresponding to each service resource information; and a mining module, configured to perform association rule mining on the basis of the service subject resource item set corresponding to each service resource information by using a frequent pattern growth algorithm, so as to obtain a target association rule between the plurality of service resource information.
[0012] The application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the industrial chain graph association rule mining method according to any one of the above-mentioned methods when executing the program.
[0013] The application further provides a non-transitory computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the industrial chain graph association rule mining method according to any one of the above-mentioned methods.
[0014] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements any one of the industry chain graph correlation rule mining methods described above.
[0015] The industry chain graph correlation rule mining method, device and storage medium provided by the application can accurately locate and analyze service subjects, production subjects and service processes in a target industry chain by obtaining a data set of a rural characteristic industry chain; a knowledge graph can be constructed based on the obtained data set, which can intuitively show the correlation between subjects in the industry chain and the service process; the service resource information in the industry chain graph can be converted into a feature vector, and dimension reduction and discretization processing can be performed thereon, which can simplify the complexity of the data, ensure that all attributes can be analyzed in a unified feature space, and improve the efficiency of subsequent correlation rule mining; the frequent pattern growth algorithm can be used to mine correlation rules based on the service subject resource item set, which can obtain target correlation rules between multiple service resource information, and reveal the potential correlation and dependency between multiple service resource information; and the efficiency and accuracy of the industry chain correlation relationship information mining can be further improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description one by one. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 is a flowchart of the industry chain graph correlation rule mining method provided by the application.
[0018] Figure 2 is a flowchart of the rural characteristic industry information crawling process provided by the application.
[0019] Figure 3 is a knowledge graph construction flowchart provided by the application.
[0020] Figure 4 is a framework diagram of constructing an industry chain graph of a target industry chain provided by the application.
[0021] Figure 5 is a structural diagram of the industry chain graph correlation rule mining device provided by the application.
[0022] Figure 6 is a physical structure diagram of the electronic device provided by the application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor belong to the protection scope of the present application.
[0024] As a visualization tool, the industry chain map can clearly show the links and their mutual relations in the industry chain, providing strong support for correlation mining. The present application relates to the technical field of industry chain analysis and data mining, aiming to deeply mine the correlation between various elements in the rural characteristic industry by constructing and analyzing the industry chain map.
[0025] Rural industry involves many sub-sectors and complex economic environment, with scattered and uneven quality data. Effective collection, cleaning and integration of these data become the difficulty of constructing accurate industry chain map. The lack of unified data standards and sharing platform leads to serious data island phenomenon, limiting the comprehensiveness and accuracy of data.
[0026] Secondly, rural industry is affected by many factors such as market, policy and environment, and its industry chain structure and correlation are constantly changing. The existing industry chain map construction method often fails to capture these changes in real time, resulting in a decline in the timeliness and accuracy of the map. The lack of dynamic updating mechanism makes it difficult for the map to adapt to the rapidly changing rural industry environment.
[0027] Finally, the existing industry chain map construction method focuses more on displaying the basic structure and main links of the industry chain, but it is still insufficient in mining the deep and complex correlation between multiple elements within and between industries. This limits the potential of the map in revealing the development law of the industry and predicting the trend of the industry.
[0028] The purpose of the present application is to build a perfect data collection, integration and verification mechanism to ensure that data from different channels and different formats can be efficiently and accurately integrated into the industry chain map. By introducing advanced data cleaning and preprocessing technology, the accuracy and usability of the data are improved, laying a solid foundation for subsequent correlation mining. Secondly, by using the FP-Growth algorithm, the method will deeply mine the deep and complex correlation between multiple elements within and between industries in rural characteristic industries. Finally, by constructing the rural characteristic industry multi-element correlation mining device and storage device, the decision-makers are provided with scientific and comprehensive reference.
[0029] Optionally, the industry chain graph association rule mining method of the embodiment of the application can be executed by a server, can also be executed by a terminal device, and can also be executed by the server and the terminal device together.
[0030] Figure 1 is a flowchart of the industry chain graph association rule mining method provided by the application, as shown in Figure 1 , the method comprises the following steps:
[0031] Step 101, obtaining the data set of the target industry chain, wherein the data set comprises a service subject, a production subject and a service process, and the target industry chain is a rural characteristic industry chain.
[0032] In the process of constructing the rural characteristic industry chain graph, the latest industry dynamics, trend analysis and market report and other information are obtained through the relevant rural characteristic information publishing website. Through the research and analysis of the industry chain, the classification of the industry is clarified, which is divided into agriculture, forestry, animal husbandry, fishery and the like. The industry chain is divided into production, processing, storage, transportation, sales and the like.
[0033] Referring to Figure 2 , Figure 2 is a flowchart of the rural characteristic industry information crawling process provided by the application, wherein the process comprises sending a request (GET), server response, parsing data, extracting social service information by using a regular expression, and storing information in a database.
[0034] The network crawler technology is adopted to extract the third-party data source and the socialized service related website, collect the data and information related to the industry chain, including text data, geographic data, image data and the like. The whole crawling process is as shown in Figure 2 .
[0035] In the embodiment of the application, the service subject data set comprises entities or institutions providing various services for the rural characteristic industry chain. For example, service institutions, agricultural cooperatives, agricultural enterprises, scientific research institutions, financial institutions and the like. The service subject data set can comprise: a service subject name, used for accurately recording the full name of the service subject; a service type, used for describing the service type provided by the service subject, such as technical training, market information, financial service, logistics support, policy guidance and the like; a service range, used for explaining the geographical range, industry range or customer group covered by the service subject; and a service effect, used for evaluating the contribution degree of the service subject to the rural characteristic industry chain, such as improving production efficiency, increasing yield, improving product quality and promoting sales and the like.
[0036] The production subject dataset mainly covers actual producers in the rural characteristic industry chain, including small farmers, family farms, agricultural enterprises, etc. The production subject dataset includes: production subject name, used to accurately record the full name or abbreviation of the production subject. Production type, used to describe the type of agricultural production engaged in by the production subject, such as planting, breeding, and agricultural product processing. Production scale, used to describe the production scale of the production subject, such as planting area, breeding quantity, and processing capacity. Product type, used to list the main types of agricultural products produced by the production subject. Yield and sales, used to record the actual yield and sales of the production subject, as well as the sales channel and market distribution.
[0037] The service process dataset mainly includes various service processes provided by service subjects for production subjects in the rural characteristic industry chain, including pre-production, production, and post-production. The service process dataset can include: service link, used to clearly describe the link to which the service process belongs, such as pre-production preparation, planting / breeding process, product processing, quality detection, packaging and transportation, and market sales. Service content, used to describe the specific service content provided by the service subject in each service link, such as the specific content of technical training, the provision method of market information, and the type and conditions of financial services. Service time, used to record the specific time or time range of service provision. Service effect, used to evaluate the degree of help of the service process to the production subject, such as improving production efficiency, reducing cost, improving product quality, and promoting sales.
[0038] Step 102, constructing a knowledge graph based on the service subject, production subject, and service process, to obtain an industry chain graph of the target industry chain.
[0039] In the embodiments of the present application, the collected data sets are cleaned, de-duplicated, formatted, and other operations to ensure the accuracy and consistency of the data.
[0040] Using natural language processing techniques, such as named entity recognition (NER), to automatically identify entities with specific meanings from text data, such as service subject name, production subject name, and service name.
[0041] Extracting relationships between entities from text, such as the service relationship between service subjects and production subjects, and the production relationship between production subjects and products. Relationship extraction can use rule-based methods, deep learning-based methods, etc.
[0042] Collecting attribute information of specific entities, such as service type and service range of service subjects, production scale and technical level of production subjects, and product type and yield. Attribute extraction helps to enrich the content of the knowledge graph, making it more complete and accurate.
[0043] The extracted entity, relationship and attribute information are integrated to eliminate contradictions and ambiguities. Through entity linking and knowledge merging, information from different sources is fused into a unified knowledge graph.
[0044] The constructed knowledge graph is stored using a graph database (such as Neo4j) or an RDF specification storage format. Graph databases have significantly improved efficiency in association queries compared to traditional relational data storage methods, and are therefore more suitable for storing complex knowledge graphs.
[0045] The rural characteristic industry takes the field, the main body and the service as the point, wherein the field is divided into characteristic agricultural products, regional characteristics, local resources, characteristic agricultural product processing, leisure agriculture, rural tourism and the like. The main body is divided into main body name and type (supply and marketing cooperative, farmer professional cooperative, service type joint society, rural collective economic organization, leading enterprise, professional service company, family farm, individual business, individual demonstration household, other), operating state, region, detailed address, contact name, telephone, service main body introduction, service project and the like. At the same time, the service main body is also associated with agricultural materials, agricultural techniques, products, markets, credit and the like. The service data refers to service process data, mainly including agricultural production service, agricultural material supply, agricultural technique service, post-processing service, agricultural machinery leasing, agricultural financial service. The service data makes the service main body and the production main body associated, and according to the service main body, the production main body and the service process relationship, a rural socialized service resource knowledge graph framework is constructed,
[0046] According to the industry chain graph atlas association rule mining method provided by the application, the knowledge graph is constructed based on the service main body, the production main body and the service process, and the industry chain graph atlas of the target industry chain is obtained, including:
[0047] Based on the service main body, the production main body and the service process, knowledge extraction is carried out, and a plurality of entities and entity relationships between the plurality of entities are obtained.
[0048] Based on the pre-defined ontology, the knowledge of the plurality of entities and the relationships between the plurality of entities is integrated, and the industry chain graph atlas of the target industry chain is obtained, wherein the pre-defined ontology is used to define the entity category, the entity attribute and the entity relationship.
[0049] Here, knowledge extraction includes automatically identifying and extracting entities such as service subjects, production subjects, etc. from text (data set) using natural language processing (NLP) techniques such as named entity recognition (NER). Identify and extract attribute information related to these entities, such as company name, address, size, main business, etc. Analyze the relationship between entities in the text and extract their service relationship, production relationship, etc. Relationships can include supply chain relationships (such as supplier-manufacturer relationships), cooperation relationships (such as manufacturer-distributor relationships), service relationships (such as service provider-customer relationships), etc.
[0050] Knowledge integration includes pre-defining an ontology to define entity categories, entity attributes, and entity relationships. The ontology should cover all target entity categories and relationship types in the industry chain to ensure the completeness and accuracy of knowledge integration. Map the extracted entities and relationships to the entity categories and relationship types in the ontology. For entities or relationships that cannot be directly mapped, new categories or types can be created and added to the ontology. Integrate the mapped entities and relationships into a unified knowledge graph. Use graph databases or RDF storage formats to store the knowledge graph for efficient querying and analysis.
[0051] Reference Figure 3 , Figure 3 is a knowledge graph construction flowchart provided by the present application, which includes: entities, relationships, attributes; ontology, relationships, attributes; knowledge extraction, knowledge integration, knowledge storage, update, control; structured data, unstructured data; structured triples and knowledge base.
[0052] In an embodiment of the present application, the collected data is used to draw an industry chain map to clarify the relationship between each link. First, clearly define the ontology to define entity categories, attributes and relationships. Mainly includes four parts: knowledge extraction, knowledge integration, knowledge storage, knowledge update maintenance and quality control. Among them, knowledge extraction needs natural language processing technology to identify entities and relationships between entities in the text. Knowledge integration integrates the extracted knowledge into a unified framework in a proper way, and needs to define an ontology to determine the categories, attributes and relationships in the knowledge graph. Knowledge storage effectively stores and indexes the knowledge for subsequent querying and analysis, which can be based on open source graph database storage. The knowledge graph construction flowchart is shown in Figure 3 .
[0053] Through the embodiment of the present application, through knowledge extraction, key information related to service subjects, production subjects and service processes can be efficiently extracted from a large amount of text data; through the pre-defined ontology, multiple entities and their relationships can be integrated into a unified knowledge system; the ontology-based knowledge integration supports semantic reasoning, and new relationships implied in the knowledge graph can be deduced, further enriching and perfecting the industry chain graph.
[0054] In step 103, the feature vector conversion is performed on the multiple service resource information in the industry chain graph, to obtain the semantic feature vector and the attribute feature vector corresponding to each service resource information.
[0055] In the embodiment of the present application, the socialized service resource information includes long text information such as service crops, service projects and service subject introductions, and each item is converted into a semantic-level feature vector through a pre-set semantic feature extraction representation model. The attribute information such as service crop type, location and time is subjected to feature engineering processing. This includes appropriate processing of different types of attributes such as numerical, categorical and time, numerical normalization, category coding and time series processing, to ensure that all attributes can be analyzed in a unified feature space.
[0056] According to the industry chain graph association rule mining method provided by the present application, the service resource information includes text information, numerical information, type information and time information, the feature vector conversion is performed on the multiple service resource information in the industry chain graph, to obtain the semantic feature vector and the attribute feature vector corresponding to each service resource information, including:
[0057] The semantic feature extraction is performed on the text information to obtain the semantic feature vector;
[0058] The numerical normalization is performed on the numerical information to obtain the numerical feature;
[0059] The type coding is performed on the type information to obtain the type feature;
[0060] The time series processing is performed on the time information to obtain the time feature;
[0061] The numerical feature, the type feature and the time feature are taken as the attribute feature vector of the service resource information.
[0062] In the embodiment of the present application, the semantic feature extraction is performed on the text information through the natural language processing technology (NLP), for example, the text is cut into words or phrases using a word segmentation tool (such as jieba); each word is converted into a high-dimensional vector using a pre-trained word vector model (such as Word2Vec, BERT); and the vectors of all words are aggregated (such as average, weighted average, TF-IDF weighting, etc.) to obtain the semantic feature vector of the entire text.
[0063] Numerical information can have different dimensions and ranges, and normalization is needed for unified processing; for example, through min-max normalization, the values are scaled to a specified range (such as 0 to 1); through Z-score normalization, the data is scaled according to the mean and standard deviation of the data, so that the data conforms to the standard normal distribution.
[0064] Type information is usually represented as discrete category labels, which can be converted to numerical features through encoding. For example, through One-Hot Encoding, a separate binary column is created for each category, with a value of 1 if the data belongs to the category and 0 otherwise. Through Label Encoding, each category is mapped to a unique integer. Through Target Encoding, the category label is encoded according to the target variable mean, which is commonly used to process classification data with target information.
[0065] Time information usually contains dates, timestamps, etc., and can be extracted through time series analysis; for example, time is decomposed into year, month, day, hour, minute, etc. components; calculate the difference between the current time and a certain reference time (such as the start time of the event); extract periodic features of time, such as day of the week, season, etc.; use time series analysis techniques (such as autoregressive models, moving averages, etc.) to extract features.
[0066] Through the embodiments of the present application, through semantic feature extraction, text information is converted into high-dimensional vectors, which can capture the semantic relationship between words and context information, thereby enhancing the machine's understanding of the content of the text; numerical normalization can eliminate the dimensional differences between different numerical features; type encoding converts discrete category information into numerical features, so that category information can participate in subsequent mathematical operations or model training; time series processing can capture trends and periodic changes in time information, providing an important basis for subsequent prediction and analysis.
[0067] Step 104, the semantic feature vector is processed by feature dimension reduction and discretization to obtain a semantic feature representation.
[0068] Feature dimension reduction is to reduce the dimension of the data while preserving as much important information as possible. For semantic feature vectors, dimension reduction helps to reduce redundant information and improve computational efficiency.
[0069] For example, according to the principal component analysis method, a set of variables that may be correlated is converted into a set of linearly uncorrelated variables, i.e. principal components, through orthogonal transformation. Applying principal component analysis to semantic feature vectors can remove redundant features and retain the most representative components.
[0070] The pre-trained word embedding model (such as Word2Vec, GloVe, etc.) is used to map the words to a low-dimensional vector space, and the semantic information of the words is represented by these vectors. In the semantic feature vector, the high-dimensional one-hot encoding can be converted into a low-dimensional dense vector using the word embedding model, thereby realizing dimension reduction.
[0071] Discretization processing is the conversion of continuous data into discrete data representation, which is usually used to simplify data representation, improve computational efficiency or meet the requirements of specific algorithms. For semantic feature vectors, discretization processing may involve converting continuous vector values into discrete category labels or binary values.
[0072] For example, by threshold method, on the semantic feature vector, a threshold can be set according to the value range of each feature to convert the continuous feature value into a discrete category label; by clustering method, K-means, hierarchical clustering and other algorithms are used to cluster the semantic feature vector, and the continuous feature value is converted into a discrete cluster label.
[0073] After feature dimension reduction and discretization processing, a simplified semantic feature representation can be obtained. The semantic feature representation not only reduces the dimension and complexity of the data, but also retains important information in the original data, which is helpful for subsequent analysis and modeling tasks.
[0074] Step 105, discretizing the attribute feature vector to obtain an attribute feature representation.
[0075] In the embodiments of the present application, the attribute feature vector is discretized to convert continuous numerical features into discrete category labels or binary values, thereby simplifying data representation and improving computational efficiency.
[0076] According to the pre-set discretization method and determined parameters, the attribute feature vector is discretized to convert continuous numerical features into discrete category labels or binary values; after discretization processing, the attribute feature vector is converted into a vector composed of discrete category labels or binary values, i.e. the attribute feature representation.
[0077] Step 106, combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information.
[0078] In the embodiments of the present application, feature dimension reduction and discretization processing are performed on the semantic feature vector, and discretization processing is performed on other attribute features, and the two are combined to form a service subject resource item set. After completing the feature engineering, the FP-Growth algorithm is used to mine the association rules of the processed data. The FP-growth algorithm is an extension of the Apriori algorithm, and since the algorithm only scans the data set twice, it has very high efficiency in the discovery process of frequent item sets.
[0079] In some embodiments, the attributes "service organization name", "service item", "service location", "service time", "service crop type", "service subject type", "service subject introduction", "annual turnover", "establishment length (years)" and the like are selected to form an attribute set. Among them, the annual turnover and the establishment length (years) are in numerical form, and the numerical values are discretized and converted according to the probability distribution. The category type attribute is formed into an item set according to each value of each category attribute, and is encoded to represent, for example, the production service type includes pruning, fertilization, pesticide spraying, weeding, harvesting, full management, and base.
[0080] The single service subject resource item set after feature engineering processing is in the form of:
[0081] {
[0082] "service item semantic vector": [0.1, 0.3, 0.5,..., 0.2],
[0083] "service subject introduction semantic vector": [0.2, 0.4, 0.1,..., 0.3],
[0084] "service industry type": 1,
[0085] "service location": 2,
[0086] "service time_year": 2024,
[0087] "service time_month": 8,
[0088] "service time_day": 16,
[0089] "annual turnover": "100-600 thousand",
[0090] "establishment length": "1-5 years"
[0091] }
[0092] According to the industry chain graph association rule mining method provided by the application, after the semantic feature representation and the attribute feature representation are combined to obtain the service subject resource item set corresponding to each service resource information, the above method further comprises:
[0093] Determine the ratio between the number of simultaneous occurrence of the first service subject resource item set and the second service subject resource item set and the number of all service subject resource item sets as the support degree between the first service subject resource item set and the second service subject resource item set;
[0094] The ratio between the number of the first service subject resource item set and the second service subject resource item set appearing simultaneously and the number of the first service subject resource item set is determined as the confidence between the first service subject resource item set and the second service subject resource item set.
[0095] The support is the ratio between the number of the first service subject resource item set and the second service subject resource item set appearing simultaneously and the total number of all service subject resource item sets. The ratio reflects the frequency of the two item sets appearing together in the total data set.
[0096] The higher the support is, the more the two service subject resource item sets appear together in the data set, and the stronger the correlation between them can be.
[0097] The confidence is the probability of the second service subject resource item set appearing in the case of the first service subject resource item set appearing. The probability reflects the possibility of the second service subject resource item set appearing under a specific condition (i.e. the first service subject resource item set appearing).
[0098] The higher the confidence is, the greater the possibility of the second service subject resource item set appearing in the case of the first service subject resource item set appearing, i.e. the higher the reliability of the association rule.
[0099] In the embodiment of the present application, all user data sets (i.e. all service subject resource item sets) are I, the support of the association rule A->B is the probability of A and B appearing simultaneously, i.e. the ratio between the number of data of A and B appearing simultaneously and the total number of records, as shown in formula (1):
[0100] (1)
[0101] wherein A represents the first service subject resource item set, B represents the second service subject resource item set, represents the support between the first service subject resource item set and the second service subject resource item set; represents the number of the first service subject resource item set and the second service subject resource item set appearing simultaneously, represents the number of all service subject resource item sets.
[0102] The confidence is the probability of B appearing in the case of A appearing, i.e. the ratio between the number of data of A and B appearing simultaneously and the number of records of A appearing, as shown in formula (2):
[0103] (2)
[0104] wherein, Confidence represents the confidence between the first service subject resource item set and the second service subject resource item set, A represents the first service subject resource item set, B represents the second service subject resource item set, Lift represents the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set, Num represents the number of the first service subject resource item set.
[0105] Through the embodiments of the present application, the frequent item set and the strong association rule can be quickly screened out by calculating the support and the confidence, so as to reduce the calculation amount and the time cost in the data mining process; the association between the service subject resource item sets can be more intuitively displayed by calculating the support and the confidence, so as to enhance the interpretability and the readability of the data.
[0106] According to the industrial chain graph atlas association rule mining method provided by the present application, the method further comprises:
[0107] The ratio of the confidence and the support of the second service subject resource item set is taken as the lift between the first service subject resource item set and the second service subject resource item set.
[0108] In the embodiments of the present application, the lift represents whether the occurrence of A has a positive or negative effect on the occurrence of B, that is, The ratio of the confidence of A and the support of B, as shown in formula (3), needs to meet the condition that the strong association rule with the lift greater than 3 is an effective association rule in application.
[0109] (3)
[0110] Wherein, Lift represents the lift between the first service subject resource item set and the second service subject resource item set, Confidence represents the confidence between the first service subject resource item set and the second service subject resource item set, Support represents the support of the second service subject resource item set (that is, the frequency of the occurrence of the second service subject resource item set).
[0111] Here, for the first service subject resource item set and the second service subject resource item set, the support (Support) represents the frequency of the simultaneous occurrence of the two item sets. The confidence (Confidence) represents the probability of the occurrence of one item set under the condition that the other item set also occurs. The lift (Lift) represents the ratio of the confidence and the support, and is used to measure the influence degree of the occurrence of one item set on the occurrence of the other item set.
[0112] The lift is used to represent whether the occurrence of A (the first service subject resource item set) has an independent, positive or negative correlation with the occurrence of B (the second service subject resource item set). If the lift is greater than 1, it means that the occurrence of A has a positive impact on the occurrence of B (i.e., A and B are positively correlated); if the lift is equal to 1, it means that A and B are independent; if the lift is less than 1, it means that the occurrence of A has a negative impact on the occurrence of B (i.e., A and B are negatively correlated).
[0113] Through the embodiments of the present application, the lift is an important indicator for measuring whether the occurrence of an item set (or rule) is independent of another item set. By calculating the lift, it can be determined whether the association between two item sets is random, independent, or there is some kind of dependency relationship.
[0114] In step 107, based on the service subject resource item set corresponding to each service resource information, the frequent pattern growth algorithm is used to perform association rule mining, and target association rules between multiple service resource information are obtained.
[0115] Through the frequent pattern growth algorithm (FPGrowth) for association rule mining, based on the service subject resource item set corresponding to each service resource information, the target association rules between multiple service resource information can be efficiently identified.
[0116] The service subject resource item set is converted into a form suitable for FPGrowth algorithm processing, such as a transaction database.
[0117] The transaction database is traversed, and the frequency of occurrence of each item set (or item in the service subject resource item set) is calculated. According to the set minimum support threshold, the frequent items are filtered out. Based on the frequent items, a frequent item set tree (FP-Tree) is constructed. The FP-Tree is a compressed data structure used to store frequent item sets and their association relationships.
[0118] Starting from the root node of the FP-Tree, all frequent patterns (i.e., frequent item sets) are recursively mined. For each frequent item, a corresponding conditional FP-Tree is constructed to further mine the frequent patterns containing the item.
[0119] For each frequent pattern, the confidence of its corresponding association rule is calculated. According to the set minimum confidence threshold, the association rules that meet the conditions are filtered out. The filtered association rules (i.e., target association rules) are sorted and optimized to make it easier to understand and apply.
[0120] According to the industry chain graph association rule mining method provided by the application, the frequent pattern growth algorithm is used to perform association rule mining on each service subject resource item set corresponding to service resource information, and target association rules between multiple service resource information are obtained, including:
[0121] All service subject resource item sets are traversed to obtain the frequency of each service subject resource item set;
[0122] The support degree of each service subject resource item set is determined based on the frequency of each service subject resource item set;
[0123] The service subject resource item sets with a support degree less than a support degree threshold are deleted to obtain a header table;
[0124] The service subject resource item sets in the header table are sorted in descending order of support degree and input into an FP tree;
[0125] Multiple frequent item sets are generated based on the header table and the FP tree;
[0126] The confidence of the association rules corresponding to each frequent item set in the multiple frequent item sets is determined, and the association rules with a confidence greater than a confidence threshold are taken as strong association rules;
[0127] The lift of each strong association rule is determined, and the strong association rules with a lift greater than a lift threshold are taken as target association rules.
[0128] In the embodiment of the application, the steps of the frequent pattern growth (FP-growth) algorithm are as follows:
[0129] Step 1, traverse the data set, and calculate the frequency according to each possible value of each attribute.
[0130] In this step, the algorithm traverses the entire data set, counts each value of each attribute, and calculates the frequency of each item (i.e., the number of times the item appears in the data set).
[0131] Step 2, calculate the minimum data item threshold according to the support degree, delete the data items less than the threshold, and arrange them in descending order in the header table.
[0132] A support degree threshold is set (usually determined according to experience or business requirements), and then the items with a frequency less than the threshold are deleted, because these items are considered to be infrequent.
[0133] The remaining frequent items are placed in the header table (Header Table) and arranged in descending order according to their frequencies (support degrees). The header table is used for quick searching and accessing frequent items.
[0134] Step 3: Traverse the dataset and remove the itemsets that have been removed in Step 2. Insert the data into the FP-tree in order of support.
[0135] Again, traverse the dataset, but this time only consider those frequent items that are in the item header table. Insert the transaction data into a frequent pattern tree (FP-Tree) based on the frequency (support) of items and their order of appearance in transactions.
[0136] The FP-Tree is a special prefix tree used to efficiently store and find frequent itemsets.
[0137] Step 4: Traverse the item header table in reverse order to find the corresponding conditional pattern bases, thus obtaining the frequent itemsets.
[0138] Traverse each frequent item in the item header table in order of frequency (support) from high to low. For each frequent item, find all its prefix paths in the FP-Tree (these paths form the conditional pattern bases of the item). Use these conditional pattern bases to build a conditional FP-Tree and recursively mine all frequent itemsets containing the current frequent item.
[0139] Step 5: For each non-empty subset of the frequent itemsets, calculate the confidence and obtain strong association rules (i.e. target association rules) based on the confidence threshold.
[0140] For each mined frequent itemset, generate all its possible non-empty subsets (these subsets form the antecedents of association rules). Calculate the confidence of each association rule (i.e. the probability of the consequent appearing given the antecedent). According to the set confidence threshold, filter out those association rules with confidence higher than the threshold, and these rules are considered strong association rules.
[0141] Step 6: For each strong association rule, calculate the lift and output the association rule if it meets the threshold.
[0142] In association rule mining, lift is an important indicator of rule effectiveness, reflecting the independence between the antecedent and the consequent of the rule.
[0143] For the antecedent and the consequent of the strong association rule, calculate their support in the dataset. At the same time, calculate the support of the antecedent and the consequent appearing together, which is usually calculated when mining frequent itemsets.
[0144] Confidence is another measure of rule strength, indicating the probability of the consequent appearing given the antecedent. For example, confidence = (support of antecedent & consequent) / support of antecedent.
[0145] Lift is the ratio of confidence to support of the consequent, which reflects the influence of the antecedent on the probability of the consequent. Lift = Confidence / Support of the consequent.
[0146] According to the set lift threshold, those rules with lift higher than the threshold are selected.
[0147] If the lift of a rule is greater than 1, it means that there is a positive correlation between the antecedent and the consequent, i.e. the occurrence of the antecedent increases the probability of the occurrence of the consequent.
[0148] If the lift of a rule is equal to 1, it means that the antecedent and the consequent are independent, i.e. the occurrence of the antecedent has no effect on the probability of the occurrence of the consequent.
[0149] If the lift of a rule is less than 1, it means that there is a negative correlation between the antecedent and the consequent, i.e. the occurrence of the antecedent decreases the probability of the occurrence of the consequent.
[0150] For the association rules that meet the lift threshold, they are output as the final effective rules.
[0151] Through the embodiments of the present application, by calculating the lift and selecting the association rules, the strong association rules mined can be further refined, and only the target association rules that can reflect the potential association in the data set are retained.
[0152] Through this method, the potential association relationship between different service resources on the rural social service platform can be revealed. For example, it can be found that certain service items are more likely to be used jointly in a specific location and time period.
[0153] Reference Figure 4 , Figure 4 is a framework schematic diagram of constructing an industry chain map of a target industry chain provided by the present application, which comprises: an information acquisition module, an association relationship establishment module, a core industry field identification module, a basic association relationship establishment module, and an association relationship generation module.
[0154] The rural characteristic industry multi-element association relationship mining device mainly comprises an information acquisition module, a service subject and a production subject based on the field, an interested industry field of the subject, and service content data. The association relationship establishment module is used to establish the basic relationship and other association relationships between the service subject and the production subject according to the field, the subject, the service content and other data. The core industry field identification module is used to identify the core field in the field according to the basic association relationship. The basic association relationship establishment module is used to establish the basic association relationship between the field and the subject and the service according to the basic association relationship based on the core field. The field-subject association relationship generation module is used to generate the enterprise customer association relationship by superimposing the other association relationships on the basis of the basic association relationship. For example Figure 4 .
[0155] Meanwhile, the method provided by the application can be implemented in a terminal environment, mainly including a processor, a memory and a display screen.
[0156] In the embodiment of the application, the rural characteristic industry graph is constructed, relevant data is collected by using a web crawler technology, an industry chain graph is drawn, and the connection between each link is clear. The FP-Growth algorithm is used to mine the association rules of the processed data, to quickly identify frequent item sets and potential association rules, and through the processing of the feature vectors, it is ensured that different types of attributes (numerical type, category type, time type) are analyzed in a unified feature space. A rural characteristic industry multi-element association relationship mining device and a storage device are constructed, the device includes an information acquisition module, an association relationship establishment module, a core industry field identification module, a basic association relationship establishment module and an association relationship generation module. These modules cooperate with each other, and through the data storage device, the mining work of the multi-element association relationship of the rural characteristic industry graph is completed.
[0157] The application is based on advanced technologies such as knowledge graph and big data processing, takes the characteristic industry field as the center, integrates internal and external related data, deeply analyzes various relationships such as industry field, subject and service content, mines various public and implicit association relationships related to the characteristic industry, and with the help of the construction of the knowledge graph, the association and comprehensiveness of the feature representation are enhanced, and the efficiency and accuracy of data analysis are significantly improved. The challenges brought by the diversity, complexity and personalized needs of rural social service resources are effectively solved, and a systematic and intelligent solution is provided. Through comprehensive analysis of semantic features and association relationships, resource allocation is optimized, and resource utilization efficiency is improved.
[0158] The industry chain graph association rule mining device provided by the application is described below, and the industry chain graph association rule mining device described below can be correspondingly referred to the industry chain graph association rule mining method described above.
[0159] Reference Figure 5 , Figure 5 FIG. 1 is a structural schematic diagram of the industry chain graph association rule mining device provided by the application.
[0160] The acquisition module 501 is used for acquiring a data set of a target industry chain, wherein the data set includes a service subject, a production subject and a service process, and the target industry chain is a rural characteristic industry chain.
[0161] The construction module 502 is configured to construct a knowledge graph based on the service subject, the production subject and the service process, and obtain an industry chain graph of the target industry chain.
[0162] The conversion module 503 is configured to convert feature vectors of a plurality of service resource information in the industry chain graph, and obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information.
[0163] The processing module 504 is configured to perform feature dimension reduction and discretization processing on the semantic feature vector, and obtain a semantic feature representation.
[0164] The processing module 504 is further configured to perform discretization processing on the attribute feature vector, and obtain an attribute feature representation.
[0165] The combination module 505 is configured to combine the semantic feature representation and the attribute feature representation, and obtain a service subject resource item set corresponding to each service resource information.
[0166] The mining module 506 is configured to perform association rule mining on the service subject resource item set corresponding to each service resource information based on a frequent pattern growth algorithm, and obtain a target association rule between the plurality of service resource information.
[0167] Specifically, the above-mentioned industry chain graph association rule mining device provided by the present application can realize all the method steps realized by the above-mentioned industry chain graph association rule mining method embodiment, and can achieve the same technical effect. Here, the same parts and beneficial effects in this embodiment as the method embodiment will not be described in detail.
[0168] Figure 6 is the entity structure schematic diagram of the electronic equipment provided by the present application, like Figure 6As shown, the electronic device can include a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the communications bus 640. The processor 610 can invoke a logic instruction in the memory 630 to execute an industry chain graph correlation rule mining method, which includes: acquiring a data set of a target industry chain, wherein the data set includes a service subject, a production subject, and a service process, and the target industry chain is a rural characteristic industry chain; performing knowledge graph construction based on the service subject, the production subject, and the service process to obtain an industry chain graph of the target industry chain; performing feature vector conversion on a plurality of service resource information in the industry chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; performing discretization processing on the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; and performing correlation rule mining on the service subject resource item set corresponding to each service resource information based on a frequent pattern growth algorithm to obtain a target correlation rule between the plurality of service resource information.
[0169] In addition, the logic instruction in the memory 630 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0170] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer-readable storage medium, and the computer program can be executed by a processor to enable a computer to perform the industry chain graph association rule mining method provided by the above-mentioned methods, which comprises: obtaining a data set of a target industry chain, wherein the data set comprises a service subject, a production subject and a service process, and the target industry chain is a rural characteristic industry chain; constructing a knowledge graph based on the service subject, the production subject and the service process to obtain an industry chain graph of the target industry chain; performing feature vector conversion on a plurality of service resource information in the industry chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; performing discretization processing on the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; and performing association rule mining on the service subject resource item set corresponding to each service resource information based on a frequent pattern growth algorithm to obtain a target association rule between the plurality of service resource information.
[0171] In still another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program can be executed by a processor to implement the industry chain graph association rule mining method provided by the above-mentioned methods, which comprises: obtaining a data set of a target industry chain, wherein the data set comprises a service subject, a production subject and a service process, and the target industry chain is a rural characteristic industry chain; constructing a knowledge graph based on the service subject, the production subject and the service process to obtain an industry chain graph of the target industry chain; performing feature vector conversion on a plurality of service resource information in the industry chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; performing discretization processing on the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; and performing association rule mining on the service subject resource item set corresponding to each service resource information based on a frequent pattern growth algorithm to obtain a target association rule between the plurality of service resource information.
[0172] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0174] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An industry chain graph association rule mining method, characterized in that, The method comprises the following steps: acquiring a data set of a target industrial chain, wherein the data set comprises a service subject, a production subject and a service process, and the target industrial chain is a rural characteristic industrial chain; constructing a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain; performing feature vector conversion on a plurality of service resource information in the industrial chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; performing feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; performing discretization processing on the attribute feature vector to obtain an attribute feature representation; combining the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; performing association rule mining on the service subject resource item set corresponding to each service resource information based on a frequent pattern growth algorithm to obtain a target association rule between the plurality of service resource information, wherein the target association rule comprises: traversing all service subject resource item sets to obtain a frequency of each service subject resource item set; determining a support degree of each service subject resource item set based on the frequency of each service subject resource item set; deleting service subject resource item sets with a support degree less than a support degree threshold to obtain a header table; sorting service subject resource item sets in the header table in descending order of support degree and inputting the service subject resource item sets into an FP tree; generating a plurality of frequent item sets based on the header table and the FP tree; determining a confidence degree of an association rule corresponding to each frequent item set in the plurality of frequent item sets, and regarding an association rule with a confidence degree greater than a confidence degree threshold as a strong association rule; determining a lift degree of each strong association rule, and regarding a strong association rule with a lift degree greater than a lift degree threshold as a target association rule.
2. The industry chain map association rule mining method according to claim 1, characterized in that, The method of constructing a knowledge graph based on the service subject, the production subject and the service process to obtain an industrial chain graph of the target industrial chain comprises the following steps: performing knowledge extraction based on the service subject, the production subject and the service process to obtain a plurality of entities and entity relationships between the plurality of entities; performing knowledge integration on the plurality of entities and the relationships between the plurality of entities based on a pre-defined ontology to obtain an industrial chain graph of the target industrial chain, wherein the pre-defined ontology is used to define entity categories, entity attributes and entity relationships.
3. The industry chain map association rule mining method according to claim 1, characterized in that, The service resource information comprises text information, numerical information, type information and time information, and the method of performing feature vector conversion on a plurality of service resource information in the industrial chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information comprises the following steps: performing semantic feature extraction on the text information to obtain a semantic feature vector; performing numerical normalization on the numerical information to obtain a numerical feature; performing type coding on the type information to obtain a type feature; performing time series processing on the time information to obtain a time feature; The numerical feature, the type feature, and the time feature are taken as an attribute feature vector of the service resource information.
4. The industry chain map association rule mining method according to claim 1, characterized in that, After the combination of the semantic feature representation and the attribute feature representation, the method further comprises: determining a ratio between a number of simultaneous occurrences of a first service subject resource item set and a second service subject resource item set and a number of all service subject resource item sets as a support degree between the first service subject resource item set and the second service subject resource item set; determining a ratio between the number of simultaneous occurrences of the first service subject resource item set and the second service subject resource item set and a number of the first service subject resource item set as a confidence degree between the first service subject resource item set and the second service subject resource item set.
5. The industry chain map association rule mining method according to claim 4, characterized in that, The method further comprises: taking a ratio between the confidence degree and the support degree of the second service subject resource item set as a promotion degree between the first service subject resource item set and the second service subject resource item set.
6. An industrial chain map association rule mining device characterized by comprising: comprises: an acquisition module, configured to acquire a data set of a target industrial chain, wherein the data set comprises a service subject, a production subject, and a service process, and the target industrial chain is a rural characteristic industrial chain; a construction module, configured to construct a knowledge graph based on the service subject, the production subject, and the service process to obtain an industrial chain graph of the target industrial chain; a conversion module, configured to perform feature vector conversion on a plurality of service resource information in the industrial chain graph to obtain a semantic feature vector and an attribute feature vector corresponding to each service resource information; a processing module, configured to perform feature dimension reduction and discretization processing on the semantic feature vector to obtain a semantic feature representation; the processing module is further configured to perform discretization processing on the attribute feature vector to obtain an attribute feature representation; a combination module, configured to combine the semantic feature representation and the attribute feature representation to obtain a service subject resource item set corresponding to each service resource information; a mining module, configured to perform association rule mining on the service subject resource item set corresponding to each service resource information based on a frequent pattern growth algorithm to obtain a target association rule between the plurality of service resource information, wherein the target association rule comprises: traversing all service subject resource item sets to obtain a frequency of each service subject resource item set; determining a support degree of each service subject resource item set based on the frequency of the service subject resource item set; deleting service subject resource item sets with a support degree less than a support degree threshold to obtain a header table; sorting service subject resource item sets in the header table in descending order of support degree and inputting the service subject resource item sets into an FP tree; generating a plurality of frequent item sets based on the header table and the FP tree; determining a confidence degree of an association rule corresponding to each frequent item set in the plurality of frequent item sets, and taking an association rule with a confidence degree greater than a confidence degree threshold as a strong association rule; determining a promotion degree of each strong association rule, and taking a strong association rule with a promotion degree greater than a promotion degree threshold as a target association rule.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor implements the industry chain graph association rule mining method according to any one of claims 1 to 5 when executing the computer program.
8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the industry chain graph association rule mining method according to any one of claims 1 to 5.
9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the industry chain graph association rule mining method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Government affair service recommendation method, device, equipment and computer readable storage medium
CN113722611A
Multi-domain knowledge fusion method based on semantic tree
CN116542332A