Method and device for acquiring industry-related data, electronic equipment and storage medium
By analyzing the combination of industry chain maps and supply chain maps, and using a classification tree model and multi-dimensional parameter calculations, the problem of insufficient accuracy in traditional assessment methods is solved, and a more accurate assessment of industry structure is achieved.
Patent Information
- Application Number
- CN202511933762.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional industrial structure assessment methods rely on macroeconomic statistics, which cannot accurately indicate the internal situation of an industry, resulting in low accuracy and reliability of the assessment results.
By analyzing the degree of alignment between the industry chain map and the supply chain map, a classification tree model is used to map the business entity nodes in the supply chain to the corresponding industrial links in the industry chain. Quantitative and qualitative analysis parameters are combined to calculate node characteristic values and business correlation values, and industry-related data are comprehensively evaluated.
It improves the accuracy and reliability of industrial structure assessment within the region, enabling a comprehensive and objective reflection of industrial operation status and development indicators.
Smart Images

Figure CN121835860A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a method and device for obtaining industry-related data, electronic equipment and storage medium. BACKGROUND
[0002] In the current social and economic background of pursuing sustainable development, the industry-related data of the industrial structure can provide important data support for formulating regional economic development plans. In traditional methods of evaluating the industrial structure, macro statistical data and economic indicators such as gross local product and employment rate are relied on. Although these data can indicate the economic situation of a region in a macro direction, they cannot accurately indicate the specific situation and coordination state within the industry, and the accuracy and reliability of the obtained industrial evaluation results are low. SUMMARY
[0003] Embodiments of the present application provide a method and device for obtaining industry-related data, electronic equipment and storage medium, which can improve the accuracy and reliability of the evaluation of the industrial structure within a region.
[0004] To achieve the above-mentioned purpose, embodiments of the present application adopt the following technical solutions: In a first aspect, a method for obtaining industry-related data is provided. The method first obtains an industry chain graph corresponding to a target industry and a supply chain graph corresponding to a target region. The supply chain graph represents a business relationship network between business entities in a supply chain through a plurality of first nodes and a plurality of first directed edges. The first node is a node corresponding to a business entity in the supply chain, and the business entity can be an enterprise, a specific business in the enterprise, or a subsidiary of the enterprise, etc. The first directed edge connects two first nodes and is used to indicate that there is a business relationship between the business entities corresponding to the two connected nodes. The direction of the directed edge can represent the direction of the business flow. The industry chain graph includes a plurality of second nodes, and each second node represents a node corresponding to an industry link in the industry chain, such as raw material supply, production and manufacturing, circulation and sales, etc.
[0005] Then, a classification tree model is used to determine a third node matching the second node of the industry chain graph from a plurality of first nodes of the supply chain graph. Through the classification tree model technology, the business entity nodes in the supply chain can be mapped to the corresponding industry links of the industry chain, realizing the association of the supply chain graph and the industry chain graph, so as to screen out the third nodes related to the target industry. Then, according to the node information of the third node, the third node is quantitatively calculated to determine the node characteristic value of the third node. The node information of the third node is the portrait information of the corresponding business entity of the node in the supply chain graph, which can reflect the market position, operating condition, development potential and other characteristics of the business entity in multiple dimensions. By analyzing the node information to evaluate the node characteristic value of the node, the business entity with a key role in the target industry can be identified.
[0006] Then, the business transaction information of each group of business entities can be obtained, and the business association value of two business entities in each group of business entities can be calculated according to the business transaction information of each group of business entities, respectively. The two business entities in each group of business entities are the business entities corresponding to the two third nodes connected by the first directed edge in the supply chain graph. The business transaction information can reflect the closeness of cooperation between business entities, and the business association value can be calculated to quantitatively evaluate the business relationship between business entities in the supply chain.
[0007] Finally, the industry-related data of the target industry can be calculated by aggregating the node characteristic value of the third node and the business association value of the two business entities in each group of business entities. The industry-related data is used to indicate the industry operation state and industry development index of the target industry in the target region.
[0008] By using the scheme, through the joint analysis of the industry chain graph and the supply chain graph, the overall structural characteristics of the industry chain are considered, and the actual association relationship between business entities in the supply chain is also considered. The classification tree model is used to realize the accurate matching of the industry links in the industry chain and the business entities in the supply chain, ensuring the accuracy of the evaluation object. The node characteristic value of each business entity node matched to the industry link and the business association value between the related business entity nodes are calculated, considering not only the node characteristic value of the individual business entity, but also the connection closeness and synergistic effect between the business entities in the supply chain network. The industry-related data calculated by the node characteristic value and the business association value can comprehensively and objectively reflect the industry operation state and industry development index of the target industry in the target region, improving the accuracy and reliability of the regional industry evaluation.
[0009] In a possible implementation manner of the first aspect, the third node matched with the second node of the industry chain graph is determined from the plurality of first nodes of the supply chain graph by using the classification tree, including: obtaining identity information of a business entity corresponding to each of the plurality of first nodes of the supply chain graph; performing classification on the plurality of first nodes according to the identity information of the business entities corresponding to the plurality of first nodes by using the classification tree to obtain a classification result; the classification result is used to indicate an industry link to which the business entity corresponding to the first node belongs; and based on the classification result of the plurality of first nodes, the third node matched with the second node of the industry chain graph is selected from the plurality of first nodes; wherein the industry link to which the business entity corresponding to the third node belongs matches the industry link corresponding to at least one second node in the industry chain graph.
[0010] Through this implementation manner, the identity information of the business entity is used as the classification basis, and the classification tree can accurately classify the business entity into the corresponding industry link according to the industry attribute, business scope, product type and other characteristics of the business entity. The automatic classification method based on the identity information avoids the subjectivity and inefficiency of manual judgment, improves the accuracy and processing efficiency of node matching. The classification result clearly indicates the positioning of each business entity in the industry chain, providing a reliable foundation for subsequent accurate screening. Through the matching mechanism of the industry link, it is ensured that the selected third node indeed belongs to the related link of the target industry, and the interference of irrelevant business entities is excluded, so that the evaluation focuses on the truly relevant part of the supply chain network.
[0011] In another possible implementation manner of the first aspect, the node information of the third node includes a quantitative analysis parameter, and the quantitative analysis parameter includes a plurality of market parameters of the business entity corresponding to the third node, and the plurality of market parameters include at least one of sales, transaction volume, market share, profit margin, market heat, ESG index and market share growth rate of the business entity corresponding to the third node. The node information of the third node is used to perform quantitative calculation on the third node to obtain a node characteristic value of the third node, including: calculating a weighted sum of the plurality of market parameters according to a preset weight combination to obtain the node characteristic value of the third node; wherein the preset weight combination includes a plurality of first weights, the plurality of first weights correspond to the plurality of market parameters one by one, and the sum of the plurality of first weights is equal to 1; and the plurality of market parameters are processed by standardization.
[0012] Through this implementation manner, a plurality of market parameters are used for quantitative analysis, which can objectively measure the market performance and market competitiveness of the business entity from different dimensions, forming a comprehensive evaluation system. Through standardization processing, the differences in the numerical ranges of different market parameters are eliminated, so that the market parameters are comparable, and then a weighted calculation method is used, so that the node evaluation is more objective and scientific, and the credibility of the evaluation result is improved.
[0013] In another possible implementation of the first aspect, the node information of the third node includes qualitative analysis parameters, which include entity labels of the business entity corresponding to the third node. Entity labels include one or more positive labels and one or more negative labels. Positive labels indicate that the business entity corresponding to the third node has high stability and positive development capability or potential, while negative labels indicate that the business entity corresponding to the third node has poor stability and no development potential. The step of quantifying the third node based on its node information to obtain its node characteristic value includes: calculating the node characteristic value of the third node based on the number of positive labels and the number of negative labels in the qualitative analysis parameters.
[0014] This approach, compared to quantitative analysis parameters, uses qualitative analysis parameters to capture characteristics of business entities that are difficult to measure directly with numerical values. Positive labels indicate a promising future for the business entity, while negative labels indicate significant operational risks. By statistically analyzing the number of positive and negative labels, the overall quality of the business entity can be directly reflected. More positive labels indicate a more pronounced advantage for the business entity, and its node characteristic value should be higher; more negative labels indicate more serious problems, and its node characteristic value should be correspondingly lower. This label-based qualitative assessment method integrates expert knowledge and industry experience into the assessment system, making the determined node characteristic values more accurate.
[0015] In another possible implementation of the first aspect, the node information of the third node includes two types of information: quantitative analysis parameters and qualitative analysis parameters. The quantitative analysis parameters include various market parameters of the business entity corresponding to the third node, including at least one of sales revenue, transaction volume, market share, profit margin, market popularity, ESG index (Environmental, Social, and Governance score), and market share growth rate. The qualitative analysis parameters include entity labels of the business entity corresponding to the third node, which include one or more positive labels and one or more negative labels. Positive labels indicate that the business entity corresponding to the third node has high stability and positive development capabilities or potential, while negative labels indicate that the business entity corresponding to the third node has poor stability and no development potential.
[0016] When determining the node characteristic value of the third node based on the node information of the third node, firstly, a weighted sum of various market parameters is calculated according to a preset weight combination to obtain the node characteristic value of the third node. Then, the node characteristic value of the third node is increased according to the adjustment coefficient corresponding to the number of positive labels in the qualitative analysis parameters, and the node characteristic value of the third node is decreased according to the adjustment coefficient corresponding to the number of negative labels in the qualitative analysis parameters, thereby obtaining the final node characteristic value.
[0017] This approach combines quantitative and qualitative analysis to construct a more comprehensive node characteristic value evaluation system. First, node characteristic values based on market parameters are obtained through weighted summation. Then, they are dynamically adjusted according to the number of positive and negative labels, incorporating expert knowledge and industry experience into the evaluation process. This ensures that the evaluation results are both data-driven and consistent with expert knowledge and industry experience. This comprehensive evaluation method, combining quantitative and qualitative analysis, makes the calculation of node characteristic values more thorough, further improving the accuracy and reliability of the final industry-related data.
[0018] In another possible implementation of the first aspect, after quantifying the third node based on its node information to obtain its node feature value, the method further includes: obtaining the node information of the second node matching the third node in the industry chain map; the node information of the second node is the information of the industry chain corresponding to the second node; determining the adjustment coefficient corresponding to the node information of the second node from a preset mapping table, as the adjustment coefficient corresponding to the third node; wherein the preset mapping table includes the identifiers of multiple industry chains and the adjustment coefficient corresponding to each industry chain; increasing the node feature value of the third node according to the adjustment coefficient corresponding to the first type of industry chain in the adjustment coefficient corresponding to the third node; wherein the multiple industry chains include the first type of industry chain and the second type of industry chain, the first type of industry chain being the industry chain in which the target region has an advantageous position in terms of economy and / or development, and the second type of industry chain being the industry chain in which the target region has a disadvantageous position in terms of economy and / or development; decreasing the node feature value of the third node according to the adjustment coefficient corresponding to the second type of industry chain in the adjustment coefficient corresponding to the third node.
[0019] Because different industrial chains typically exhibit varying development statuses within specific regions, business entities within advantageous industrial chains contribute more significantly to the region's industrial operation and development indicators, while those within disadvantaged industrial chains contribute relatively less. By establishing a pre-defined mapping table to correspondence between industrial chains and adjustment coefficients, the characteristic values of the third node corresponding to the industrial segment in the first type of industrial chain can be increased to highlight the role of business entities in advantageous industrial chains. Conversely, the characteristic values of the third node corresponding to the industrial segment in the second type of industrial chain can be decreased to avoid overestimating the contribution of business entities in disadvantaged industrial chains. This dynamic adjustment method based on industrial chain type makes the assessment of node characteristic values more aligned with the actual industrial structure of the target region, further improving the accuracy of industry-related data.
[0020] In another possible implementation of the first aspect, the business transaction information includes one or more of the following: enterprise cooperation information, transaction frequency information, capital flow information, and logistics information between the two business entities in each group of business entities. Each item of business transaction information corresponds to a second weight. The business association value between the two business entities in each group of business entities is calculated based on their respective business transaction information, including: calculating the weighted value of the business association between the two business entities in each group of business entities by weighting each item of business transaction information according to one or more items of business transaction information and the corresponding second weight.
[0021] This approach uses multi-dimensional business transaction information to measure the relationships between business entities. Different types of business transaction information depict the characteristics of business relationships from different perspectives. By comprehensively considering this information, a more accurate assessment of the degree of relationship can be obtained. Then, a weighted calculation method is used to integrate multiple business transaction information, which not only avoids the one-sidedness of single business transaction information, but also reflects the relative importance of each business transaction information through weight settings. The calculated business relationship value can truly reflect the connection strength between business entities in the supply chain network.
[0022] In another possible implementation of the first aspect, the industry-related data of the target industry is calculated by aggregating the node feature values of the third node and the business association values of the two business entities in each group of business entities. This includes: adding the node feature values of all third nodes and the business association values of the two business entities in each group of business entities to obtain the industry-related data of the target industry.
[0023] This approach, which combines node characteristic values and business-related values through summation, enables a quantitative assessment of the target industry as a whole, while ensuring the interpretability of the assessment results.
[0024] Secondly, a device for acquiring industry-related data is provided, the device comprising: The graph acquisition module is used to acquire the industrial chain graph corresponding to the target industry and the supply chain graph corresponding to the target region. The supply chain graph includes multiple first nodes and multiple first directed edges. The first nodes are the nodes corresponding to the business entities in the supply chain. The first directed edges connect two first nodes and are used to indicate the business relationship between the business entities corresponding to the two first nodes. The business entity is at least one of the following: an enterprise, a business in an enterprise, or a subsidiary of an enterprise. The industrial chain graph includes multiple second nodes, which are the nodes corresponding to each industrial link in the industrial chain. The node matching module is used to determine the third node that matches the second node of the industrial chain map from multiple first nodes of the supply chain map using a classification tree model; The node calculation module is used to determine the node feature value of the third node based on the node information of the third node; the node information of the third node is the profile information of the business entity corresponding to the third node in the supply chain graph. The association calculation module is used to obtain the business transaction information of each group of business entities in multiple groups of business entities, and calculate the business association value of two business entities in each group of business entities based on the business transaction information of each group of business entities; wherein, the two business entities in each group of business entities are the business entities corresponding to two third nodes connected by the first directed edge in the supply chain graph. The industry assessment module is used to calculate industry-related data of the target industry based on the node characteristic value of the third node and the business association value of the two business entities in each group of business entities; among them, the industry-related data is used to indicate the industry operation status and industry development indicators of the target industry in the target region.
[0025] Thirdly, an electronic device is provided, the method comprising: a memory and at least one processor. The memory is communicatively connected to the processor. The memory is used to store computer program code, the computer program code including computer instructions. When the processor executes the computer instructions, it causes the electronic device to perform the method of the first aspect and any possible implementation thereof.
[0026] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions. When executed by a processor, these computer instructions are used to implement the method described in the first aspect and any possible implementation thereof.
[0027] Fifthly, embodiments of this application provide a computer program product that, when run on a computer or executed by a computer's processor, implements the method described in the first aspect and any possible design thereof. The computer may be the electronic device described in the second aspect and any possible implementation thereof.
[0028] It is understood that the beneficial effects achieved by the apparatus for acquiring industry-related data as described in the second aspect above, the electronic device as described in the third aspect, the computer-readable storage medium as described in the fourth aspect, and the computer program product as described in the fifth aspect can be referred to as the beneficial effects in the first aspect and any possible implementation thereof, and will not be repeated here. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating a method for acquiring industry-related data provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a node matching process provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a method for determining node feature values provided in an embodiment of this application; Figure 4 This is a flowchart illustrating another method for obtaining industry-related data provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an apparatus for acquiring industry-related data provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0030] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0031] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0032] The collection, storage, use, processing, transmission, provision, and disclosure of node information and other related information in the technical solutions provided in this application comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0033] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, they do not mean that the applicant has used or necessarily used the solution.
[0034] In related technologies, there are methods based on network analysis that construct industry networks to assess inter-industry connections and thus evaluate industry structure. However, these methods often neglect data on individual firms within the industry, leading to generally low accuracy and reliability of the assessment results. For example, if a key firm in the industry is in poor financial condition, analyzing only industry data while ignoring that firm's data fails to assess the impact on the industry's future development should that firm face funding shortfalls.
[0035] Based on this, embodiments of this application provide a method for obtaining industry-related data, which can improve the accuracy and reliability of the rationality assessment of the industrial structure within a region by analyzing the degree of consistency between the industrial chain map and the supply chain map.
[0036] The method for acquiring industry-related data provided in this application can be applied to electronic devices with data processing capabilities, such as servers. Alternatively, the electronic device may include a personal computer (PC), tablet computer, laptop computer, portable computer (such as a mobile phone), wearable electronic device (such as a smartwatch), augmented reality (AR) / virtual reality (VR) device, in-vehicle computer, etc. The following embodiments do not impose special limitations on the specific form of the electronic device. The execution subject of the method for acquiring industry-related data provided in this application can be the aforementioned electronic device, or it can be a device for determining and acquiring industry-related data. This device for acquiring industry-related data can be integrated into the electronic device or the processor of the electronic device. For ease of explanation, the subsequent content of this application embodiment uses a server as the execution subject to illustrate the risk response process.
[0037] The server first acquires the industry chain map corresponding to the target industry and the supply chain map corresponding to the target region. Using a classification tree model, it categorizes and matches business entity nodes in the supply chain map to identify those that match industry segment nodes in the industry chain map. Then, based on the node information of the matched business entity nodes, the server determines the node characteristic value of each matched business entity node within the industry structure. Furthermore, the server acquires business transaction information between related business entities that are connected in the supply chain map and calculates business correlation values between these entities. Finally, based on the node characteristic values of all matched business entity nodes and the business correlation values between all related business entities, the server comprehensively calculates industry-related data for the target industry, indicating the industry's operational status and industry development indicators within the target region.
[0038] In this way, through joint quantitative assessment of the macro-industrial chain and the micro-supply chain, not only is the existence of business entities matching the industrial links in the region considered, but also the node characteristic values of the business entities and the business relationship values between the various business entities are further quantified. This enables a more accurate and comprehensive assessment of the actual role and development health of the target industry in the target region, and improves the accuracy and reliability of the industrial structure assessment in the region.
[0039] Please refer to Figure 1 ,Figure 1 This is a flowchart illustrating a method for acquiring industry-related data according to an embodiment of this application. The method includes the following steps.
[0040] S101, the server obtains the industrial chain map corresponding to the target industry and the supply chain map corresponding to the target region.
[0041] In this embodiment of the application, the industrial chain refers to the collection of all industrial links involved in the entire process of an industry, from upstream raw materials and R&D design, to midstream processing and manufacturing, and then to downstream distribution, sales, and after-sales service. An industrial chain map is a visual representation of the industrial chain, including multiple second nodes, each corresponding to an industrial link in the industrial chain.
[0042] It should be noted that the industrial chain segments can include various aspects such as R&D and design, raw material supply, production and manufacturing, logistics and transportation, marketing and sales, technical services, and after-sales maintenance. For example, in the new energy vehicle industry chain, the industrial chain segments can include: upstream lithium mining, battery material R&D, and battery material processing; midstream cell production, battery assembly, vehicle design, and vehicle manufacturing; and downstream vehicle sales, charging infrastructure construction, after-sales service, and battery recycling. Each segment corresponds to a second node in the industry chain diagram.
[0043] A supply chain refers to a network structure formed by business relationships among multiple business entities. A supply chain graph is a visual representation of a supply chain, including multiple first nodes and multiple first directed edges. A first node is a node corresponding to a business entity in the supply chain. A business entity can be a company, a business function within a company, or a subsidiary of a company. A first directed edge connects two first nodes and indicates the business relationship between the corresponding business entities. For example, a first directed edge can indicate a business relationship where one business entity provides raw materials, components, technical services, or sales channels to another business entity.
[0044] In some embodiments, the server can retrieve the industry chain map corresponding to the target industry from an industry database. The industry database can store the industry chain maps corresponding to various industries. The server can retrieve the industry chain map corresponding to the target industry from the industry database based on the identification information of the target industry input by the user.
[0045] In some embodiments, the server can obtain the supply chain graph corresponding to the target region from the enterprise database and the transaction database. The enterprise database can store basic information of each enterprise in the target region, including enterprise name, enterprise type, business scope, etc. The transaction database can store transaction records between enterprises, including the parties involved, transaction content, transaction amount, etc. The server can create the first node based on the enterprise information in the enterprise database, and create the first directed edge between the first nodes corresponding to the relevant enterprises based on the transaction records in the transaction database, thereby constructing the supply chain graph corresponding to the target region.
[0046] S102, the server uses a classification tree to determine the third node that matches the second node of the industrial chain map from multiple first nodes of the supply chain map.
[0047] In this embodiment, the server needs to match the business entity nodes in the supply chain graph with the industry link nodes in the industry chain graph to determine the third node that matches the second node in the industry chain graph. This matching process is implemented using a classification tree model.
[0048] A classification tree is a machine learning model that can classify data based on input features. In this embodiment, the server can use a classification tree model to classify the first node into different industry segments in the industry chain based on the identity information of the business entity corresponding to the first node, thereby achieving matching between the first node and the second node.
[0049] In some embodiments, please refer to Figure 2 , Figure 2 This is a flowchart illustrating a node matching process provided in an embodiment of this application. Step S102 may include the following sub-steps.
[0050] S1021, For each of the multiple first nodes in the supply chain graph, the server obtains the identity information of the business entity corresponding to that first node.
[0051] In this embodiment, the identity information of a business entity refers to information that indicates the industry, business type, and scope of operations of the business entity. For example, identity information may include the industry classification of the business entity, its business scope, entity category, geographical location, and stage in the industry chain. As a specific example, for a battery manufacturing company, its identity information may include: industry classification as battery manufacturing, business scope including lithium battery production and sales, category as a manufacturing company, geographical location in an industrial park, and stage in the midstream manufacturing segment of the industry chain.
[0052] In some embodiments, the server can obtain the identity information of a business entity from an enterprise database. The enterprise database stores the registration and operational information of each business entity, and the server can directly obtain this registration and operational information from the enterprise database and use it as identity information.
[0053] In other embodiments, the server can also extract identity information from publicly available data of the business entity using natural language processing technology. For example, the server can extract information such as the business entity's industry classification and business scope from the business entity's official website, annual reports, news reports, etc., as identity information.
[0054] S1022, the server uses a classification tree to classify multiple first nodes according to the identity information of the business entities corresponding to the multiple first nodes, and obtains the classification results.
[0055] In this embodiment, the classification result is used to indicate the industry segment to which the business entity corresponding to the first node belongs. The server uses identity information as input features and classifies the first node using a classification tree model. The classification result output by the classification tree model indicates the industry segment to which the business entity corresponding to the first node belongs.
[0056] In some embodiments, the server may employ a gradient boosting-based classification tree model, such as the LightGBM model. LightGBM is a gradient boosting decision tree model in which the server inputs identity information of a business entity, such as its industry classification, business scope, category, geographical location, and industry chain stage, into the LightGBM model. The LightGBM model can then output the industry segment to which the business entity belongs.
[0057] When training a classification tree model, the server can use labeled training data. This training data includes the identity information of multiple business entities and the industry segment labels to which those entities belong. The server uses this training data to train the classification tree model, which then learns the correspondence between different identity information and different industry segments.
[0058] In some embodiments, for enterprises operating multiple businesses, the server can divide the enterprise into multiple business nodes, with each business node corresponding to one of the enterprise's businesses. The server can then allocate the enterprise's total output value proportionally across these business entity nodes based on the proportion of output value from each business. This ensures comprehensiveness and accuracy in matching, avoiding ambiguity in classification caused by the enterprise operating multiple businesses.
[0059] For example, for a home appliance company that manufactures both air conditioners and refrigerators, the server can split the company into two business nodes: one corresponding to the air conditioner production business and the other to the refrigerator production business. If the company's air conditioner business accounts for 60% of its output value and the refrigerator business accounts for 40%, then the server can allocate 60% of the company's total output value to the air conditioner business node and 40% to the refrigerator business node.
[0060] In other embodiments, for group enterprises, the server can identify each subsidiary of the group enterprise within a target region and treat each subsidiary as an independent business entity node. This avoids overemphasizing the role of the group enterprise within the region while also reflecting the collaborative relationships between the various subsidiaries of the group enterprise.
[0061] For example, for a group with multiple subsidiaries within a target region, the server can identify the group's subsidiaries such as air conditioner manufacturing, refrigerator manufacturing, smart home, and solar energy equipment, and treat each subsidiary as an independent first node. Specifically, the air conditioner and refrigerator manufacturing subsidiaries belong to the home appliance manufacturing segment, the smart home subsidiary to the smart device R&D and sales segment, and the solar energy equipment subsidiary to the new energy equipment manufacturing segment. The server can obtain the identity information of each subsidiary separately and classify each subsidiary node, ensuring that each subsidiary belongs to a different industry segment.
[0062] S1023, the server, based on the classification results of multiple first nodes, selects the third node that matches the second node of the industry chain map from the multiple first nodes.
[0063] In this embodiment, the third node refers to the first node whose industry segment to which the corresponding business entity belongs matches the industry segment corresponding to at least one second node in the industry chain diagram. That is, the business entity corresponding to the third node is a business entity participating in the target industry.
[0064] In some embodiments, the server can compare the classification results with the second nodes in the industry chain map. For each first node, if the industry segment indicated by the classification result of the first node is the same as the industry segment corresponding to a certain second node in the industry chain map, then the server can determine that first node as a third node. For example, if there is a second node in the industry chain map corresponding to the battery cell production segment, and the classification result of a certain first node indicates that the corresponding business entity belongs to the battery cell production segment, then the server can determine that first node as a third node.
[0065] S103, the server performs quantization calculations on the third node based on the node information of the third node to obtain the node feature value of the third node.
[0066] In this embodiment, the node information of the third node is the profile information of the corresponding business entity in the supply chain graph. The profile information is a comprehensive description of various characteristics of the business entity, which may include information such as the business entity's market performance, financial status, development potential, and industry position.
[0067] The node characteristic value of the third node is a numerical value used to quantify the importance of the business entity corresponding to that third node in the industry structure. The higher the node characteristic value, the greater the contribution of the business entity to industry development, and the higher its weight should be in the industry structure assessment.
[0068] In some embodiments, the node information of the third node may include quantitative analysis parameters. Quantitative analysis parameters refer to parameters that can be quantified numerically, including various market parameters of the business entity corresponding to the third node. Market parameters may include, but are not limited to: sales revenue, transaction volume, market share, profit margin, market popularity, ESG (environmental, social, and governance) score, and market share growth rate.
[0069] Sales revenue refers to the total sales revenue of a business entity within a certain period; transaction volume refers to the number of transactions completed by a business entity within a certain period; market share refers to the proportion of a business entity's sales revenue to the total sales revenue of the entire market; profit margin refers to the proportion of a business entity's net profit to its total sales revenue; market popularity refers to the number of searches for a business entity on various platforms; ESG index refers to the comprehensive score of a business entity in terms of environment, society, and governance; and market share growth rate refers to the change in a business entity's current market share compared to its previous market share.
[0070] In one specific embodiment, please refer to Figure 3 , Figure 3 This is a flowchart illustrating a method for determining node feature values provided in an embodiment of this application. When the node information of the third node includes quantitative analysis parameters, step S103 may include the following sub-steps.
[0071] S1031, the server standardizes various market parameters in the quantitative analysis parameters.
[0072] In this embodiment of the application, since different market parameters have different numerical ranges, the server can standardize multiple market parameters in the quantitative analysis parameters so that the values of various market parameters are within the same range, which facilitates subsequent weighted calculations.
[0073] In some embodiments, the server may employ a max-min normalization method. For each market parameter, the server calculates the maximum and minimum values of that market parameter across all third-party nodes, and then normalizes the parameter value for each third-party node according to the following formula: Normalized value = (Original value - Minimum value) / (Maximum value - Minimum value). After normalization, the values of all market parameters are between 0 and 1.
[0074] S1032, the server calculates the weighted sum of multiple market parameters according to the preset weight combination to obtain the node characteristic value of the third node.
[0075] In this embodiment, the preset weight combination includes multiple first weights, each corresponding to a different market parameter, and the sum of the first weights equals 1. The server multiplies the standardized value of each market parameter by its corresponding first weight, and then adds all the products together to obtain the node feature value of the third node. The preset weight combination can be pre-set based on expert experience.
[0076] Specifically, the server can calculate the node characteristic value of the third node according to the following formula: Node characteristic value = w1 × Standardized sales revenue + w2 × Standardized transaction volume + w3 × Standardized market share + w4 × Standardized profit margin + w5 × Standardized market popularity + w6 × Standardized ESG index + w7 × Standardized market share growth rate. Where w1, w2, w3, w4, w5, w6, and w7 are the first weights corresponding to sales revenue, transaction volume, market share, profit margin, market popularity, ESG index, and market share growth rate, respectively, and w1 + w2 + w3 + w4 + w5 + w6 + w7 = 1. For example, w1 = 0.3, w2 = 0.2, w3 = 0.15, w4 = 0.1, w5 = 0.1, w6 = 0.1, and w7 = 0.05.
[0077] In other embodiments, the node information of the third node may also include qualitative analysis parameters. Qualitative analysis parameters are non-numerical parameters described by labels or categories. For example, qualitative analysis parameters include the entity label of the business entity corresponding to the third node.
[0078] Entity labels include one or more positive labels and one or more negative labels. Positive labels indicate that the business entity corresponding to the third node is highly stable and has positive development capabilities or potential. Negative labels indicate that the business entity corresponding to the third node is unstable and lacks development potential.
[0079] In some embodiments, positive labels may include: high-tech enterprise, specialized and innovative enterprise, technology leader, internationalized, diversified, green enterprise, cash cow, century-old brand, etc. These positive labels indicate the business entity's advantages in technological innovation, market expansion, and sustainable development. Negative labels may include: sunset industry, labor-intensive, resource-dependent, monopolistic, etc. These negative labels indicate the risks the business entity may face, such as industry decline risk, cost fluctuation risk, resource depletion risk, insufficient innovation capabilities, etc.
[0080] In one specific embodiment, please refer to [link / reference]. Figure 3 If the node information of the third node includes qualitative analysis parameters, step S103 may include the following sub-steps.
[0081] S1033, the number of positive labels and the number of negative labels in the server's statistical qualitative analysis parameters.
[0082] In this embodiment, the server can extract all entity tags from the profile information of business entities and classify all entity tags into positive or negative tags according to the tag library. The tag library stores various tags and their corresponding categories, and the server can determine the category of each tag by querying the tag library.
[0083] S1034, the server calculates the node feature value of the third node based on the number of positive labels and the number of negative labels.
[0084] In some embodiments, when the number of positive tags is greater than the number of negative tags, the server calculates the node characteristic value of the third node according to the following formula: Node characteristic value = (Number of positive tags - Number of negative tags) / (Number of positive tags + Number of negative tags). When the number of negative tags is greater than the number of positive tags, the server calculates the node characteristic value of the third node according to the following formula: Node characteristic value = (Number of negative tags - Number of positive tags) / (Number of positive tags + Number of negative tags) × (-1).
[0085] Thus, the node feature values calculated by the server range from -1 to 1. The more positive labels a node has, the closer its feature value is to 1; the more negative labels a node has, the closer its feature value is to -1.
[0086] In other embodiments, the server can also comprehensively analyze quantitative and qualitative analysis parameters to calculate the node characteristic value of the third node. The server can first calculate two scores based on the quantitative and qualitative analysis parameters respectively, and then take a weighted average of the two scores to obtain the final node characteristic value.
[0087] Alternatively, in some embodiments, when the node information of the third node includes both quantitative and qualitative analysis parameters, the server can comprehensively analyze the quantitative and qualitative analysis parameters to determine the node characteristic value of the third node. Specifically, the server calculates a weighted sum of multiple market parameters according to a preset weight combination to obtain the node characteristic value of the third node; this calculation process is the same as in the embodiments described above. Then, the server can obtain the number of positive labels and the number of negative labels in the qualitative analysis parameters.
[0088] For positive labels, the server can increase the node feature value of the third node according to the adjustment coefficient corresponding to the number of positive labels. In this embodiment, the server can pre-establish a mapping relationship between the number of positive labels and the adjustment coefficient. The more positive labels there are, the larger the corresponding adjustment coefficient. For example, the server can multiply the number of positive labels by a preset positive adjustment step size to determine the adjustment coefficient, where the positive adjustment step size is a preset parameter. Then, the server multiplies the quantitative node feature value by the positive adjustment coefficient to obtain the node feature value adjusted by the positive labels.
[0089] For negative labels, the server can adjust the node feature value of the third node according to the adjustment coefficient corresponding to the number of negative labels. In this embodiment, the server can also pre-establish a mapping relationship between the number of negative labels and the adjustment coefficient. The more negative labels there are, the smaller the corresponding adjustment coefficient. For example, the server can multiply the number of negative labels by a preset negative adjustment step size to determine the adjustment coefficient, where the negative adjustment step size is also a preset parameter. Then, the server multiplies the quantitative node feature value by the negative adjustment coefficient to obtain the node feature value adjusted by the negative labels.
[0090] In this way, the server can comprehensively consider both quantitative and qualitative analysis parameters to obtain more accurate third-node feature values.
[0091] S104, the server obtains the business transaction information of each of the multiple business entities.
[0092] In this embodiment, the two business entities in each group are the business entities corresponding to two third nodes connected by a first directed edge in the supply chain graph. In other words, in this embodiment, the server only needs to obtain the business transaction information between pairs of business entities that have a business relationship in the supply chain graph and are both identified as third nodes.
[0093] Business transaction information refers to the business interactions between two business entities, which may include one or more of the following: business cooperation information, transaction frequency information, capital flow information, and logistics information.
[0094] Specifically, enterprise cooperation information refers to information such as whether there are collaborative projects between two business entities and the degree of closeness of their cooperation. For example, the number of R&D projects jointly participated in, the number of patents jointly applied for, and shared technology platforms. Transaction frequency information refers to information such as the number of transactions and contracts between two business entities. For example, the number of purchase orders and service contracts between two business entities within a year. Fund flow information refers to information such as the amount of funds transferred and investment between two business entities. For example, the total amount of payments made by one business entity to another, and the amount of investment made. Logistics information refers to information such as transportation time, transportation costs, and logistics service quality between two business entities. For example, the average time it takes to transport goods from one business entity to another, and the proportion of transportation costs to the transaction amount.
[0095] In one example, the server can retrieve transaction frequency and fund flow information from a transaction database. The transaction database stores historical transaction records between partner companies, including transaction time, amount, and details. The server can count the number of transactions between two business entities within a certain period, as transaction frequency information; and it can sum the transaction amounts between the two business entities, as fund flow information.
[0096] In one example, the server can retrieve logistics information from a logistics database. This database stores logistics records between businesses, including shipping time, delivery time, transportation distance, and transportation costs. The server can then calculate the average transportation time and average transportation cost between the two business entities as logistics information.
[0097] In one example, the server can retrieve enterprise collaboration information from a collaborative project database. This database stores records of collaborative projects between enterprises, including project name, project type, participating enterprises, and project amount. The server can then count the number of projects jointly participated in by two business entities, which can be used as enterprise collaboration information.
[0098] In other embodiments, when certain business transaction information is difficult to obtain directly, the server can use the industry average level as the value for that business transaction information. The server can query the industry database for the average business transaction parameters of the industry to which the two business entities belong, and use this as the business transaction information between the two entities. For example, for two business entities belonging to the electronic components industry, if logistics cost data between the two business entities cannot be obtained, the server can use the average logistics cost ratio of the electronic components industry as an estimate.
[0099] S105, the server calculates the business association value between two business entities in each group of business entities based on the business transaction information of each group of business entities.
[0100] In this embodiment, the business association value is a numerical value used to quantify the degree of business connection between two business entities. The higher the business association value, the more frequent and closer the business interactions between the two business entities, and the stronger the synergistic effect on the industry chain.
[0101] In some embodiments, each piece of business transaction information corresponds to a second weight. The server calculates the business association value between two business entities in each group of business entities by weighting the data based on one or more pieces of business transaction information for each group of business entities and the second weight corresponding to each piece of business transaction information.
[0102] Specifically, the server can calculate the business association value using the following formula: Business association value = ×Enterprise Cooperation Information+ ×Transaction Frequency Information+ ×Capital Flow Information+ × Logistics transaction information. Among them, arrive These are the second weights corresponding to enterprise cooperation information, transaction frequency information, capital flow information, and logistics information, respectively. The second weight can be determined based on expert experience.
[0103] In other embodiments, to make the calculation results more clearly reflect the differences, the server can employ an exponential function to enhance the influence of important features and obtain business-related values. The exponential function amplifies more important features, making their impact more significant. Specifically, the server can calculate the business-related values according to the following formula: ; in, It is the degree of association between node i and node j. The value represents the enterprise cooperation information between node i and node j. It is the value of the transaction frequency information. It is the value of information on fund flows. This is the value for logistics information.
[0104] In some other embodiments, to ensure relatively stable calculation results, the server can use the Sigmoid function to normalize the business association values, ensuring that the normalized values remain within a certain range. Specifically, the server can calculate the business association values according to the following formula: ; After processing with the Sigmoid function, the normalized business correlation values are between 0 and 1, which facilitates subsequent calculations of industry-related data.
[0105] S106, the server performs quantitative calculations based on the node characteristic value of the third node and the business association value of the two business entities in each group of business entities to obtain the industry-related data of the target industry.
[0106] In this embodiment, industry-related data is used to indicate the operational status and industry development indicators of the target industry in the target region. The industry-related data comprehensively considers the node characteristic values of each business entity in the industrial chain, as well as the business relationship values between business entities, thus comprehensively reflecting the development status of the target industry in the target region.
[0107] In some embodiments, the server adds the node characteristic values of all third nodes to the business association values of two business entities in each group of business entities to obtain the industry-related data of the target industry. This calculation method considers both the importance of each business entity within the industry chain and the collaborative relationships between them, thus more accurately reflecting the overall role of the target industry in the target region.
[0108] For example, if there are 10 business entities identified as third nodes in the target area, and there are 15 first directed edges between these 10 business entities, that is, there are 15 business-related values, then the server adds the node feature values of these 10 third nodes together, and adds the sum of the 15 business-related values to obtain the industry-related data of the target industry.
[0109] In some embodiments, the server can also perform standardization processing on industry-related data. For example, the server can use a min-max standardization method to standardize industry-related data, as described in the embodiments above.
[0110] In some embodiments, the server can also evaluate the industrial structure of a target region based on industry-related data. For example, the server can set a threshold for industry-related data. When the industry-related data is higher than the threshold, it is determined that the target industry is developing well in the target region; when the industry-related data is lower than the threshold, it is determined that the target industry is developing poorly in the target region and measures need to be taken to improve it.
[0111] In some embodiments, the server generates a visual display interface for industry-related data. The server can graphically display industry-related data and related information, allowing users to intuitively understand the development status of the target industry in the target region.
[0112] In some embodiments, to further improve the accuracy of the assessment, the server may also employ a multi-dimensional assessment method. In addition to the matching degree assessment based on the industry chain map and supply chain map, the server may also combine information from other dimensions, such as technological innovation capabilities, environmental impact, and social benefits, to conduct a comprehensive assessment of the target industry, and determine the final assessment result based on the assessment results of other dimensions and the industry-related data obtained in the above embodiments.
[0113] In some embodiments, after S103, the server may further adjust the node's characteristic value; please refer to [reference needed]. Figure 4 , Figure 4 This is a flowchart illustrating another method for obtaining industry-related data provided in this application embodiment. The content of S101-S106 can be referred to in the above embodiment; S201-S204 are described below.
[0114] S201, the server obtains the node information of the second node that matches the third node in the industry chain map.
[0115] In this embodiment, the node information of the second node is the information of the industrial chain corresponding to the second node. Here, the industrial chain corresponding to the second node refers to the industrial chain to which the industrial segment corresponding to the second node belongs.
[0116] S202, the server determines the adjustment coefficient corresponding to the node information of the second node from the preset mapping table, and uses it as the adjustment coefficient corresponding to the third node.
[0117] In this embodiment, the preset mapping table includes identifiers for multiple industry chains and adjustment coefficients for each industry chain. The adjustment coefficients are used to adjust the node characteristic values of the third node based on the development status of the industry chain in the target area.
[0118] In some embodiments, multiple industrial chains can be categorized into a first type and a second type. The first type of industrial chain refers to those in the target region that hold an advantageous position in terms of economy and / or development, such as core industrial chains occupying a central position in the regional economy, high value-added industrial chains with advantages in contributing to regional economic growth, technological innovation, and employment, and strategic emerging industrial chains capable of achieving industrial upgrading and innovation. The second type of industrial chain refers to those in the target region that hold a disadvantageous position in terms of economy and / or development, such as traditional industrial chains facing decline, low value-added industrial chains with low contribution to the regional economy and limited development potential, environmentally sensitive industrial chains that negatively impact the environment, and industrial chains that the regional economy is overly reliant on.
[0119] In the mapping table, the adjustment coefficient for the first type of industry chain is greater than 1, which is used to increase the node characteristic value of the third node belonging to this industry chain. The adjustment coefficient for the second type of industry chain is less than 1, which is used to decrease the node characteristic value of the third node belonging to this industry chain.
[0120] For example, in one specific embodiment, the mapping table may include the following: For the machinery and equipment industry chain, the target region has advantages in technological innovation, large market demand, and obvious intelligent manufacturing trend. The adjustment coefficient corresponding to the machinery and equipment industry chain is 1.115, which means that the node characteristic value of the third node of the industrial link that is matched to the machinery and equipment industry chain is increased by 11.5%.
[0121] For the electronic component industry chain, in the target region where technological progress is rapid and market demand is large, the adjustment coefficient corresponding to the electronic component industry chain is 1.165, which means that the node characteristic value of the third node in the industrial link that matches the electronic component industry chain is increased by 16.5%.
[0122] For the pharmaceutical industry chain, given the high R&D investment and stable market demand in the target region, the adjustment coefficient for the pharmaceutical industry chain is 1.165, which means that the node characteristic value of the third node in the industrial link of the pharmaceutical industry chain is increased by 16.5%.
[0123] For the information technology service industry chain, given the accelerated digital transformation and widespread application of cloud computing and big data in the target region, the adjustment coefficient for the information technology service industry chain is 1.165, which means that the node characteristic value of the third node in the industry chain that is matched with the information technology service industry chain is increased by 16.5%.
[0124] For the healthcare service industry chain, with the increase in health awareness, demand, and technology application in the target area, the adjustment coefficient for the healthcare service industry chain is 1.165, which means that the node characteristic value of the third node in the industry chain that is matched with the healthcare service industry chain is increased by 16.5%.
[0125] For the petrochemical industry chain, facing significant environmental policy pressure and poor market prospects in the target region, the adjustment coefficient for the petrochemical industry chain is 0.865, which means that the node characteristic value of the third node in the industrial segment that is matched to the petrochemical industry chain is reduced by 13.5%.
[0126] For the chemical industry chain, given the significant environmental policy pressure and clear demand for transformation and upgrading in the target region, the adjustment coefficient for the chemical industry chain is 0.815, which means that the node characteristic value of the third node in the industrial link of the chemical industry chain will be reduced by 18.5%.
[0127] For the textile and apparel industry chain, facing severe homogenization and fierce price competition in the target region, the adjustment coefficient corresponding to the textile and apparel industry chain is 0.765, which means that the node characteristic value of the third node of the industrial link that is matched to the textile and apparel industry chain is reduced by 23.5%.
[0128] For the agricultural, forestry, animal husbandry and fishery industry chain, where the level of modernization in the target area has improved but the overall added value is relatively low, the adjustment coefficient corresponding to the agricultural, forestry, animal husbandry and fishery industry chain is 0.965. This means that the node characteristic value of the third node in the industrial chain that is matched to the agricultural, forestry, animal husbandry and fishery industry chain is reduced by 3.5%. The server looks up the corresponding adjustment coefficient from the mapping table based on the industry chain identifier corresponding to the second node, and uses it as the adjustment coefficient corresponding to the third node.
[0129] S203, the server increases the node characteristic value of the third node according to the adjustment coefficient of the first type of industrial chain in the adjustment coefficient corresponding to the third node.
[0130] In this embodiment of the application, if the third node belongs to the first type of industry chain, the server multiplies the node feature value of the third node by the corresponding adjustment coefficient. Since the adjustment coefficient is greater than 1, the node feature value will be increased.
[0131] S204, the server lowers the node characteristic value of the third node according to the adjustment coefficient of the second type of industrial chain in the adjustment coefficient corresponding to the third node.
[0132] In this embodiment of the application, if the third node belongs to the second type of industry chain, the server multiplies the node feature value of the third node by the corresponding adjustment coefficient. Since the adjustment coefficient is less than 1, the node feature value will be lowered.
[0133] The above adjustments can make the node characteristic values of the third node more consistent with the actual development of the industrial chain in the target area, thereby improving the accuracy of industry-related data.
[0134] In another specific embodiment, the server can further acquire environmental impact information of the business entity corresponding to the third node, including energy consumption, pollutant emissions, and resource utilization efficiency. The server can calculate an environmental impact score based on the environmental impact information and use the environmental impact score as an adjustment factor for the node characteristic value of the third node. For environmentally friendly business entities, the server can further increase the node characteristic value of the corresponding third node; for business entities with significant environmental impact, the server can decrease the node characteristic value of the corresponding third node.
[0135] In another specific embodiment, the server can further obtain social contribution information of the business entity corresponding to the third node, including the number of employees, tax contributions, and participation in public welfare activities. The server can calculate a social contribution score based on the social contribution information and use the social contribution score as an adjustment factor for the node characteristic value of the third node. For business entities with high social contributions, the server can further increase their node characteristic value.
[0136] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of a device for acquiring industry-related data provided in an embodiment of this application. Figure 5 As shown, the device for acquiring industry-related data includes: a map acquisition module 501, a node matching module 502, a node calculation module 503, a correlation calculation module 504, and an industry assessment module 505.
[0137] The graph acquisition module 501 is used to acquire the industrial chain graph corresponding to the target industry and the supply chain graph corresponding to the target region. The supply chain graph includes multiple first nodes and multiple first directed edges. The first nodes are the nodes corresponding to the business entities in the supply chain. The first directed edges connect two first nodes and are used to indicate the business relationship between the business entities corresponding to the two first nodes. The business entity is at least one of the following: an enterprise, a business in an enterprise, or a subsidiary of an enterprise. The industrial chain graph includes multiple second nodes, which are the nodes corresponding to each industrial link in the industrial chain. The node matching module 502 is used to determine the third node that matches the second node of the industrial chain map from multiple first nodes of the supply chain map using a classification tree model; The node calculation module 503 is used to perform quantitative calculations on the third node based on the node information of the third node to obtain the node feature value of the third node; the node information of the third node is the profile information of the business entity corresponding to the third node in the supply chain graph. The association calculation module 504 is used to obtain the business transaction information of each group of business entities in multiple groups of business entities, and calculate the business association value of two business entities in each group of business entities based on the business transaction information of each group of business entities; wherein, the two business entities in each group of business entities are the business entities corresponding to two third nodes connected by the first directed edge in the supply chain graph. The industry assessment module 505 is used to perform aggregate calculations based on the node characteristic values of the third node and the business association values of two business entities in each group of business entities to obtain industry-related data of the target industry; among which, the industry-related data is used to indicate the industry operation status and industry development indicators of the target industry in the target region.
[0138] In other embodiments, the node matching module 502 is further configured to: obtain the identity information of the business entity corresponding to each of the multiple first nodes in the supply chain graph; classify the multiple first nodes according to the identity information of the business entities corresponding to the multiple first nodes using a classification tree model to obtain classification results; use the classification results to indicate the industry segment to which the business entity corresponding to the first node belongs; and based on the classification results of the multiple first nodes, select a third node that matches a second node in the supply chain graph from the multiple first nodes; wherein the industry segment to which the business entity corresponding to the third node belongs matches the industry segment corresponding to at least one second node in the supply chain graph.
[0139] In other embodiments, the node information of the third node includes quantitative analysis parameters, which include multiple market parameters of the business entity corresponding to the third node. These multiple market parameters include at least one of the following: sales revenue, transaction volume, market share, profit margin, market popularity, EGS score, and market share growth rate of the business entity corresponding to the third node. The node calculation module 503 is further configured to determine the node characteristic value of the third node based on its node information, including: calculating a weighted sum of multiple market parameters according to a preset weight combination to obtain the node characteristic value of the third node; wherein the preset weight combination includes multiple first weights, each corresponding one-to-one with a multiple market parameter, and the sum of the multiple first weights equals 1; the multiple market parameters are standardized.
[0140] In other embodiments, the node information of the third node includes qualitative analysis parameters, which include entity tags of the business entity corresponding to the third node. Entity tags include one or more positive tags and one or more negative tags. Positive tags indicate that the business entity corresponding to the third node has high stability and positive development capability or potential, while negative tags indicate that the business entity corresponding to the third node has poor stability and no development potential. The node calculation module 503 is further configured to calculate the node characteristic value of the third node based on the number of positive tags and the number of negative tags in the qualitative analysis parameters.
[0141] In other embodiments, the apparatus for acquiring industry-related data may further include a data adjustment module for acquiring node information of a second node matching a third node in an industry chain map; the node information of the second node is the information of the industry chain corresponding to the second node; determining the adjustment coefficient corresponding to the node information of the second node from a preset mapping table as the adjustment coefficient corresponding to the third node; wherein the preset mapping table includes the identifiers of multiple industry chains and the adjustment coefficient corresponding to each industry chain; increasing the node characteristic value of the third node according to the adjustment coefficient corresponding to the first type of industry chain in the adjustment coefficient corresponding to the third node; wherein the multiple industry chains include the first type of industry chain and the second type of industry chain, the first type of industry chain being the industry chain in which the target region has an advantageous position in terms of economy and / or development, and the second type of industry chain being the industry chain in which the target region has a disadvantageous position in terms of economy and / or development; decreasing the node characteristic value of the third node according to the adjustment coefficient corresponding to the second type of industry chain in the adjustment coefficient corresponding to the third node.
[0142] In other embodiments, the business transaction information includes one or more of the following: enterprise cooperation information, transaction frequency information, fund flow information, and logistics information between two business entities in each group of business entities, with each item of business transaction information corresponding to a second weight. The aforementioned association calculation module 504 is further configured to calculate the business association value between two business entities in each group of business entities by weighting the calculation based on one or more items of business transaction information for each group of business entities and the second weight corresponding to each item of business transaction information.
[0143] In other embodiments, the aforementioned industry assessment module is also used to add the node feature values of all third nodes and the business association values of two business entities in each group of business entities to obtain industry-related data of the target industry.
[0144] The apparatus for acquiring industry-related data provided in this application embodiment can execute the method shown in the above method embodiment. Its implementation principle and beneficial effects can be referred to the relevant description in the method embodiment, and will not be repeated here.
[0145] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 6 As shown, the electronic device includes: a memory 601, a transceiver 602, and at least one processor 603.
[0146] Transceiver 602 is used to interact with other devices to send and receive data.
[0147] The memory 601 stores computer program code, which includes computer instructions. These computer instructions run in the described electronic device to implement the method shown in the above-described method embodiments. For example, the memory may include high-speed random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, portable hard drive, read-only memory, disk, or optical disc, etc.
[0148] Processor 603 can be a general-purpose processor, including a Central Processing Unit (CPU), a network processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Processor 603 can also be other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor.
[0149] The memory 601, transceiver 602, and processor 603 are communicatively connected. For example, the memory 601 and transceiver 602 can be connected to the processor 603 via a system bus and communicate with each other. The system bus can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, an industry standard architecture (ISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the figure, but this does not mean that there is only one bus or one type of bus.
[0150] Optionally, the memory 601 can be either standalone or integrated with the processor 603. When the memory 601 is set up independently, it is connected to the processor 603 via a system bus.
[0151] This application also provides a chip for executing instructions, which is used to execute the technical solution of the method for obtaining industry-related data in the above embodiments.
[0152] This application also provides a computer-readable storage medium storing computer instructions. When these computer instructions are executed by a processor, they are used to implement the technical solution of the method for acquiring industry-related data described in the above embodiments. Specifically, when the computer instructions are executed by a processor, the electronic device can perform the technical solution of the method for acquiring industry-related data described in the above embodiments.
[0153] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, it can implement the technical solution of the method for obtaining industry-related data in the above embodiments.
[0154] The aforementioned computer-readable storage media can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage media can be any available medium accessible to a general-purpose or special-purpose computer.
[0155] An exemplary computer-readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the computer-readable storage medium can also be a component of the processor. The processor and the computer-readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the computer-readable storage medium can exist as discrete components in an electronic control unit or main control device; this application does not limit this.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0157] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.
[0158] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.
[0159] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0160] It should be understood that the steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0161] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for acquiring industry-related data, characterized in that, include: Obtain the industry chain map corresponding to the target industry and the supply chain map corresponding to the target region; the supply chain map includes multiple first nodes and multiple first directed edges, the first nodes are nodes corresponding to business entities in the supply chain, and the first directed edges connect two first nodes to indicate the business relationship between the business entities corresponding to the two first nodes; the business entity is at least one of the following: an enterprise, a business in an enterprise, or a subsidiary of an enterprise; the industry chain map includes multiple second nodes, the second nodes are nodes corresponding to each industrial link in the industry chain. A classification tree model is used to determine a third node that matches a second node in the industrial chain map from multiple first nodes in the supply chain map. Based on the node information of the third node, the third node is quantized to obtain the node feature value of the third node; The node information of the third node is the profile information of the business entity corresponding to the third node in the supply chain map; Obtain business transaction information for each group of business entities in multiple groups of business entities, and calculate the business association value between two business entities in each group of business entities based on the business transaction information of each group of business entities; wherein, the two business entities in each group of business entities are the business entities corresponding to two third nodes connected by a first directed edge in the supply chain graph; The industry-related data of the target industry is obtained by aggregating and calculating the node feature value of the third node and the business association value of two business entities in each group of business entities; wherein, the industry-related data is used to indicate the industry operation status of the target industry in the target region.
2. The method according to claim 1, characterized in that, The step of using a classification tree model to determine a third node that matches a second node in the supply chain map from multiple first nodes includes: For each of the multiple first nodes in the supply chain graph, obtain the identity information of the business entity corresponding to the first node; Using the classification tree model, the multiple first nodes are classified according to the identity information of the business entities corresponding to the multiple first nodes to obtain classification results; the classification results are used to indicate the industry segment to which the business entity corresponding to the first node belongs; Based on the classification results of the plurality of first nodes, a third node that matches the second node of the industry chain map is selected from the plurality of first nodes; wherein, the industry segment to which the business entity corresponding to the third node belongs matches the industry segment corresponding to at least one second node in the industry chain map.
3. The method according to claim 1, characterized in that, The node information of the third node includes quantitative analysis parameters, which include multiple market parameters of the business entity corresponding to the third node. These multiple market parameters include at least one of the following: sales revenue, transaction volume, market share, profit margin, market popularity, ESG index, and market share growth rate of the business entity corresponding to the third node. The ESG index is an environmental, social, and governance score. The step of quantizing the third node based on its node information to obtain its node feature value includes: The weighted sum of the various market parameters is calculated according to a preset weight combination to obtain the node feature value of the third node; The preset weight combination includes multiple first weights, each of which corresponds one-to-one with the various market parameters, and the sum of the multiple first weights equals 1; the various market parameters are after standardization.
4. The method according to claim 1, characterized in that, The node information of the third node includes qualitative analysis parameters, which include entity tags of the business entity corresponding to the third node. The entity tags include one or more positive tags and one or more negative tags. The positive tags indicate that the business entity corresponding to the third node has high stability and positive development capabilities or potential, while the negative tags indicate that the business entity corresponding to the third node has poor stability and no development potential. The step of quantizing the third node based on its node information to obtain its node feature value includes: The node feature value of the third node is calculated based on the number of positive labels and the number of negative labels in the qualitative analysis parameters.
5. The method according to claim 1, characterized in that, The node information of the third node includes quantitative analysis parameters and qualitative analysis parameters; The quantitative analysis parameters include multiple market parameters of the business entity corresponding to the third node. These multiple market parameters include at least one of the following: sales revenue, transaction volume, market share, profit margin, market popularity, ESG index, and market share growth rate of the business entity corresponding to the third node. The ESG index is an environmental, social, and governance score. The qualitative analysis parameters include entity tags of the business entities corresponding to the third node. The entity tags include one or more positive tags and one or more negative tags. The positive tags indicate that the business entities corresponding to the third node have high stability and positive development capabilities or potential. The negative tags indicate that the business entities corresponding to the third node have poor stability and no development potential. The step of quantizing the third node based on its node information to obtain its node feature value includes: The weighted sum of the various market parameters is calculated according to a preset weight combination to obtain the node feature value of the third node; wherein, the preset weight combination includes multiple first weights, each of which corresponds one-to-one with the various market parameters, and the sum of the multiple first weights equals 1; the various market parameters are after standardization. The node feature value of the third node is increased according to the adjustment coefficient corresponding to the number of positive labels in the qualitative analysis parameters; the node feature value of the third node is decreased according to the adjustment coefficient corresponding to the number of negative labels in the qualitative analysis parameters.
6. The method according to any one of claims 1, 3-5, characterized in that, After performing quantization calculations on the third node based on its node information to obtain the node feature value of the third node, the method further includes: Obtain the node information of the second node that matches the third node in the industry chain map; the node information of the second node is the information of the industry chain corresponding to the second node; The adjustment coefficient corresponding to the node information of the second node is determined from a preset mapping table and used as the adjustment coefficient corresponding to the third node; wherein, the preset mapping table includes the identifiers of multiple industry chains and the adjustment coefficient corresponding to each industry chain; According to the adjustment coefficient corresponding to the first type of industrial chain in the adjustment coefficient corresponding to the third node, the node characteristic value of the third node is increased; wherein, the multiple industrial chains include the first type of industrial chain and the second type of industrial chain, the first type of industrial chain is the industrial chain in the target region that is in an advantageous position in terms of economy and / or development, and the second type of industrial chain is the industrial chain in the target region that is in a disadvantageous position in terms of economy and / or development. According to the adjustment coefficient corresponding to the second type of industrial chain in the adjustment coefficient corresponding to the third node, the node characteristic value of the third node is lowered.
7. The method according to any one of claims 1, 3-5, characterized in that, The business transaction information includes one or more of the following: business cooperation information, transaction frequency information, capital flow information, and logistics information between two business entities in each group of business entities. Each item of the business transaction information corresponds to a second weight. The step of calculating the business association value between two business entities in each group of business entities based on their business transaction information includes: Based on one or more business transaction information items of each group of business entities and the second weight corresponding to each item of business transaction information, the business association value of the two business entities in each group of business entities is calculated by weighting.
8. The method according to any one of claims 1, 3-5, characterized in that, The process involves aggregating and calculating the industry-related data of the target industry based on the node feature value of the third node and the business association values of two business entities in each group of business entities, including: The node feature values of all the third nodes and the business association values of the two business entities in each group of business entities are added together to obtain the industry-related data of the target industry.
9. An apparatus for acquiring industry-related data, characterized in that, include: The graph acquisition module is used to acquire the industrial chain graph corresponding to the target industry and the supply chain graph corresponding to the target region. The supply chain graph includes multiple first nodes and multiple first directed edges. The first nodes are nodes corresponding to business entities in the supply chain. The first directed edges connect two first nodes and are used to indicate the business relationship between the business entities corresponding to the two first nodes. The business entity is at least one of the following: an enterprise, a business within an enterprise, or a subsidiary of an enterprise. The industrial chain graph includes multiple second nodes, which are nodes corresponding to various industrial links in the industrial chain. The node matching module is used to determine, from multiple first nodes of the supply chain graph, a third node that matches the second node of the industrial chain graph using a classification tree model; The node calculation module is used to perform quantitative calculations on the third node based on the node information of the third node to obtain the node feature value of the third node; The node information of the third node is the profile information of the business entity corresponding to the third node in the supply chain map; The association calculation module is used to obtain the business transaction information of each group of business entities in multiple groups of business entities, and calculate the business association value of two business entities in each group of business entities based on the business transaction information of each group of business entities; wherein, the two business entities in each group of business entities are the business entities corresponding to two third nodes connected by a first directed edge in the supply chain graph. The industry assessment module is used to perform aggregate calculations based on the node characteristic values of the third node and the business association values of two business entities in each group of business entities to obtain industry-related data of the target industry; wherein, the industry-related data is used to indicate the industry operation status and industry development indicators of the target industry in the target region.
10. An electronic device, characterized in that, include: A memory and at least one processor; the memory is communicatively connected to the processor; the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, are used to implement the method as described in any one of claims 1-8.
12. A computer program product, characterized in that, When the computer program product is run on a computer / executed by the computer's processor, it implements the method as described in any one of claims 1-8.