Text information processing method and device and related equipment
By generating an enterprise's industrial chain relationship diagram and calculating the similarity, the problem of the existing technology being unable to accurately measure the industrial distribution of enterprises is solved, more reliable analysis results are achieved, and the understanding of the industrial chain structure and correlation is enhanced.
Patent Information
- Application Number
- CN202510691785.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, enterprise industrial distribution analysis lacks effective capture of the structural characteristics of the industrial chain, cannot accurately reflect the distribution of enterprises in multiple industrial chains, and ignores the technological correlation between different industrial chains, resulting in inaccurate evaluation results.
By obtaining the industrial data and scope attribute data of enterprises, an industrial chain relationship graph is generated, graph relationship features are extracted, and the similarity of industrial chain features is calculated. By using large language models and graph convolution calculation methods, internal and external features of the industrial chain are integrated to generate more accurate industrial distribution information.
It improves the reliability of industrial distribution analysis, can more accurately measure the distribution and diversification of enterprises in multiple industrial chains, and enhances the understanding of the internal structure and correlation of industrial chains.
Smart Images

Figure CN120633631A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of text processing technology, and in particular to a text information processing method, apparatus, and related equipment. Background Art
[0002] Currently, simple text matching or basic keyword extraction methods are often used to represent a company's industrial distribution. However, the vector representations generated by these methods lack effective capture of the structural characteristics of the company's industrial chain, making it difficult to accurately reflect the company's industrial distribution across multiple industrial chains. Furthermore, when measuring a company's industrial distribution, many rely on crude industry classifications, ignoring the technological connections between different industrial chains. This makes it difficult to accurately assess a company's industrial distribution and degree of diversification. Summary of the Invention
[0003] The present disclosure provides a text information processing method, apparatus, and related equipment, which improve the reliability of industry distribution analysis results at least to a certain extent.
[0004] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0005] According to one aspect of the present disclosure, a text information processing method is provided, including: obtaining text information, the text information including industry data and industry scope attribute data of an object, wherein the industry data includes multiple industries, content descriptions of each industry, and the industrial chain to which each industry belongs, and the industry scope attribute data includes multiple industries of the object; determining multiple target industries from the industry data and the industry scope attribute data, and determining the frequency of occurrence of each target industry in the industry scope attribute data, wherein each target industry is an industry included in both the industry data and the industry scope attribute data; determining the content description of each target industry and the industrial chain to which each target industry belongs from the industry data, and extracting embedded features of the content description of each target industry; using the industry data to generate a relationship graph of the industrial chain to which each target industry belongs, and extracting graph relationship features of the relationship graph of the industrial chain to which each target industry belongs; determining the industrial chain features corresponding to the industrial chain to which each target industry belongs based on the embedded features and frequencies corresponding to each target industry and the graph relationship features corresponding to the industrial chain to which each target industry belongs; determining the similarity of the industrial chain features corresponding to the industrial chain to which each target industry belongs, and determining the industrial distribution information of the object based on the determined similarity.
[0006] In one embodiment of the present disclosure, based on the embedding features and frequencies corresponding to each target industry and the graph relationship features corresponding to the industrial chain to which each target industry belongs, the industrial chain features corresponding to the industrial chain to which each target industry belongs are determined, including: grouping the target industries belonging to the same industrial chain together; performing graph convolution calculations on the embedding features and frequencies corresponding to the target industries belonging to the same industrial chain and the graph relationship features corresponding to the industrial chain to obtain the industrial chain features corresponding to the industrial chain.
[0007] In one embodiment of the present disclosure, industrial data is used to generate a relationship diagram of the industrial chain to which each target industry belongs, including: determining all industries under the industrial chain to which each target industry belongs and the level of each industry from the industrial data; for the industrial chain to which each target industry belongs: generating nodes and node names corresponding to each industry under the industrial chain; generating identification numbers of nodes corresponding to each industry based on the level of each industry; and connecting the nodes corresponding to each industry based on the level of each industry to obtain a relationship diagram of the industrial chain.
[0008] In one embodiment of the present disclosure, based on the level of each industry, the nodes corresponding to each industry are connected to obtain a relationship diagram of the industrial chain, including: based on the level of each industry, the nodes corresponding to each industry are connected to obtain edges between the nodes corresponding to each industry and the nodes corresponding to the industries of the previous level of each industry; based on the frequency corresponding to each industry, the weights of the edges between the nodes corresponding to each industry and the nodes corresponding to the industries of the previous level of each industry are determined to obtain a relationship diagram of the industrial chain.
[0009] In one embodiment of the present disclosure, the similarity of the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs is determined, including: calculating the similarity between the industrial chain characteristics corresponding to any two industrial chains in the industrial chain to which each target industry belongs, so as to obtain a similarity matrix corresponding to the industrial chain to which each target industry belongs; taking an upper triangular matrix without a diagonal in the similarity matrix, and using the similarity contained in the upper triangular matrix as the determined similarity.
[0010] In one embodiment of the present disclosure, determining the industry distribution information of the object based on the determined similarities includes: calculating an average value of the determined similarities; and determining the industry distribution information based on the average value.
[0011] According to another aspect of the present disclosure, a text information processing device is provided, characterized in that it includes: an acquisition module, configured to acquire text information, the text information including industry data and industry-wide attribute data of an object, wherein the industry data includes multiple industries, content descriptions of each industry, and the industrial chain to which each industry belongs, and the industry-wide attribute data includes multiple industries of the object; a first determination module, configured to determine multiple target industries from the industry data and the industry-wide attribute data, and determine the frequency of occurrence of each target industry in the industry-wide attribute data, wherein each target industry is an industry included in both the industry data and the industry-wide attribute data; a second determination module, configured to determine multiple target industries from the industry data Determine the content description of each target industry and the industrial chain to which each target industry belongs, and extract the embedded features of the content description of each target industry; the generation module is configured to use industrial data to generate a relationship graph of the industrial chain to which each target industry belongs, and extract the graph relationship features of the relationship graph of the industrial chain to which each target industry belongs; the third determination module is configured to determine the industrial chain features corresponding to the industrial chain to which each target industry belongs based on the embedded features and frequency corresponding to each target industry and the graph relationship features corresponding to the industrial chain to which each target industry belongs; the evaluation module is configured to determine the similarity of the industrial chain features corresponding to the industrial chain to which each target industry belongs, and determine the industrial distribution information of the object based on the determined similarity.
[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any of the above methods by executing the executable instructions.
[0013] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any of the above methods is implemented.
[0014] According to another aspect of the present disclosure, a computer program product is provided, including computer instructions stored in a computer-readable storage medium, and the computer instructions implement operating instructions of any of the above methods when executed by a processor.
[0015] In the embodiments of the present disclosure, industrial data is used to generate a relationship graph of the industrial chains to which each target industry belongs, and the graph relationship features of the relationship graph of each industrial chain are extracted, and then the industrial chain features corresponding to each industrial chain are determined, and finally the industrial distribution information is determined, so as to solve the problem in the existing technology that conventional vector representation cannot fully represent the industrial chain information and ignores the technical correlation between different industrial chains, resulting in the inability to accurately measure the industrial distribution, thereby improving the reliability of the industrial distribution analysis results.
[0016] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0018] Figure 1 A schematic diagram of a text information processing system in an embodiment of the present disclosure is shown.
[0019] Figure 2 A flowchart of a text information processing method in an embodiment of the present disclosure is shown.
[0020] Figure 3 A flow chart of a method for calculating industrial chain characteristics in an embodiment of the present disclosure is shown.
[0021] Figure 4 A flow chart illustrating a method for generating an industrial chain relationship diagram in an embodiment of the present disclosure is shown.
[0022] Figure 5 A flowchart illustrating another method for generating an industrial chain relationship diagram in an embodiment of the present disclosure is shown.
[0023] Figure 6 A flowchart of a similarity calculation method in an embodiment of the present disclosure is shown.
[0024] Figure 7 A flowchart of a method for determining industry distribution information in an embodiment of the present disclosure is shown.
[0025] Figure 8 A schematic diagram showing an industrial chain relationship diagram in an embodiment of the present disclosure is shown.
[0026] Figure 9 A schematic diagram showing another industrial chain relationship diagram in an embodiment of the present disclosure.
[0027] Figure 10 A schematic diagram of a text information processing device in an embodiment of the present disclosure is shown.
[0028] Figure 11 A schematic diagram of an electronic device provided in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0030] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0031] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0033] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0034] It should be pointed out that, in the absence of conflict, the embodiments of the present disclosure and the technical features therein may be combined with each other.
[0035] For ease of understanding, several terms involved in this disclosure are explained below:
[0036] Torch-Geometric is an open-source library based on PyTorch, designed for Graph Neural Networks (GNNs). It provides a rich set of tools and modules for processing graph-structured data, such as node classification, graph classification, link prediction, and other tasks.
[0037] The Multilayer Perceptron (MLP) is a classic feedforward neural network consisting of multiple fully connected layers, typically including an input layer, hidden layers, and an output layer. MLP is one of the foundational models of deep learning and is widely used in classification, regression, and other tasks.
[0038] Large language models are deep learning models with a large number of parameters, typically based on the Transformer architecture, that can understand and generate natural language. These models excel in natural language processing (NLP) tasks such as text generation, translation, and question answering.
[0039] Categorical encoding is a technique for converting categorical variables (such as product ID, category labels, etc.) into numerical form.
[0040] mapping_id usually represents a unique identifier for a specific mapping relationship.
[0041] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.
[0042] Figure 1 A schematic diagram of a text information processing system in an embodiment of the present disclosure is shown.
[0043] The text information processing system includes an object server 101 and a data processing center 102. Object server 101 stores and manages internal data of objects, including industry-specific attribute data. Data processing center 102 stores or can access industry data and determines industry distribution information for the objects based on the industry-specific attribute data and the industry data.
[0044] Among them, an application program can be installed in the data processing center 102 to perform: obtaining text information, the text information includes industry data and industry scope attribute data of the object, wherein the industry data includes multiple industries, content descriptions of each industry and the industrial chain to which each industry belongs, and the industry scope attribute data includes multiple industries of the object; determining multiple target industries from the industry data and the industry scope attribute data, and determining the frequency of occurrence of each target industry in the industry scope attribute data, wherein each target industry is an industry included in both the industry data and the industry scope attribute data; determining the content description of each target industry and the industrial chain to which each target industry belongs from the industry data, and extracting the embedded features of the content description of each target industry; using the industry data to generate a relationship graph of the industrial chain to which each target industry belongs, and extracting the graph relationship features of the relationship graph of the industrial chain to which each target industry belongs; based on the embedded features and frequencies corresponding to each target industry and the graph relationship features corresponding to the industrial chain to which each target industry belongs, determining the industrial chain features corresponding to the industrial chain to which each target industry belongs; determining the similarity of the industrial chain features corresponding to the industrial chain to which each target industry belongs, and determining the industrial distribution information of the object based on the determined similarity.
[0045] The object server 101 and the data processing center 102 are connected via a communication network. Optionally, the communication network is a wired network or a wireless network.
[0046] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0047] Figure 2 A flowchart of a text information processing method according to an embodiment of the present disclosure is shown. Figure 2 As shown, the following steps are included:
[0048] S201, obtaining text information, the text information including industry data and industry-scope attribute data of an object, wherein the industry data includes multiple industries, a description of each industry, and the industrial chain to which each industry belongs, and the industry-scope attribute data includes multiple industries of the object;
[0049] Industry data, available from public databases or websites, is a collection of publicly known industries, their descriptions, and the industry chains to which they belong. Industry scope attribute data refers to the business scope of an entity (e.g., a company), specifically encompassing the multiple industries it operates in. This data is used to identify the industries a company is currently operating in or plans to enter.
[0050] S202, determining multiple target industries from the industry data and the industry scope attribute data, and determining the frequency of each target industry appearing in the industry scope attribute data, wherein each target industry is an industry included in both the industry data and the industry scope attribute data;
[0051] Target industries are those that appear in both industry data and industry-wide attribute data, determined through comparative analysis. They are key points in measuring a company's diversified operations.
[0052] S203, determining the content description of each target industry and the industrial chain to which each target industry belongs from the industrial data, and extracting the embedded features of the content description of each target industry;
[0053] The industrial chain feature is an embedded vector, which is a feature representation obtained by processing the content description of the target industry using a large language model, and is usually expressed in the form of a high-dimensional vector.
[0054] S204, using the industry data to generate a relationship graph of the industrial chain to which each target industry belongs, and extracting graph relationship features of the relationship graph of the industrial chain to which each target industry belongs;
[0055] Determine from the industrial data all the industries under the industrial chain to which each target industry belongs and the level of each industry, and then generate a relationship diagram of the industrial chain based on all the contents of the industrial chain to which each target industry belongs. The large model can be used to extract the graph relationship features of the relationship diagram of the industrial chain to which each target industry belongs.
[0056] S205, determining the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs based on the embedding characteristics and frequencies corresponding to each target industry and the graph relationship characteristics corresponding to the industrial chain to which each target industry belongs;
[0057] The industrial chain characteristics corresponding to the industrial chain to which the target industry belongs are calculated using the embedding features and frequencies corresponding to each target industry and the graph relationship features corresponding to the industrial chain to which the target industry belongs.
[0058] S206 , determining the similarity of the industrial chain characteristics corresponding to the industrial chains to which each target industry belongs, and determining the industrial distribution information of the object based on the determined similarity.
[0059] Industrial distribution information includes information on whether a company's industrial distribution is concentrated, dispersed, single, or diversified. This information can be used to measure the degree of diversification within a company's operations. The similarity between the industry chain characteristics corresponding to each target industry is calculated, and ultimately, the target's industrial distribution information can be determined based on the determined similarity.
[0060] The disclosed embodiment uses the above-mentioned technical means to interpret the content in the industrial data using a large language model and encode it into embedded features. Next, the target industry that matches the industrial data is identified from the industry scope attribute data, and its frequency of occurrence is counted. Then, a relationship graph of the industrial chain is constructed, and the graph relationship features are extracted. Based on the graph relationship features of the industrial chain, the features of the industrial chain to which each target industry belongs are calculated, and the similarity between these industrial chain features is further calculated. Finally, the industrial distribution information of the enterprise is determined based on the similarity results. This solves the problem in the prior art that conventional vector representations cannot fully represent industrial chain information and ignores the technical correlation between different industrial chains, resulting in the inability to accurately measure industrial distribution, thereby improving the reliability of industrial distribution analysis results.
[0061] Figure 3 A flow chart of a method for calculating industrial chain characteristics according to an embodiment of the present disclosure is shown. Figure 3 As shown, the following steps are included:
[0062] S301, group together target industries belonging to the same industrial chain;
[0063] S302, performing graph convolution calculation on the embedded features and frequencies corresponding to the target industries belonging to the same industrial chain, as well as the graph relationship features corresponding to the industrial chain, to obtain the industrial chain features corresponding to the industrial chain. Graph convolution calculation: A feature extraction method based on graph structured data, which can combine node features and their connection relationships to generate new feature representations. In this disclosure, it is mainly used to integrate the embedded features, frequency information, and graph relationship features of each target industry within the industrial chain to obtain a feature vector representing the characteristics of the entire industrial chain.
[0064] In this embodiment, the target industries belonging to the same industrial chain are first grouped together based on the industrial chain information to which the target industries belong. For each specific industrial chain, the embedded features and their frequencies of all target industries in the group are collected, and combined with the relationship graph features of the industrial chain. Then, this information is processed using the graph convolution calculation method to obtain the industrial chain characteristics that reflect the characteristics of the entire industrial chain. This process involves using the open source tool torch-geometric or other large language models to implement graph convolution calculations, and finally generate industrial chain characteristics. The above technical means enhance the understanding of the internal structure of the industrial chain because it not only takes into account the characteristics of a single industry, but also analyzes the correlation between these industries.
[0065] For example, suppose a company's industry-wide attribute data indicates it's involved in 5G communications equipment manufacturing and related RF device development. Both areas fall within the 5G industry chain. By performing graph convolution on the embedded features and frequency information of specific nodes within the 5G industry chain, combined with the relationship graph features of the 5G industry chain, a comprehensive 5G industry chain profile is derived. This enables the company to accurately grasp its specific position within the 5G industry chain, including key information such as the closeness of its connections with upstream and downstream companies, thereby optimizing its resource allocation and development strategy.
[0066] In an optional embodiment, another machine learning model, such as a multi-layer perceptron (MLP), is used for feature fusion. Specifically, after identifying target industries belonging to the same industrial chain, their embedded features, frequency information, and graph relationship features are directly input into the MLP, and the industrial chain characteristics are output through the trained MLP model. This method can also effectively integrate information from different sources to generate a feature vector that can reflect the overall characteristics of the industrial chain. This embodiment maintains the effective capture of industrial chain structural characteristics while providing different technical implementation paths.
[0067] In an optional embodiment, each target industry is divided into multiple groups, wherein each group corresponds to an industrial chain, and each group contains at least one target industry; graph convolution calculation is performed on the characteristics and frequency of the target industry in each group and the graph relationship characteristics corresponding to the industrial chain corresponding to the group to obtain the graph relationship characteristics corresponding to the industrial chain corresponding to the group.
[0068] The above technical means enhance the understanding of the internal structure of the industrial chain because it not only takes into account the characteristics of individual industries, but also analyzes the correlation between these industries.
[0069] In an optional embodiment, each target industry is divided into multiple groups, where each group corresponds to an industrial chain and each group contains at least one target industry; graph convolution calculation is performed on the characteristics and frequency of the target industry in each group to obtain the graph relationship characteristics corresponding to the industrial chain corresponding to each group.
[0070] The above technical means are used to improve the efficiency of computing the graph relationship features corresponding to the industrial chain.
[0071] Figure 4 A flow chart showing a method for generating an industrial chain relationship diagram in an embodiment of the present disclosure is shown. Figure 4 As shown, the following steps are included:
[0072] S401, determining all industries under the industrial chain to which each target industry belongs and the level of each industry from the industrial data;
[0073] S402, for the industrial chain to which each target industry belongs:
[0074] S403, generating nodes and node names corresponding to each industry under the industrial chain;
[0075] S404, generating identification numbers of nodes corresponding to each industry based on the level of each industry;
[0076] S405 , based on the level of each industry, connecting the nodes corresponding to each industry to obtain a relationship diagram of the industry chain.
[0077] In the industry chain relationship diagram, each industry is represented as a node, and the node name is the specific name or description of the industry. For example, in the 5G industry chain, "RF Devices" is a specific node, and its node name is "RF Devices." Identification number: For ease of management and identification, each industry's corresponding node is assigned a unique identification number (mapping_id). This number is determined based on the industry's level in the industry chain. Generally, the highest level of the industry chain has the highest number, and the numbers of the nodes in the lower levels increase in sequence.
[0078] In this embodiment, for each specific industrial chain, nodes and node names corresponding to each industry under the industrial chain are generated, and identification numbers are assigned to each industry based on its level in the industrial chain. Then, these nodes are connected according to the hierarchical relationship of each industry to construct a relationship diagram of the industrial chain, which can be connected according to the hierarchical relationship from high to low. In this process, the nodes of the upper level always have a smaller identification number than the nodes of the lower level, ensuring the logic and hierarchy of the relationship diagram structure. The above technical means enhance the understanding of the internal structure of the industrial chain, and the relationship between the various parts of the industrial chain is clearly displayed through clear node connections and hierarchical divisions.
[0079] For example, suppose a company's industry-wide attribute data indicates involvement in 5G communications equipment manufacturing and RF device development. First, all industry information related to these two areas is extracted from the industry data, including their position and rank within the 5G industry chain. Next, nodes are created for each industry within the 5G industry chain, such as "RF Devices" and "Optical Modules," and identification numbers are assigned based on their rank. For example, "5G Industry Chain," as the top-level node, receives the highest identification number, 74, while the specific industry nodes beneath it are numbered sequentially. Finally, based on this hierarchical information, each node is connected to form a complete 5G industry chain relationship diagram. This allows companies to clearly understand their position within the 5G industry chain and their upstream and downstream relationships, helping to formulate more precise development strategies.
[0080] In an optional embodiment, similarities between industries are calculated, and the connection method between nodes is determined based on the similarities. For example, if there is a high degree of technological or market correlation between two industries, they will be closely connected in the relationship diagram, even if they are at different levels in the traditional sense. This method can help reveal potential, non-traditional connections in the industrial chain, providing companies with a new perspective to understand the dynamic changes in their industry chain. This method can also effectively construct an industrial chain relationship diagram, but it provides a different perspective to understand and analyze the industrial chain structure.
[0081] Figure 5 A flow chart showing another method for generating an industrial chain relationship diagram in an embodiment of the present disclosure is shown. Figure 5 As shown, the following steps are included:
[0082] S501, based on the level of each industry, connect the nodes corresponding to each industry to obtain the edges between the nodes corresponding to each industry and the nodes corresponding to the industry at the next higher level;
[0083] S502 , based on the frequencies corresponding to each industry, determine the weights of the edges between the nodes corresponding to each industry and the nodes corresponding to the industries at the next higher level, and obtain a relationship diagram of the industry chain.
[0084] In this example, the nodes corresponding to each industry are connected in descending order according to their hierarchy, thus generating edges between each node and the node corresponding to the industry in the next higher level. The frequency corresponding to each industry can be used as the weight of the edge between the node corresponding to that industry and the node corresponding to the industry in the next higher level, thus generating a relationship diagram for the industry chain. Generating an industry chain relationship diagram through the above technical means not only demonstrates the internal structure and hierarchical relationships of the industry chain, but also reflects the different strategic importance of different industries to the enterprise.
[0085] In addition, the preset multiples of the frequency corresponding to each industry may be used as the weight of the edge between the node corresponding to the industry and the node corresponding to the industry at the previous level.
[0086] Figure 6 A flowchart of a similarity calculation method according to an embodiment of the present disclosure is shown. Figure 6 As shown, the following steps are included:
[0087] S601, calculating the similarity between the industrial chain features corresponding to any two industrial chains in the industrial chain to which each target industry belongs, to obtain a similarity matrix corresponding to the industrial chain to which each target industry belongs;
[0088] S602 : Take an upper triangular matrix without a diagonal line in the similarity matrix, and use the similarity included in the upper triangular matrix as the determined similarity.
[0089] Each element in the similarity matrix represents the similarity between two industry chain features. For n industry chain features, the similarity matrix is an n*n square matrix. The elements on the diagonal represent the similarity between a certain industry chain feature and itself (usually 1), while the off-diagonal elements represent the similarity between different industry chain features.
[0090] Upper triangular matrix: Extract the upper triangular part without diagonal lines from the similarity matrix, that is, only retain the upper right part of the matrix. This part contains the similarity information between all different industrial chain features, avoiding repeated calculations (because the similarity matrix is symmetrical).
[0091] In this embodiment, the similarity between the industrial chain features corresponding to any two industrial chains in the industrial chain to which each target industry belongs is first calculated to generate a similarity matrix. Then, the upper triangular matrix without diagonal lines is taken from the similarity matrix, and all similarity values contained in the upper triangular matrix are used as the determined similarity. This evaluates the degree of similarity between the different industrial chains in which the enterprise is involved, helping the enterprise understand the true situation of its diversified operations. Specifically, by analyzing these similarity values, it is possible to identify in which areas the enterprise has overlaps or associations, and in which areas it has achieved true diversified development. Through the above-mentioned technical means, the understanding of the similarities and differences between the different industrial chains in which the enterprise is involved is enhanced, and the accuracy of measuring the degree of enterprise diversification is improved.
[0092] In an optional embodiment, cluster analysis is performed on all industry chain characteristics to group multiple industry chains. Within each group, the average similarity is calculated. This technology simplifies the similarity calculation process, highlighting highly similar industry chain clusters, and providing companies with a more macro perspective to understand and plan their diversification strategies.
[0093] Figure 7 A flow chart of a method for determining industry distribution information in an embodiment of the present disclosure is shown. Figure 7 As shown, the following steps are included:
[0094] S701, calculating the average value of the determined similarities;
[0095] S702: Determine industry distribution information based on the average value. In this embodiment, the average value of the determined similarities is first calculated, that is, the average of all similarity values extracted from the upper triangular portion of the similarity matrix. Then, based on this average value, the enterprise's industry distribution information is determined. This technical approach provides a comprehensive evaluation index by quantifying the similarities between different industry chain characteristics, thereby improving the accuracy of assessing the industry distribution status of enterprises.
[0096] For example, let's assume a company called "X Technology" is involved in three different industry chains: 5G communication equipment manufacturing, RF device development, and cloud computing. We calculate the similarity between the characteristics of these three industry chains, obtaining a 3x3 similarity matrix. Assume the results are as follows:
[0097]
[0098] Take the upper triangular matrix without diagonal values, namely 0.83, 0.76, and 0.92, and calculate the average value of the upper triangular matrix to get 0.837. 1-0.837=0.163 is the company's industrial distribution information. The higher the score, the more balanced the company's industrial distribution and the higher the degree of business diversification.
[0099] In an alternative embodiment, each similarity value is weighted based on the importance of different industry chains to the enterprise (for example, by revenue ratio or investment scale), and then a weighted average similarity is calculated. This technical approach not only considers the similarities between industry chains but also incorporates the actual investment of the enterprise in each industry chain, providing a more realistic enterprise diversification assessment.
[0100] In one embodiment of the present disclosure, the content description of each industry in the industry data is generated using a large language model.
[0101] Through the above technical means, a large language model is used to provide detailed descriptions of industrial chain content, making the content representation of the industrial chain more accurate and comprehensive. This provides a solid foundation for subsequent feature extraction of industrial chain content, graph convolution calculation, and industrial distribution assessment.
[0102] Taking the "RF structural parts" industry chain as an example: "5G industry chain" refers to the complete industrial system surrounding the research and development, production, application and services of the fifth generation mobile communication technology (5G). Among them, "device materials" is a link in the industry chain, which covers all kinds of materials and compounds needed to manufacture various electronic devices. "RF devices" are a specific type of product under "device materials", mainly referring to devices that work within the radio frequency range, which are responsible for signal transmission, reception, processing and conversion. "RF structural parts" refers specifically to those structural parts in RF devices that are used to support, fix or connect RF components.
[0103] So the levels from high to low are: 5G industry chain > device materials > RF devices > RF structural parts.
[0104] RF structural components have a wide range of uses, mainly including the following aspects:
[0105] Improve signal quality: Through precise mechanical design and material selection, RF structural components can ensure stable signal transmission and reduce interference and loss.
[0106] Enhanced system performance: The design and optimization of RF structural components help improve the performance of the entire 5G system, such as increasing coverage and improving data transmission rates.
[0107] Reduce costs: Efficient RF structural component design can reduce the use of raw materials, thereby reducing production costs.
[0108] Simplify the assembly process: Standardized and modular RF structural component design can simplify the equipment assembly process and improve production efficiency.
[0109] In one embodiment, taking the 5G industry chain as an example, the identification numbers of the nodes corresponding to each industry in the 5G industry chain are generated, and the category codes product_id and mapping_id can be generated. The product_id facilitates unique identification when matching with the business scope content. The mapping_id is an internal sequential number of the industry chain. The number of all nodes in the industry chain can be added by 1 as the identification number of the industry chain, and 1 represents the node of the industry chain itself. For example, the 5G industry chain has a total of 73 content nodes, so the "5G industry chain" is numbered 74. Other industries in the 5G industry chain can be numbered in ascending order, starting from 0. The general principle is that the number of the upper level must be smaller than the number of the lower level. Table 1 takes the first 10 numbers of the 5G industry chain as an example, as follows:
[0110]
[0111]
[0112] Table 1
[0113] Figure 8 A schematic diagram showing an industrial chain relationship diagram in an embodiment of the present disclosure is shown as follows: Figure 8 As shown:
[0114] This industry chain relationship diagram is about the 5G industry chain. The 5G industry chain has a mapping_id of 74; the node name is optical devices and modules, and its mapping_id is 1; the node name is optical module, and its mapping_id is 2; the node name is optical transceiver interface component, and its mapping_id is 3; the node name is fiber laser device, and its mapping_id is 4; the node name is fiber connection device, and its mapping_id is 5; the node name is fiber transceiver, and its mapping_id is 6; the node name is fiber ceramic ferrule and sleeve, and its mapping_id is 7; the node name is RF device, and its mapping_id is 8; the node name is power amplifier component, and its mapping_id is 9. In addition, mapping_id 0 can represent a concept between the 5G industry chain and optical devices and modules, such as device materials, etc. The other mapping_ids each correspond to a node of an industry. The 5G industry chain mapping_id of 74 means that there are 74 nodes in the 5G industry chain. Figure 8 Some nodes are not shown and will not be described here.
[0115] Figure 9 A schematic diagram showing another industrial chain relationship diagram in an embodiment of the present disclosure is shown. Figure 9 As shown:
[0116] This industry chain relationship diagram shows the 5G industry chain after the industry scope attribute data of a company is matched. The matched nodes in the 5G industry chain are marked in gray. The 5G industry chain mapping_id is 74, which means that there are 74 nodes in the 5G industry chain. Figure 9 Some nodes are not shown, so we will not go into details here.
[0117] The number of hits on a node is the frequency of the target industry represented by the node. The details are as follows:
[0118] First, the industry-wide attribute data is segmented. These words are matched to specific nodes in the industry chain graph to determine the intensity of the company's activities in different industry chains. For example, the business scope text of "X Technology" may include the following words: "chips," "wireless telecommunications business," "optical modules," etc. These words hit the nodes in the relationship graph of the corresponding industry chain (equivalent to the frequency of each target industry in the industry-wide attribute data), as follows:
[0119] Node 30: The node name is "chip" and the number of hits is 3.
[0120] Node 44: The node name is "Wireless Telecommunications Service" and the number of hits is 1.
[0121] Node 25: The node name is "Optical Module" and the number of hits is 2.
[0122] Node 3: The node name is "RF structural parts" and the number of hits is 1.
[0123] Node 33: The node name is "Fiber Laser Device" and the number of hits is 4.
[0124] In actual operation, these hit nodes will be marked (such as using gray markings) to facilitate subsequent analysis.
[0125] The weight of the edge between each node and its parent node is determined by the node name and the number of times the node is hit. The specific calculation method is as follows:
[0126] Node 30: The weight of the edge from node 27 to node 30 is 3 (because "chip" hits 3 times).
[0127] Node 44: The weight of the edge from node 43 to node 44 is 1 (because "Wireless Telecommunications Service" is hit once).
[0128] Node 25: The weight of the edge from node 24 to node 25 is 2 (because the “light module” is hit 2 times).
[0129] Node 3: The weight of the edge from node 1 to node 3 is 1 (because "RF structure component" is hit once).
[0130] Node 33: The weight of the edge from node 30 to node 33 is 4 (because "fiber laser device" is hit 4 times).
[0131] Edge weights play a key role in graph convolution computations. Edge weights reflect a company's activity or focus on the industry corresponding to the node. A higher weight indicates that the industry is more important to the company. By incorporating this weight information into the graph convolutional network model, we can more accurately capture the relationships and characteristics within the industry chain, thereby generating more precise industry distribution information.
[0132] Based on the same inventive concept, the present disclosure also provides a text information processing device, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0133] Figure 10 A schematic diagram of a text information processing device according to an embodiment of the present disclosure is shown. Figure 10 As shown, the text information processing device may include:
[0134] The receiving module 1001 is configured as an acquisition module, configured to acquire text information, the text information including industry data and industry-wide attribute data of the object, wherein the industry data includes multiple industries, a description of each industry, and the industrial chain to which each industry belongs, and the industry-wide attribute data includes multiple industries of the object;
[0135] A first determination module 1002 is configured to determine a plurality of target industries from the industry data and the industry scope attribute data, and determine the frequency of occurrence of each target industry in the industry scope attribute data, wherein each target industry is an industry included in both the industry data and the industry scope attribute data;
[0136] The second determination module 1003 is configured to determine the content description of each target industry and the industrial chain to which each target industry belongs from the industrial data, and extract the embedded features of the content description of each target industry;
[0137] A generating module 1004 is configured to generate a relationship graph of the industrial chain to which each target industry belongs using the industrial data, and extract graph relationship features of the relationship graph of the industrial chain to which each target industry belongs;
[0138] The third determination module 1005 is configured to determine the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs based on the embedding characteristics and frequencies corresponding to each target industry and the graph relationship characteristics corresponding to the industrial chain to which each target industry belongs;
[0139] The evaluation module 1006 is configured to determine the similarity of the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs, and determine the industrial distribution information of the object based on the determined similarity.
[0140] According to the technical solution provided by the embodiment of the present disclosure, text information is obtained, and the text information includes industry data and industry scope attribute data of the object, wherein the industry data includes multiple industries, content descriptions of each industry, and the industrial chain to which each industry belongs, and the industry scope attribute data includes multiple industries of the object; multiple target industries are determined from the industry data and the industry scope attribute data, and the frequency of occurrence of each target industry in the industry scope attribute data is determined, wherein each target industry is an industry included in both the industry data and the industry scope attribute data; the content description of each target industry and the industrial chain to which each target industry belongs are determined from the industry data, and the embedded features of the content description of each target industry are extracted; the industry data is used to generate a relationship graph of the industrial chain to which each target industry belongs, and the graph relationship features of the relationship graph of the industrial chain to which each target industry belongs are extracted; based on the embedded features and frequencies corresponding to each target industry and the graph relationship features corresponding to the industrial chain to which each target industry belongs, the industrial chain features corresponding to the industrial chain to which each target industry belongs are determined; the similarity of the industrial chain features corresponding to the industrial chain to which each target industry belongs is determined, and the industrial distribution information of the object is determined based on the determined similarity. Solve the problem in related technologies that conventional vector representations cannot fully represent industrial chain information and ignore the technical correlation between different industrial chains, resulting in the inability to accurately measure industrial distribution.
[0141] In some embodiments, the third determination module 1005 is further configured to group target industries belonging to the same industrial chain together; perform graph convolution calculations on the embedded features and frequencies corresponding to the target industries belonging to the same industrial chain and the graph relationship features corresponding to the industrial chain to obtain the industrial chain features corresponding to the industrial chain.
[0142] In some embodiments, the third determination module 1005 is further configured to employ another machine learning model, such as a multi-layer perceptron (MLP), to perform feature fusion. Specifically, after determining target industries belonging to the same industrial chain, their embedded features, frequency information, and graph relationship features are directly input into the MLP, and the trained MLP model outputs the industrial chain features.
[0143] In some embodiments, the third determination module 1005 is further configured to divide each target industry into multiple groups, wherein each group corresponds to an industrial chain, and each group contains at least one target industry; graph convolution calculation is performed on the characteristics and frequency of the target industry in each group and the graph relationship characteristics corresponding to the industrial chain corresponding to the group to obtain the graph relationship characteristics corresponding to the industrial chain corresponding to the group.
[0144] In some embodiments, the third determination module 1005 is further configured to divide each target industry into multiple groups, wherein each group corresponds to an industrial chain, and each group contains at least one target industry; graph convolution calculation is performed on the characteristics and frequency of the target industry in each group to obtain the graph relationship characteristics corresponding to the industrial chain corresponding to each group.
[0145] In some embodiments, the generation module 1004 is further configured to determine all industries under the industrial chain to which each target industry belongs and the level of each industry from the industrial data; for the industrial chain to which each target industry belongs: generate nodes and node names corresponding to each industry under the industrial chain; generate identification numbers for the nodes corresponding to each industry based on the level of each industry; based on the level of each industry, connect the nodes corresponding to each industry to obtain a relationship diagram of the industrial chain.
[0146] In some embodiments, the generation module 1004 is further configured to use a similarity-based method to construct an industrial chain relationship diagram. Specifically, the similarity between industries is first calculated, and the connection method between nodes is determined based on the similarity rather than the level. For example, if there is a high degree of technical or market correlation between two industries, they will be closely connected in the relationship diagram, even if they are at different levels in the traditional sense. This method can help reveal potential, non-traditional connections in the industrial chain and provide companies with a new perspective to understand the dynamic changes in their industrial chain. This method can also effectively construct an industrial chain relationship diagram, but provides a different perspective to understand and analyze the industrial chain structure.
[0147] In some embodiments, the generation module 1004 is further configured to connect the nodes corresponding to each industry based on the level of each industry, and obtain the edges between the nodes corresponding to each industry and the nodes corresponding to the industries at the previous level; based on the frequency corresponding to each industry, determine the weights of the edges between the nodes corresponding to each industry and the nodes corresponding to the industries at the previous level, and obtain a relationship diagram of the industrial chain.
[0148] In some embodiments, the evaluation module 1006 is further configured to calculate the similarity between the industrial chain characteristics corresponding to any two industrial chains in the industrial chain to which each target industry belongs, so as to obtain a similarity matrix corresponding to the industrial chain to which each target industry belongs; take an upper triangular matrix without a diagonal in the similarity matrix, and use the similarity contained in the upper triangular matrix as the determined similarity.
[0149] In some embodiments, the evaluation module 1006 is further configured to use a cluster analysis-based method to perform similarity evaluation. Specifically, first, cluster analysis is performed on all industry chain features, and similar industry chains are grouped according to their feature vectors; and the average similarity is calculated within each group.
[0150] In some embodiments, the evaluation module 1006 is further configured to calculate an average value of the determined similarities; and determine the industry distribution information based on the average value.
[0151] In some embodiments, the evaluation module 1006 is further configured to assign a corresponding weight to each similarity value according to the importance of different industrial chains to the enterprise (for example, according to the revenue ratio or investment scale), and then calculate the weighted average similarity.
[0152] In one embodiment of the present disclosure, the content description of each industry in the industry data is generated using a large language model.
[0153] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."
[0154] Refer to the following Figure 11 1100 according to this embodiment of the present disclosure will be described. Figure 11 The electronic device 1100 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0155] like Figure 11 As shown, electronic device 1100 is implemented as a general-purpose computing device. Components of electronic device 1100 may include, but are not limited to, the aforementioned at least one processing unit 1110, the aforementioned at least one storage unit 1120, and a bus 1130 connecting various system components (including storage unit 1120 and processing unit 1110).
[0156] The storage unit stores program codes, which can be executed by the processing unit 1110, so that the processing unit 1110 executes the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 1110 can execute the following steps of the above method embodiment: obtaining text information, the text information includes industry data and industry scope attribute data of the object, wherein the industry data includes multiple industries, content descriptions of each industry and the industrial chain to which each industry belongs, and the industry scope attribute data includes multiple industries of the object; determining multiple target industries from the industry data and the industry scope attribute data, and determining the frequency of occurrence of each target industry in the industry scope attribute data, wherein each target industry is an industry included in both the industry data and the industry scope attribute data; determining the content description of each target industry and the industrial chain to which each target industry belongs from the industry data, and extracting the embedded features of the content description of each target industry; using the industry data to generate a relationship graph of the industrial chain to which each target industry belongs, and extracting the graph relationship features of the relationship graph of the industrial chain to which each target industry belongs; based on the embedded features and frequencies corresponding to each target industry and the graph relationship features corresponding to the industrial chain to which each target industry belongs, determining the industrial chain features corresponding to the industrial chain to which each target industry belongs; determining the similarity of the industrial chain features corresponding to the industrial chain to which each target industry belongs, and determining the industrial distribution information of the object based on the determined similarity.
[0157] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 11201 and / or a cache memory unit 11202 , and may further include a read-only memory unit (ROM) 11203 .
[0158] The storage unit 1120 may also include a program / utility 11204 having a set (at least one) of program modules 11205, such program modules 11205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0159] The bus 1130 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0160] Electronic device 1100 may also communicate with one or more external devices 1140 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more text messaging devices that enable a user to interact with electronic device 1100, and / or any device that enables electronic device 1100 to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication may occur via input / output (I / O) interface 1150. Furthermore, electronic device 1100 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 1160. As shown, network adapter 1160 communicates with other modules of electronic device 1100 via bus 1130. It should be understood that, although not shown, other hardware and / or software modules may be used in conjunction with electronic device 1100, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0161] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0162] In the disclosed exemplary embodiments, a computer-readable storage medium is also provided. The computer-readable storage medium may be a readable signal medium or a readable storage medium.
[0163] In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary implementations of the present disclosure described in the above "Specific Implementation Methods" section of this specification.
[0164] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0165] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0166] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0167] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0168] The present disclosure provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text information processing method provided in any of the optional embodiments of the present disclosure.
[0169] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0170] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0171] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0172] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope of the present disclosure being indicated by the appended claims.
Claims
1. A text information processing method, characterized in that: include: Obtaining text information, the text information including industry data and industry-wide attribute data of the object, wherein the industry data includes multiple industries, a description of each industry, and the industrial chain to which each industry belongs, and the industry-wide attribute data includes multiple industries of the object; Determining a plurality of target industries from the industry data and the industry-scope attribute data, and determining a frequency of occurrence of each target industry in the industry-scope attribute data, wherein each target industry is an industry included in both the industry data and the industry-scope attribute data; Determining the content description of each target industry and the industrial chain to which each target industry belongs from the industrial data, and extracting the embedded features of the content description of each target industry; generating a relationship graph of the industrial chain to which each target industry belongs by using the industrial data, and extracting graph relationship features of the relationship graph of the industrial chain to which each target industry belongs; Based on the embedding characteristics and frequencies corresponding to each target industry and the graph relationship characteristics corresponding to the industrial chain to which each target industry belongs, determine the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs; The similarity of the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs is determined, and the industrial distribution information of the object is determined based on the determined similarity.
2. The method according to claim 1, characterized in that Determining the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs based on the embedding characteristics and frequencies corresponding to each target industry and the graph relationship characteristics corresponding to the industrial chain to which each target industry belongs includes: Group together target industries belonging to the same industrial chain; Graph convolution calculation is performed on the embedding features and frequencies corresponding to the target industries belonging to the same industrial chain and the graph relationship features corresponding to the industrial chain to obtain the industrial chain features corresponding to the industrial chain.
3. The method according to claim 1, characterized in that The generating of a relationship diagram of the industrial chain to which each target industry belongs by utilizing the industrial data includes: Determine all industries under the industrial chain to which each target industry belongs and the level of each industry from the industrial data; For the industrial chain to which each target industry belongs: Generate nodes and node names corresponding to each industry under the industrial chain; Generate identification numbers of nodes corresponding to each industry based on the level of each industry; Based on the level of each industry, the corresponding nodes of each industry are connected to obtain the relationship diagram of the industrial chain.
4. The method according to claim 3, characterized in that Based on the level of each industry, the nodes corresponding to each industry are connected to obtain the relationship diagram of the industrial chain, including: Based on the level of each industry, connect the nodes corresponding to each industry to obtain the edges between the nodes corresponding to each industry and the nodes corresponding to the industry at the next level above; Based on the frequency corresponding to each industry, the weight of the edge between the node corresponding to each industry and the node corresponding to the industry at the next higher level is determined to obtain the relationship diagram of the industrial chain.
5. The method according to claim 1, wherein Determining the similarity of the industrial chain characteristics corresponding to the industrial chains to which each target industry belongs includes: Calculate the similarity between the industrial chain features corresponding to any two industrial chains in the industrial chain to which each target industry belongs, so as to obtain the similarity matrix corresponding to the industrial chain to which each target industry belongs; An upper triangular matrix without a diagonal line is taken from the similarity matrix, and the similarity included in the upper triangular matrix is used as the determined similarity.
6. The method according to claim 5, characterized in that The determining of the industry distribution information of the object based on the determined similarity includes: Calculating an average of the determined similarities; The industry distribution information is determined based on the average value.
7. A text information processing device, characterized in that: include: an acquisition module configured to acquire text information, the text information including industry data and industry-wide attribute data of an object, wherein the industry data includes multiple industries, a description of each industry, and the industrial chain to which each industry belongs, and the industry-wide attribute data includes multiple industries of the object; a first determination module configured to determine a plurality of target industries from the industry data and the industry-scope attribute data, and determine a frequency of occurrence of each target industry in the industry-scope attribute data, wherein each target industry is an industry included in both the industry data and the industry-scope attribute data; A second determination module is configured to determine the content description of each target industry and the industrial chain to which each target industry belongs from the industrial data, and extract the embedded features of the content description of each target industry; a generation module configured to generate a relationship graph of the industrial chain to which each target industry belongs using the industrial data, and extract graph relationship features of the relationship graph of the industrial chain to which each target industry belongs; The third determination module is configured to determine the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs based on the embedding characteristics and frequencies corresponding to each target industry and the graph relationship characteristics corresponding to the industrial chain to which each target industry belongs; The evaluation module is configured to determine the similarity of the industrial chain characteristics corresponding to the industrial chain to which each target industry belongs, and determine the industrial distribution information of the object based on the determined similarity.
8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Enterprise division method and device
CN119066492A
Method for constructing chemical-plastic industry chain knowledge graph by using graph convolutional network
CN119250172A
Industrial chain risk assessment method and system based on co-occurrence emotion and graph attention
CN119398486A
Method and apparatus for determining industrial chain node of enterprise, and terminal and storage medium
WO2023093116A1