A method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology
By using a multi-task text recognition model to process multi-source text data of the enterprise supply chain in NLP technology, the accuracy problem when identifying the upstream and downstream relationships of the enterprise supply chain in the prior art is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510293901.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-13
AI Technical Summary
When the prior art uses NLP technology to automatically identify the upstream and downstream relationships of the enterprise supply chain, it is impossible to accurately process the abbreviation and alias of the enterprise and the product, and the rules-based methods are difficult to process complex sentence patterns, resulting in a reduction in recognition accuracy.
By obtaining the multi-source text data of the target supplier for pre-processing, combining the alias indicator characteristics, syntactic relationship characteristics and event indicator characteristics in the multi-task text recognition model, a context embedding vector is generated, and similarity calculation and threshold judgment are performed with the matching entity text data, entity alias text data is obtained, and entity alias text data is finally used to process syntactic relationship characteristics and event indicator characteristics, and entity relationship information of the target supplier is obtained.
It improves the accuracy of enterprise supply relationship identification, can more accurately identify alias and process complex sentence patterns, thereby improving the recognition effect of upstream and downstream relationships in the supply chain.
Smart Images

Figure CN119808787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of enterprise supply relationship identification, and specifically to a method for automatically identifying upstream and downstream relationships in an enterprise supply chain by using NLP technology. Background Art
[0002] A key part of supply chain analysis is to obtain the supply relationship between enterprises. With the development of natural language processing technology, NLP technology can be used to identify the names of companies and products from text data, and obtain the supply relationship between enterprises and the corresponding relationship between enterprises and products.
[0003] At present, the existing technology still has shortcomings in using NLP technology to automatically identify the upstream and downstream relationships in the enterprise supply chain. On the one hand, the existing technology does not take into account the existence of abbreviations and aliases for enterprises and products, which will lead to the inability to accurately identify the names of enterprises and products; on the other hand, the existing technology uses rule-based methods and cannot handle complex sentences, resulting in misjudgment of the initiator and receiver of the action in the sentence, thereby reducing the accuracy of identifying the upstream and downstream relationships in the enterprise supply chain.
[0004] Therefore, a method using NLP technology to automatically identify the upstream and downstream relationships in an enterprise supply chain is proposed. Summary of the invention
[0005] The purpose of the present invention is to provide a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology. The method first obtains multi-source text data of a target supplier and pre-processes it to obtain pre-processed multi-source text data and matching entity text data; the pre-processed multi-source text data and matching entity text data are input into a multi-task text recognition model to obtain alias indicator features, syntactic relationship features and event indicator features; the alias indicator features and syntactic relationship features are combined to obtain a context embedding vector and perform similarity calculation and threshold judgment with the matching entity text data to obtain entity alias text data; the syntactic relationship features, event indicator features, pre-processed multi-source text data and entity text data groups are processed using a multi-task text recognition model to obtain entity relationship information of the target supplier. The present invention can effectively improve the accuracy of enterprise supply relationship identification.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology, comprising:
[0008] Obtain multi-source text data of target suppliers;
[0009] Preprocessing the multi-source text data to obtain preprocessed multi-source text data and matching entity text data;
[0010] Inputting the preprocessed multi-source text data and the matching entity text data into a multi-task text recognition model to obtain alias indicator features, syntactic relationship features, and event indicator features; obtaining a context embedding vector based on the alias indicator features, the syntactic relationship features, and the multi-task text recognition model;
[0011] Calculate similarity and threshold value of the context embedding vector and the matching entity text data to obtain entity alias text data; update the enterprise knowledge base using the entity alias text data to obtain an enterprise knowledge update base;
[0012] The multi-task text recognition model is used to process the syntactic relationship features, the event indicator features, the pre-processed multi-source text data and the entity text data group to obtain entity relationship information of the target supplier.
[0013] Furthermore, the multi-source text data of the target supplier includes: company website text data, bidding documents, contracts and industry report data; the matching entity text data includes: matching enterprise text data and matching product text data; the entity text data group includes: the matching entity text data and the entity alias text data.
[0014] Furthermore, the process of preprocessing the multi-source text data to obtain preprocessed multi-source text data and matching entity text data includes:
[0015] Performing data integration, data cleaning and text denoising on the multi-source text data to obtain the pre-processed multi-source text data;
[0016] The pre-processed multi-source text data and the entity text data in the enterprise knowledge base are identified and matched using a fuzzy matching algorithm to obtain the matching entity text data.
[0017] Furthermore, the multi-task text recognition model includes: an input layer, an alias indicator recognition layer, a syntactic relationship recognition layer, a context vector extraction layer, an event indicator recognition layer, an entity relationship recognition layer, an entity relationship adjustment layer and an output layer.
[0018] Further, the preprocessed multi-source text data and the matching entity text data are input into a multi-task text recognition model to obtain alias indicator features, syntactic relationship features and event indicator features; the specific process of obtaining a context embedding vector according to the alias indicator features, the syntactic relationship features and the multi-task text recognition model includes:
[0019] Inputting the preprocessed multi-source text data and the matching entity text data into an input layer of the multi-task text recognition model to obtain multi-source text features and matching entity text features;
[0020] The multi-source text features are respectively input into the alias indicator recognition layer, the syntactic relationship recognition layer and the event indicator recognition layer in the multi-task text recognition model to obtain the alias indicator features, the syntactic relationship features and the event indicator features; wherein the matching entity text features participate in the attention calculation process in the syntactic relationship recognition layer;
[0021] The alias indicator feature and the syntactic relationship feature are sequentially input into the context vector extraction layer and the output layer in the multi-task text recognition model to obtain the context embedding vector.
[0022] Furthermore, the process of performing similarity calculation and threshold judgment on the context embedding vector and the matching entity text data to obtain entity alias text data includes:
[0023] Performing text vectorization processing on the matching entity text data to obtain a matching entity text vector;
[0024] Calculate the similarity between the context embedding vector and the matching entity text vector to obtain text vector similarity;
[0025] The text vector similarity is compared with a preset similarity threshold, and if the text vector similarity exceeds the preset similarity threshold, the entity alias text data corresponding to the context embedding vector is obtained.
[0026] Furthermore, the calculation formula of the text vector similarity is:
[0027] ;
[0028] in, is the text vector similarity; is the dimension of the context embedding vector and the matching entity text vector; is the learnable weight value of the i-th dimension; is the context embedding vector value of the i-th dimension; is the matching entity text vector value of the i-th dimension; and are the L2 norms of the context embedding vector and the matching entity text vector respectively.
[0029] Furthermore, the process of using the multi-task text recognition model to process the syntactic relationship features, the event indicator features, the pre-processed multi-source text data and the entity text data group to obtain the entity relationship information of the target supplier includes:
[0030] Inputting the preprocessed multi-source text data and the entity text data group into the input layer of the multi-task text recognition model to obtain multi-source text features and entity text features;
[0031] Inputting the syntactic relationship feature, the multi-source text feature and the entity text feature into the entity relationship recognition layer in the multi-task text recognition model to obtain entity relationship features;
[0032] The entity relationship features and the event indicator word features are sequentially input into the entity relationship adjustment layer and the output layer in the multi-task text recognition model to obtain the entity relationship information of the target supplier.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1. By combining multi-source text data with a multi-task text recognition model, and inputting the output features of the alias indicator recognition layer and the syntactic relationship recognition layer in the multi-task text recognition model into the context vector extraction layer, the context embedding vector is obtained; the combination of alias indicator features and syntactic relationship features can provide more accurate context information, thereby helping the model to better locate the alias position; this process can provide data support for subsequent similarity calculations to improve the accurate recognition of aliases, thereby improving the accuracy of enterprise supply relationship identification.
[0035] 2. The text vector similarity is obtained by calculating the weighted similarity between the context embedding vector output by the multi-task text recognition model and the known matching entity text vector; the accurate recognition of aliases is further improved based on the threshold judgment result; the use of learnable weights of different dimensions in the calculation process of text vector similarity can pay more attention to the dimensions that are conducive to distinguishing aliases; this process can effectively improve the accuracy of enterprise supply relationship identification.
[0036] 3. By inputting syntactic relationship features, multi-source text features and entity text features into the entity relationship recognition layer in the multi-task text recognition model, entity relationship features are obtained; combining rich contextual information with syntactic structure information can quickly obtain the relationship between entities; then, the entity relationship features are adjusted using event indicator features and the entity relationship adjustment layer to accurately correct the supply relationship between entities, which can effectively improve the accuracy of enterprise supply relationship identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1A flow chart of a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology according to the present invention;
[0038] Figure 2 Schematic diagram of the structure of the multi-task text recognition model of the present invention. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] The present invention proposes a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology. The method proposed by the present invention will be specifically described below in conjunction with Example 1 and Example 2, as follows:
[0041] Embodiment 1
[0042] In order to accurately obtain the upstream and downstream relationships of the supply chain, a supplier enterprise uses a method proposed in this invention that uses NLP technology to automatically identify the upstream and downstream relationships of the enterprise supply chain. The process of this method is shown as follows: Figure 1 As shown, including:
[0043] Obtain multi-source text data of target suppliers;
[0044] Furthermore, the target supplier's multi-source text data includes: company website text data, bidding documents, contracts, and industry report data;
[0045] By collecting multi-source text data, a solid data foundation can be provided for subsequent text recognition. Rich sentences related to upstream and downstream relationships in the supply chain can be extracted from multi-source text data, thereby improving the accuracy of enterprise supply relationship identification.
[0046] Preprocessing the multi-source text data to obtain preprocessed multi-source text data and matching entity text data;
[0047] Furthermore, the process of preprocessing the multi-source text data to obtain the preprocessed multi-source text data and the matching entity text data includes:
[0048] Perform data integration, data cleaning and text denoising on multi-source text data to obtain pre-processed multi-source text data;
[0049] Use fuzzy matching algorithm to identify and match preprocessed multi-source text data with entity text data in the enterprise knowledge base to obtain matching entity text data;
[0050] Furthermore, data integration includes: data format unification, data structure standardization and data merging; data cleaning includes: missing processing, duplicate processing and exception processing; text denoising includes: label processing, character processing, URL and Email address removal, spelling check, etc.;
[0051] Furthermore, matching entity text data includes: matching enterprise text data and matching product text data.
[0052] The quality of text data can be improved by preprocessing multi-source text data. At the same time, the matching entity text data obtained by fuzzy matching can be used as guiding data for subsequent text recognition, which can help the model quickly locate the scope of entities and aliases, thereby improving the accuracy of enterprise supply relationship identification.
[0053] Input the preprocessed multi-source text data and the matching entity text data into the multi-task text recognition model to obtain alias indicator features, syntactic relationship features and event indicator features; obtain a context embedding vector based on the alias indicator features, syntactic relationship features and the multi-task text recognition model;
[0054] Furthermore, the structure of the multi-task text recognition model can be referred to Figure 2 , including: input layer, alias indicator recognition layer, syntactic relationship recognition layer, context vector extraction layer, event indicator recognition layer, entity relationship recognition layer, entity relationship adjustment layer and output layer;
[0055] In this embodiment, the multi-task text recognition model mainly realizes alias indicator recognition, syntactic relationship recognition, event indicator recognition and entity relationship recognition; the use of alias indicator features and syntactic relationship features can help accurately determine the alias scope, and the use of syntactic relationship features and event indicator features can respectively identify and adjust entity relationships. The model uses the above recognition functions to effectively improve the accuracy of enterprise supply relationship recognition.
[0056] Furthermore, the preprocessed multi-source text data and the matching entity text data are input into the multi-task text recognition model to obtain alias indicator features, syntactic relationship features and event indicator features; the specific process of obtaining the context embedding vector according to the alias indicator features, syntactic relationship features and the multi-task text recognition model includes:
[0057] Input the preprocessed multi-source text data and matching entity text data into the input layer of the multi-task text recognition model to obtain multi-source text features and matching entity text features;
[0058] The multi-source text features are respectively input into the alias demonstrative word recognition layer, the syntactic relation recognition layer and the event demonstrative word recognition layer in the multi-task text recognition model to obtain the alias demonstrative word features, the syntactic relation features and the event demonstrative word features; among which, the matching entity text features participate in the attention calculation process in the syntactic relation recognition layer;
[0059] The alias indicator features and syntactic relationship features are sequentially input into the context vector extraction layer and output layer in the multi-task text recognition model to obtain the context embedding vector.
[0060] Furthermore, the syntactic relation recognition layer is a gated attention recognition network, which obtains contextual features related to syntactic relations by using a gated weighted attention module, and then the attention recognition network uses the contextual features to identify syntactic relations related to the matching entity text;
[0061] Further, the syntactic relationship objects associated with the matching entity text include both aliases of the matching entity text and supply relationship objects;
[0062] Furthermore, the attention weight calculation formula in the gated weighted attention module is:
[0063] ;
[0064] in, represents the attention weight; represents the normalized exponential function; A query matrix representing multi-source text features; Represents a matrix product operation; represents the gated weighted key matrix after transposition; Indicates square root calculation; represents the dimension of the key matrix;
[0065] Furthermore, the calculation process of the gated weighted key matrix can be expressed as:
[0066] ;
[0067] ;
[0068] in, represents the gated weighted bond matrix; Represents the feature matrix concatenation operation; A key matrix representing multi-source text features; represents the gate value; A key matrix representing the text features of the matching entity; Represents the sigmoid activation function; represents the gating weight matrix; represents the bias vector;
[0069] Furthermore, the network structures of the alias indicator recognition layer and the event indicator recognition layer are consistent, both of which are self-attention networks, that is, the residual network after removing the gating and weighting mechanism in the syntactic relationship recognition layer.
[0070] The context embedding vector is obtained by inputting the output features of the alias indicator recognition layer and the syntactic relationship recognition layer into the context vector extraction layer; the combination of alias indicator features and syntactic relationship features can provide more accurate context information, thereby helping the model to better locate the alias position; this process can provide data support for subsequent similarity calculations to improve the accurate recognition of aliases, thereby improving the accuracy of enterprise supply relationship identification.
[0071] The context embedding vector and the matching entity text data are subjected to similarity calculation and threshold judgment to obtain entity alias text data; the enterprise knowledge base is updated using the entity alias text data to obtain an enterprise knowledge update base;
[0072] Furthermore, the process of calculating the similarity and threshold value of the context embedding vector and the matching entity text data to obtain the entity alias text data includes:
[0073] Perform text vectorization processing on the matching entity text data to obtain the matching entity text vector;
[0074] Calculate the similarity between the context embedding vector and the matching entity text vector to obtain the text vector similarity;
[0075] Compare the text vector similarity with a preset similarity threshold, and if the text vector similarity exceeds the preset similarity threshold, obtain the entity alias text data corresponding to the context embedding vector;
[0076] Furthermore, the text vectorization processing process includes: first preprocessing the matching entity text data to obtain preprocessed matching entity text data; then inputting the preprocessed matching entity text data into the pre-trained BERT model for extraction to obtain a context vector; adjusting the context vector to keep it consistent with the dimension size of the context embedding vector to obtain a matching entity text vector;
[0077] Furthermore, the preset similarity threshold is set to 0.85.
[0078] The text vector similarity is obtained by calculating the weighted similarity between the context embedding vector output by the multi-task text recognition model and the known matching entity text vector; the accurate recognition of aliases is further improved based on the threshold judgment result.
[0079] Furthermore, the calculation formula for text vector similarity is:
[0080] ;
[0081] in, is the text vector similarity; is the dimension of the context embedding vector and the matching entity text vector; is the learnable weight value of the i-th dimension; is the context embedding vector value of the i-th dimension; is the matching entity text vector value of the i-th dimension; and are the L2 norms of the context embedding vector and the matching entity text vector, respectively.
[0082] In order to illustrate the text vector similarity calculation proposed in the present invention, three groups of different text data are selected for text vector similarity test, which are recorded as test one, test two and test three; the corresponding matching entity text data are obtained through preprocessing operations; the corresponding context embedding vector is obtained using a multi-task text recognition model; the text vector similarity test results of each group are obtained by performing text vector similarity calculation and threshold judgment on the context embedding vector of each group and the matching entity text vector, which can be referred to in Table 1.
[0083] Table 1. Text vector similarity test results
[0084] Test samples Text vector similarity Whether to obtain the corresponding entity alias text data Test 1 0.91 yes Test 2 0.63 no Test Three 0.89 yes
[0085] In the process of calculating the text vector similarity in this embodiment, the use of learnable weights of different dimensions can pay more attention to the dimensions that are conducive to distinguishing names; this process can effectively improve the accuracy of enterprise supply relationship identification.
[0086] The multi-task text recognition model is used to process the syntactic relationship features, event indicator features, pre-processed multi-source text data and entity text data groups to obtain the entity relationship information of the target supplier.
[0087] Further, the entity text data group includes: matching entity text data and entity alias text data;
[0088] Furthermore, the process of using the multi-task text recognition model to process the syntactic relationship features, event indicator features, pre-processed multi-source text data and entity text data group to obtain entity relationship information of the target supplier includes:
[0089] Input the preprocessed multi-source text data and entity text data group into the input layer of the multi-task text recognition model to obtain multi-source text features and entity text features;
[0090] Inputting syntactic relationship features, multi-source text features and entity text features into the entity relationship recognition layer in the multi-task text recognition model to obtain entity relationship features;
[0091] The entity relationship features and event indicator word features are sequentially input into the entity relationship adjustment layer and the output layer in the multi-task text recognition model to obtain the entity relationship information of the target supplier;
[0092] Furthermore, the network structure of the entity relationship recognition layer is basically the same as that of the syntactic relationship recognition layer. The difference is that in the attention weight calculation, not only the entity text features are involved, but also the syntactic relationship features are added to improve the accuracy of entity relationship recognition;
[0093] Furthermore, the entity relationship adjustment layer is a fusion network based on a multi-layer perceptron for nonlinearly transforming and fusing entity relationship features and event indicator word features.
[0094] By inputting syntactic relationship features, multi-source text features and entity text features into the entity relationship recognition layer in the multi-task text recognition model, entity relationship features are obtained; combining rich contextual information with syntactic structure information can quickly obtain the relationship between entities; then, the entity relationship features are adjusted using event indicator features and the entity relationship adjustment layer to accurately correct the supply relationship between entities, which can effectively improve the accuracy of enterprise supply relationship identification.
[0095] This embodiment proposes a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology; the method first obtains multi-source text data of a target supplier and pre-processes it to obtain pre-processed multi-source text data and matching entity text data; the pre-processed multi-source text data and matching entity text data are input into a multi-task text recognition model to obtain alias indicator features, syntactic relationship features and event indicator features; the alias indicator features and syntactic relationship features are combined to obtain a context embedding vector and perform similarity calculation and threshold judgment with the matching entity text data to obtain entity alias text data; the syntactic relationship features, event indicator features, pre-processed multi-source text data and entity text data groups are processed using a multi-task text recognition model to obtain entity relationship information of the target supplier. The present invention can effectively improve the accuracy of enterprise supply relationship identification.
[0096] Embodiment 2
[0097] The present invention proposes a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology. In order to further verify the effectiveness of the method proposed in the present invention in entity alias identification and supply relationship identification, the present invention conducts alias identification accuracy tests and supply relationship identification accuracy tests for different schemes. The present invention selects two supplier companies A and B to conduct the above two groups of accuracy tests.
[0098] The present invention collects multi-source text data related to supplier enterprise A as test data, which are recorded as data one, data two and data three; each data is different; the multi-source text data of each enterprise is preprocessed to obtain preprocessed multi-source text data and matching entity text data of each group; different schemes are used to process the preprocessed multi-source text data and matching entity text data of each group, and the context embedding vectors output by different schemes are used to perform similarity calculation and threshold judgment to obtain alias recognition data of each group, and then through manual verification, the alias recognition accuracy test results of each group are obtained.
[0099] The comparison schemes adopted in this embodiment are respectively recorded as Scheme 1, Scheme 2 and Scheme 3. Among them, Scheme 1 is proposed in the present invention to combine the alias indicator features and syntactic relationship features output by the alias indicator recognition layer and the syntactic relationship recognition layer; Scheme 2 is to remove the gating and weighting mechanism of the syntactic relationship recognition layer of the present invention, that is, the matching entity text features do not participate in the attention calculation process in the syntactic relationship recognition layer; Scheme 3 is not to use syntactic relationship features, but directly use alias indicator features for recognition.
[0100] The test results of alias recognition accuracy are shown in Table 2.
[0101] Table 2 Alias recognition accuracy test results
[0102] Test Data plan Alias recognition accuracy Data 1 Solution 1 91.58% Data 2 Solution 2 89.27% Data 3 Option 3 85.44%
[0103] From the results in Table 2, it can be seen that the alias identification scheme proposed in the present invention, namely scheme 1, has better test results in terms of alias identification accuracy than the test results of other schemes; thus, it can be shown that the method proposed in the present invention can accurately identify entity alias data and further improve the accuracy of enterprise supply relationship identification;
[0104] In order to further test the accuracy of supply relationship identification, this embodiment selects multi-source text data related to supplier enterprise B as test data, recorded as sample one, sample two and sample three; the test data are all different; preprocessing, alias indicator recognition, syntactic relationship recognition and event indicator recognition are performed on each group of test data to obtain alias indicator features, syntactic relationship features and event indicator features; combining text vector similarity calculation and threshold judgment process, the entity text data group of each group is obtained; according to three different supply relationship identification test schemes, the entity supply relationship under each group scheme is obtained, and the supply relationship identification test results are obtained by manual investigation; test scheme one is the combination of entity relationship identification and entity relationship adjustment strategy proposed by the present invention; test scheme two is not to add entity text data group in the entity relationship identification process; test scheme three is to remove the entity relationship adjustment process; the supply relationship identification accuracy test results are shown in Table 3;
[0105] Table 3 Test results of supply relationship identification accuracy
[0106] Test Data plan Supply relationship identification accuracy Sample 1 Test plan 1 90.93% Sample 2 Test plan 2 88.76% Sample 3 Test plan three 86.05%
[0107] It can be seen from the results in Table 3 that the accuracy test results obtained by using test scheme 1, that is, the supply relationship identification method proposed in the present invention, are better than those of other schemes. This shows that it is necessary to combine entity relationship identification with entity relationship adjustment, which can improve the accuracy of enterprise supply relationship identification.
[0108] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology, characterized in that: include: Obtain multi-source text data of target suppliers; Preprocessing the multi-source text data to obtain preprocessed multi-source text data and matching entity text data; Inputting the preprocessed multi-source text data and the matching entity text data into a multi-task text recognition model to obtain alias indicator features, syntactic relationship features, and event indicator features; obtaining a context embedding vector based on the alias indicator features, the syntactic relationship features, and the multi-task text recognition model; Calculate similarity and threshold value of the context embedding vector and the matching entity text data to obtain entity alias text data; update the enterprise knowledge base using the entity alias text data to obtain an enterprise knowledge update base; Using the multi-task text recognition model, the syntactic relationship features, the event indicator features, the pre-processed multi-source text data and the entity text data group are processed to obtain entity relationship information of the target supplier; The multi-source text data of the target supplier includes: company website text data, bidding documents, contracts and industry report data; the matching entity text data includes: matching enterprise text data and matching product text data; the entity text data group includes: the matching entity text data and the entity alias text data; The process of preprocessing the multi-source text data to obtain the preprocessed multi-source text data and the matching entity text data includes: Performing data integration, data cleaning and text denoising on the multi-source text data to obtain the pre-processed multi-source text data; The pre-processed multi-source text data and the entity text data in the enterprise knowledge base are identified and matched using a fuzzy matching algorithm to obtain the matching entity text data.
2. According to claim 1, a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology is characterized in that: The multi-task text recognition model includes: an input layer, an alias indicator recognition layer, a syntactic relationship recognition layer, a context vector extraction layer, an event indicator recognition layer, an entity relationship recognition layer, an entity relationship adjustment layer and an output layer.
3. According to claim 1, a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology is characterized in that: Inputting the preprocessed multi-source text data and the matching entity text data into a multi-task text recognition model to obtain alias indicator features, syntactic relationship features, and event indicator features; and obtaining a context embedding vector according to the alias indicator features, the syntactic relationship features, and the multi-task text recognition model includes: Inputting the preprocessed multi-source text data and the matching entity text data into an input layer in the multi-task text recognition model to obtain multi-source text features and matching entity text features; The multi-source text features are respectively input into the alias indicator recognition layer, the syntactic relationship recognition layer and the event indicator recognition layer in the multi-task text recognition model to obtain the alias indicator features, the syntactic relationship features and the event indicator features; wherein the matching entity text features participate in the attention calculation process in the syntactic relationship recognition layer; The alias indicator feature and the syntactic relationship feature are sequentially input into the context vector extraction layer and the output layer in the multi-task text recognition model to obtain the context embedding vector.
4. According to claim 1, a method for automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology is characterized in that: The process of performing similarity calculation and threshold judgment on the context embedding vector and the matching entity text data to obtain entity alias text data includes: Performing text vectorization processing on the matching entity text data to obtain a matching entity text vector; Calculate the similarity between the context embedding vector and the matching entity text vector to obtain text vector similarity; The text vector similarity is compared with a preset similarity threshold, and if the text vector similarity exceeds the preset similarity threshold, the entity alias text data corresponding to the context embedding vector is obtained.
5. The method of automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology according to claim 4, characterized in that: The calculation formula of the text vector similarity is: ; in, is the text vector similarity; is the dimension of the context embedding vector and the matching entity text vector; is the learnable weight value of the i-th dimension; is the context embedding vector value of the i-th dimension; is the matching entity text vector value of the i-th dimension; and are the L2 norms of the context embedding vector and the matching entity text vector respectively.
6. The method of automatically identifying upstream and downstream relationships in an enterprise supply chain using NLP technology according to claim 1, characterized in that: The process of using the multi-task text recognition model to process the syntactic relationship features, the event indicator features, the pre-processed multi-source text data and the entity text data group to obtain the entity relationship information of the target supplier includes: Inputting the preprocessed multi-source text data and the entity text data group into the input layer of the multi-task text recognition model to obtain multi-source text features and entity text features; Inputting the syntactic relationship feature, the multi-source text feature and the entity text feature into the entity relationship recognition layer in the multi-task text recognition model to obtain entity relationship features; The entity relationship features and the event indicator word features are sequentially input into the entity relationship adjustment layer and the output layer in the multi-task text recognition model to obtain the entity relationship information of the target supplier.
Citation Information
Patent Citations
Method and system for constructing business social network
CN101853292A
A NLP-based method for automatic extraction and analysis of enterprise supply relationship
CN109376202A