Supplier risk prediction method based on knowledge graph
By constructing and completing the knowledge graph and applying a pre-trained risk prediction model, the problem of insufficient supplier risk assessment in the existing technology is solved, and more accurate and efficient risk prediction is achieved.
Patent Information
- Application Number
- CN202411970129.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
There is a lack of an effective supplier risk management method in the prior art. The traditional method relies on the information and manual experience provided by the supplier and is susceptible to information beautification and subjective factors, resulting in insufficient risk assessment.
The supplier's risk prediction method based on the knowledge graph is adopted to build a knowledge graph by obtaining the supplier's current operation data and historical operation data, and use the preset knowledge graph completion method and the pre-trained risk prediction model to make risk prediction.
By considering more abundant risk-related data sources, the accuracy and efficiency of risk prediction are improved, providing a more reliable basis for supplier risk assessment.
Smart Images

Figure CN120069515A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of supplier risk management, and in particular, to a supplier risk prediction method based on a knowledge graph. Background Art
[0002] In the field of engineering construction, suppliers are indispensable for projects, and suppliers provide necessary professional services for enterprises. When enterprises select suppliers, they will consider many factors: one is the construction quality risk, such as the quality of the project entity, the implementation of construction technology, and the quality inspection situation, etc.; they will also pay attention to litigation risks, such as whether the supplier has contract breaches or labor disputes, etc. The above factors may all affect the smooth progress of the project and the reputation of the enterprise. Therefore, before the project starts, it is necessary to conduct a risk assessment of the supplier.
[0003] Traditional supplier risk assessment methods mainly rely on the information provided by the suppliers themselves. However, suppliers may beautify the information or conceal unfavorable information in order to obtain cooperation opportunities, resulting in the information obtained by enterprises being untrue and incomplete. At the same time, some assessments rely on manual experience, and judgments are made based on past work experience and understanding of the suppliers. However, this method has obvious drawbacks. It is difficult to consider everything with manual experience, and the assessment results are easily affected by subjective factors, resulting in an insufficient risk assessment of the suppliers, making it difficult for construction enterprises to carry out effective risk management work and increasing the uncertainty and potential risks during the project implementation process. In addition, some existing enterprises have tried to introduce enterprise resource planning (ERP) systems or customer relationship management (CRM) systems to manage supplier information, but these systems often focus on process control and data recording and cannot conduct in-depth analysis. In addition, there have also been attempts to use big data technology for risk monitoring, but most of them have not been deeply integrated with the specific business scenarios of enterprises and have insufficient utilization of knowledge. Therefore, an effective supplier risk management method is needed. Summary of the Invention
[0004] In view of the deficiencies in the prior art, the present invention provides a supplier risk prediction method based on a knowledge graph, which mainly solves the problem of the lack of an effective supplier risk management method in the prior art.
[0005] The object of the present invention is achieved through the following solutions:
[0006] According to an embodiment of the present invention, a method for predicting supplier risks based on a knowledge graph is provided, including the following steps: S1. Obtain the current operation data and historical operation data of the supplier, and construct a knowledge graph based on the obtained data, where the current operation data includes professional certificate data, training development data, and financial status data, and the historical operation data includes material quality inspection data, quality and safety inspection data, contract default data, industry violation record data, historical labor dispute records, and historical labor compliance data; S2. Use a preset knowledge graph completion method to complete the constructed knowledge graph to obtain an updated knowledge graph; S3. Based on the updated knowledge graph, use a pre-trained risk prediction model to perform risk prediction.
[0007] According to an embodiment of the present invention, constructing a knowledge graph based on the obtained data is as follows: S11. Extract the obtained data, and form explicit triples and implicit triples through entity recognition and relationship recognition, where the current operation data corresponds to the explicit triples, and the historical operation data corresponds to the implicit triples; S12. Align entities through a translation model to obtain a fused knowledge graph.
[0008] According to an embodiment of the present invention, extracting the obtained data and forming explicit triples and implicit triples through entity recognition and relationship recognition is as follows: S111. Represent the obtained data with word vectors, part-of-speech vectors, and word length vectors; S112. Calculate the correlation of adjacent elements in the data through a conditional random field model, learn entity representations and perform entity recognition, and further extract the relationships between entities to form entity, entity, and relationship triples; S113. Align entities through a translation model to obtain a fused knowledge graph.
[0009] According to an embodiment of the present invention, aligning entities through a translation model to obtain a fused knowledge graph is as follows: S1131. Use a translation model to calculate the distance between each pair of mapped entities; S1132. Judge the similarity of entities based on a preset distance threshold, integrate similar entities and construct new relationships to obtain a fused knowledge graph.
[0010] According to an embodiment of the present invention, the distance between two entities is:
[0011]
[0012] where E a indicates entity a, E b indicates entity b, indicates the translation relationship between two entities a and b.
[0013] According to an embodiment of the present invention, the preset knowledge graph completion method includes knowledge graph completion based on the TransE model.
[0014] According to an embodiment of the present invention, a preset knowledge graph completion method is used to complete the constructed knowledge graph to obtain an updated knowledge graph as follows: S21. Extract a batch of triples from the constructed knowledge graph as correct triples, and generate incorrect triples based on the correct triples; S22. Calculate the vector distance between the entity vector plus the relationship vector and the tail entity vector of each triple; S23. Calculate the loss based on the vector distance of any correct triple and the vector distance of any incorrect triple, and update the parameters of the TransE model; S24. Repeat steps S21-S23 until a predetermined number of iterations is reached or the loss function converges, and use the trained TransE model to complete the incomplete triples and find and delete the incorrect triples.
[0015] According to an embodiment of the present invention, the loss is:
[0016] L = ∑ (h,r,t)∈S ∑ (h′,r,t′)∈S (d(h + r, t)+d(h′ + r, t″)),
[0017] where (h, r, t) indicates a correct triple, (h′, r, t′) indicates an incorrect triple, d(h + r, t) indicates the vector distance of the correct triple, d(h′ + r, t′) indicates the vector distance of the incorrect triple, and S indicates the triple set.
[0018] According to an embodiment of the present invention, based on the updated knowledge graph, risk prediction is performed using a pre-trained risk prediction model as follows: S31. Based on the updated knowledge graph, construct an entity feature matrix and a relationship feature matrix; S32. Calculate the weights between the associated entity and all connected neighbor entities under any relationship path through the entity-level attention network in the pre-trained risk prediction model, and linearly combine the weights with the neighbor entities to obtain the entity features after feature fusion under any relationship path;
[0019] S33. Fuse the semantic information under different relationship paths, learn the relationship-level attention coefficients of each relationship path, and use them as the input of the relationship-level attention network to obtain a new feature vector of any entity after passing through the relationship-level attention network layer; S34. After normalizing the new feature vector, obtain the predicted risk level of any entity.
[0020] According to an embodiment of the present invention, the weights between the associated entity and all connected neighbor entities under any relationship path are:
[0021]
[0022] where For the relationship path r m All neighbor entities of the entity i below, is the weight vector of the entity-level attention network, e i 、e j and e n represent any entity feature, r m represents the path feature between different entities, and the relationship-level attention coefficient of each relationship path is:
[0023]
[0024] where g is the attention mechanism vector, W b is the parameterized weight matrix, and b is the bias vector, indicates the entity feature after feature fusion under the path r m
[0025] Compared with the prior art, the present invention has the following beneficial effects: By considering richer data sources that are more relevant to risks, the prediction of risks is made more accurate; in addition, through a series of data and feature processing measures, the efficiency and accuracy of the prediction are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The following further describes the embodiments of the present invention with reference to the drawings:
[0027] Figure 1 is a schematic flowchart of the method for predicting supplier risks based on a knowledge graph according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0028] Before proceeding with the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "coupled," "connected," and their derivatives refer to any direct or indirect communication or connection between two or more elements, whether or not those elements are in physical contact with each other. The terms "transmit," "receive," and "communicate," and their derivatives, encompass both direct and indirect communication. The terms "comprise" and "include," and their derivatives, mean including but not limited to. The term "or" is inclusive, meaning and / or. The phrase "associated with," and its derivatives, mean including, included within, interconnected, containing, contained within, connected or connected to, coupled or coupled to, communicating with, cooperating with, interlacing, juxtaposed, proximate, bound or bound to, having, having an attribute, having a relationship or having a relationship with, etc. The term "controller" refers to any device, system, or part thereof that controls at least one operation. Such a controller may be implemented in hardware, or in a combination of hardware and software and / or firmware. The functions associated with any particular controller may be centralized or distributed, whether local or remote. The phrase "at least one," when used in conjunction with a list of items, means that different combinations of one or more of the listed items may be used, and only one item from the list may be required. For example, "at least one of A, B, C" includes any one of the following combinations: A, B, C, A and B, A and C, B and C, A and B and C.
[0029] Definitions of other specific words and phrases are provided throughout this patent document. One of ordinary skill in the art should understand that in many instances, if not most instances, such definitions apply to the prior and future use of such defined words and phrases.
[0030] In this patent document, the application combinations of modules and the hierarchical division of sub-modules are for illustrative purposes only. Without departing from the scope of the present disclosure, the application combinations of modules and the hierarchical division of sub-modules may have different forms.
[0031] In order to make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] As described in the background art, there is a lack of an effective supplier risk management method in the prior art. To address the above problem, according to an embodiment of the present invention, a supplier risk prediction method based on a knowledge graph is provided, including the following steps: S1. Obtain the current operation data and historical operation data of the supplier, and construct a knowledge graph based on the obtained data. Among them, the current operation data includes professional certificate data, training data, and financial status data, and the historical operation data includes material quality inspection data, quality and safety inspection data, contract default data, industry violation record data, historical labor dispute records, and historical labor compliance data; S2. Use a preset knowledge graph completion method to complete the constructed knowledge graph to obtain an updated knowledge graph; S3. Based on the updated knowledge graph, use a pre-trained risk prediction model to perform risk prediction. By considering richer and more risk-related data sources, the risk prediction is made more accurate; in addition, through a series of data and feature processing measures, the efficiency and accuracy of the prediction are improved.
[0033] As is well known, multiple factors can reflect the risks of suppliers. Some obvious factors, such as current operation data including professional certificate data, training data, and financial status data (such as debt-to-asset ratio, net profit rate, etc.), can more intuitively reflect the qualifications of suppliers, and these factors are often conveniently provided by suppliers. For some implicit factors, such as material quality inspection data, quality and safety inspection data (such as process standard compliance rate), contract default data (such as the number of defaults, the proportion of default amount), industry violation record data (such as the number of violation penalties), historical labor dispute records (dispute types, the number of people involved, settlement status, etc.), and historical labor compliance data (compliance with labor laws, compliance of social insurance payment, etc.), these factors are also related to risks, but are often not provided by suppliers. However, fully considering these factors and expanding to more relevant factors will undoubtedly make the risk assessment of suppliers more accurate.
[0034] According to an embodiment of the present invention, constructing a knowledge graph based on the obtained data is as follows: S11. Extract the obtained data, and form explicit triples and implicit triples through entity recognition and relationship recognition. Among them, the current operation data corresponds to the explicit triples, and the historical operation data corresponds to the implicit triples; S12. Align the entities through a translation model to obtain a fused knowledge graph.
[0035] Specifically, in order to be applicable to structured and unstructured data and find entities and relationships in this data, according to an embodiment of the present invention, the acquired data is extracted, and explicit triples and implicit triples are formed through entity recognition and relationship recognition as follows: S111. Represent the acquired data with word vectors, part-of-speech vectors, and word length vectors; S112. Calculate the correlation between adjacent elements in the data through a conditional random field model, learn entity representations and perform entity recognition, and further extract the relationships between entities to form entity, entity, and relationship triples; S113. Align entities through a translation model to obtain a fused knowledge graph. For example, for the case of "a supplier providing false materials to seek winning the bid and being punished", convert the words in the text into vectors, and at the same time analyze the part of speech of the words (such as nouns, verbs), which helps to understand the structure and semantics of the text. In addition, represent the length vector of the words to capture the complexity of the text.
[0036] In addition, in order to ensure the simplicity of the knowledge graph and remove redundancy, according to an embodiment of the present invention, entities are aligned through a translation model, and the fused knowledge graph is obtained as follows: S1131. Use the translation model to calculate the distance between the mapped entities; S1132. Judge the similarity of entities based on a preset distance threshold, integrate similar entities and construct new relationships to obtain a fused knowledge graph. For example, the closer the distance between entities, the greater the similarity, and they can be considered duplicate entities. Specifically, according to an embodiment of the present invention, the distance between two entities is:
[0037]
[0038] where E a denotes entity a, and E b denotes entity b, denotes the translation relationship between two entities a and b. In addition, in order to store entities and the relationships between entities, the Neo4j graph database is used to store the fused knowledge graph KG a .
[0039] According to an embodiment of the present invention, the preset knowledge graph completion method includes knowledge graph completion based on the TransE model. In addition, other methods can also be adopted. Traditional knowledge graph completion methods are usually divided into two categories: methods based on representation learning (such as TransE, ConvE, Rotate, etc.) and methods based on predicate logic rules (such as AMIE, Neural-LP, etc.). The former requires rich training data, while the latter depends on the number of high-quality rules. Therefore, it can be selected according to different application scenarios.
[0040] To further ensure the effectiveness of the knowledge graph, according to an embodiment of the present invention, a preset knowledge graph completion method is used to complete the constructed knowledge graph to obtain an updated knowledge graph as follows: S21. Extract a batch of triples from the constructed knowledge graph as correct triples, and generate incorrect triples based on the correct triples; S22. Calculate the vector distance between the entity vector plus the relationship vector and the tail entity vector of each triple; S23. Calculate the loss based on the vector distance of any correct triple and the vector distance of any incorrect triple, and update the parameters of the TransE model;
[0041] S24. Repeat steps S21 - S23 until a predetermined number of iterations is reached or the loss function converges, and use the trained TransE model to complete incomplete triples and find and delete incorrect triples. According to an embodiment of the present invention, the loss is:
[0042] L = ∑ (h,r,t)∈S ∑ (h′,r,t′)∈S (d(h + r, t)+d(h′ + r, t″)),
[0043] where (h, r, t) indicates a correct triple, (h′, r, t′) indicates an incorrect triple, d(h + r, t) indicates the vector distance of the correct triple, d(h′ + r, t′) indicates the vector distance of the incorrect triple, and S indicates the set of triples.
[0044] After multiple iterations of training, the model learns the low - dimensional vector representations of entities and relationships. At this time, these vectors can be used for reasoning. For incomplete triples in the knowledge graph, for example, given the head entity and the relationship, try to find the tail entity vector that is closest to the head entity vector plus the relationship vector to complete the triple. For example, if (Supplier B, relationship,?) is known, the distances between all possible tail entity vectors and the vector of Supplier B plus the relationship vector can be calculated, and the tail entity with the minimum distance is selected as the completion result. In addition, for triples that may be incorrect, calculate their distances. If the distance is too large and exceeds a preset threshold, it is considered that the triple may be incorrect and is deleted from the knowledge graph. Through the above measures, the knowledge graph KG b is finally obtained after reasoning by the TransE model, where incomplete triples are completed, incorrect triples are deleted, and the knowledge graph is more accurate and complete, providing a more reliable basis for supplier risk assessment.
[0045] According to an embodiment of the present invention, based on the updated knowledge graph, risk prediction is performed using a pre-trained risk prediction model as follows: S31. Based on the updated knowledge graph, construct an entity feature matrix and a relationship feature matrix; S32. Through the entity-level attention network in the pre-trained risk prediction model, calculate the weights between the associated entity and all connected neighbor entities under any relationship path, and linearly combine the weights with the neighbor entities to obtain the entity features after feature fusion under any relationship path;
[0046] S33. Fuse the semantic information under different relationship paths, learn the relationship-level attention coefficients of each relationship path, and use them as the input of the relationship-level attention network to obtain a new feature vector of any entity after passing through the relationship-level attention network layer; S34. After the new feature vector is normalized, obtain the predicted risk level of any entity. Take the feature dimension C as the risk category to be classified. According to the risk level division standard, the risk categories are mainly divided into high risk, medium risk and low risk, that is, C = 3. Then the final classification feature representation of the entity features is Q′ i =[f 1 ,f 2 ,f 3 , corresponding to the final classification feature values of the entity respectively. After normalization through the softmax function, obtain the probability of the risk level where entity i is located. The formula is as follows:
[0047]
[0048] Among them, fx x is a feature value in the feature vector of entity i, is the probability value that entity i belongs to f x . Finally, select the category with the largest probability value as the risk level of entity i, obtain the risk assessment level of the supplier entity, and propose countermeasure suggestions for the evaluation results.
[0049] According to an embodiment of the present invention, under any relationship path, the weights between the associated entity and all connected neighbor entities are:
[0050]
[0051] Among them, are all neighbor entities of entity i under relationship path r m , is the weight vector of the entity-level attention network, e i , e j and e n represent any entity feature, r m represents the path feature between different entities. The relationship-level attention coefficients of each relationship path are:
[0052]
[0053] Among them, g is the attention mechanism vector, W b is a parameterized weight matrix, and b is a bias vector. indicates the entity feature after feature fusion under the path r m In addition, the entity feature matrix is represented as E = [e 1 , e 2 ,... e n , and the path feature matrix between different entities is represented as R = [r 1 , r 2 ,..., r m . Additionally, the entity-level attention mechanism head is set to k = 6, and the entity feature vector after feature fusion under the path r m is Similarly, the relation-level attention mechanism head is also set to k = 6, and the new feature vector representation Q i = [q 1 , q 2 ,..., q n of entity i after passing through the relation-level attention network layer.
[0054] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0055] The computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing.
[0056] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A supplier risk prediction method based on knowledge graph, characterized in that: The following steps are involved: S1. Obtain the supplier's current operating data and historical operating data, and build a knowledge graph based on the acquired data, wherein the current operating data includes professional certificate data, training data and financial status data, and the historical operating data includes material quality inspection data, quality and safety inspection data, contract breach data, industry violation record data, historical labor dispute records and historical labor compliance data; S2. Use a preset knowledge graph completion method to complete the constructed knowledge graph to obtain an updated knowledge graph; S3. Based on the updated knowledge graph, risk prediction is performed using the pre-trained risk prediction model.
2. According to a method for predicting supplier risk based on knowledge graph according to claim 1, it is characterized in that: The knowledge graph constructed based on the acquired data is: S11, extracting the acquired data, and forming explicit triples and implicit triples through entity recognition and relationship recognition, wherein the current operation data corresponds to the explicit triples, and the historical operation data corresponds to the implicit triples; S12. Align entities through the translation model to obtain the fused knowledge graph.
3. According to a method for predicting supplier risk based on knowledge graph according to claim 2, it is characterized in that: The acquired data is extracted, and explicit triples and implicit triples are formed through entity recognition and relationship recognition as follows: S111, performing word vector, part-of-speech vector and word length vector representation on the acquired data; S112, calculating the correlation of adjacent elements in the data through a conditional random field model, learning entity representation and performing entity recognition, and further extracting the relationship between entities to form entity, entity and relationship triplets; S113. Align entities through the translation model to obtain a fused knowledge graph.
4. According to claim 3, a supplier risk prediction method based on knowledge graph is characterized in that: By aligning entities through the translation model, the fused knowledge graph is obtained as follows: S1131, using the translation model to calculate the distance between the mapped entities; S1132. Determine the similarity of entities based on a preset distance threshold, integrate similar entities, and construct new relationships to obtain a fused knowledge graph.
5. The supplier risk prediction method based on knowledge graph according to claim 4 is characterized in that: The distance between the two entities is: Among them, E a Indicates entity a, E b Indicates entity b, Indicates a translation relationship between two entities a and b.
6. According to claim 1, a supplier risk prediction method based on knowledge graph is characterized in that: The preset knowledge graph completion method includes knowledge graph completion based on the TransE model.
7. The supplier risk prediction method based on knowledge graph according to claim 6 is characterized in that: The constructed knowledge graph is completed using the preset knowledge graph completion method to obtain the updated knowledge graph: S21, extracting a batch of triples from the constructed knowledge graph as correct triples, and generating error triples based on the correct triples; S22, calculating the vector distance between the entity vector of each triple plus the relationship vector and the tail entity vector; S23, calculating the loss based on the vector distance of any correct triplet and the vector distance of any incorrect triplet, and updating the parameters of the TransE model; S24. Repeat steps S21-S23 until a predetermined number of iterations is reached or the loss function converges, and use the trained TransE model to complete incomplete triplets, and find and delete erroneous triplets.
8. The supplier risk prediction method based on knowledge graph according to claim 7 is characterized in that: The losses are: L=∑ (h,r,t)∈S ∑ (h′,r,t′)∈S (d(h+r,t)+d(h′+r,t ′ ′)), Among them, (h, r, t) indicates the correct triple, (h ′ ,r,t ′ ) indicates an erroneous triple, d(h+r,t) indicates the vector distance of a correct triple, d(h'+r,t') indicates the vector distance of an erroneous triple, and S indicates a triplet set.
9. The supplier risk prediction method based on knowledge graph according to claim 1 is characterized in that: Based on the updated knowledge graph, the risk prediction is performed using the pre-trained risk prediction model: S31. Based on the updated knowledge graph, construct an entity feature matrix and a relationship feature matrix; S32, calculating the weights between the associated entity and all connected neighbor entities under any relationship path through the entity-level attention network in the pre-trained risk prediction model, and linearly combining the weights with the neighbor entities to obtain entity features after feature fusion under any relationship path; S33, fusing the semantic information under different relational paths, learning the relational-level attention coefficient of each relational path, and using it as the input of the relational-level attention network to obtain a new feature vector of any entity after passing through the relational-level attention network layer; S34. After the new feature vector is normalized, the predicted risk level of any entity is obtained.
10. A supplier risk prediction method based on knowledge graph according to claim 9, characterized in that: Under any relationship path, the weight between the associated entity and all connected neighbor entities is: in, is the relationship path r m All neighbor entities of entity i, is the weight vector of the entity-level attention network, e i 、e j and e n Represents any entity feature, r m Represents the path features between different entities, and the relation-level attention coefficient of each relation path is: Among them, g is the attention mechanism vector, W b is the parameterized weight matrix, b is the bias vector, Indicates the path r m The entity features after the following features are fused.