Supplier name correction method and system
By building a knowledge graph and supplier name correction model, the problem of inconsistent supplier names in different business scenarios is solved, and the accurate correction and accuracy of supplier data is achieved, ensuring the smooth and efficient operation of the enterprise.
Patent Information
- Application Number
- CN202510143619.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In the complex and changeable business ecosystem, supplier names are inconsistent in different business scenarios. Traditional string matching technology cannot effectively identify the true identity of the supplier, resulting in a decrease in data accuracy.
By collecting multi-source data, building a knowledge graph covering all-round information of suppliers, and building a supplier name correction model, accurately identifying supplier names in different business scenarios, and achieving effective correction.
It improves the accuracy of supplier data, ensures the smooth and efficient corporate supply chain management and decision-making, and reduces communication costs, management costs and operational risks caused by name chaos.
Smart Images

Figure CN120045723A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method and system for correcting supplier names. Background Art
[0002] In today's complex, ever-changing and highly digital business ecosystem, there is intensive data exchange between suppliers and enterprises. The sources of enterprise-acquired supplier data are extensive and complex, including but not limited to internal procurement records, contract documents, external business databases, industry reports, and data provided by suppliers themselves.
[0003] The same supplier may have inconsistent names in different business scenarios or data sources. For example, in the procurement business scenario, to ensure the rigor and accuracy of contracts, enterprises usually use the full name of the supplier, such as "XX Technology Co., Ltd. in XX City"; however, in the logistics business scenario, for the convenience of logistics personnel to quickly identify and record, an abbreviated name may be used, such as "XX Technology". In the financial business scenario, the supplier name needs to match bank account information, invoice information, etc., and may be the complete registered name plus detailed information such as the unified social credit code, for example, "XX Electronic Co., Ltd. in XX City (91XXXXXXXXXXXXXX)". But in the after-sales service business scenario, for the convenience of customer communication and record, a more user-friendly name may be used, like "XX Electronic After-sales".
[0004] Under the traditional data processing paradigm, enterprises attempt to solve the problem of supplier name correction through simple string matching techniques, usually using the trade name as an identifier to identify enterprises. This method only stays at the surface character comparison and cannot insight into the hidden clues of the true identity of the supplier behind the business scenario. For example, the supplier name recorded in the production and manufacturing scenario may be its trademark name, while the full registered name of the enterprise is used in financial settlement. Simple string matching is extremely likely to misjudge them as two different suppliers, and traditional means cannot organically integrate the supplier name with rich business associations, qualifications, finances, credit, etc. information behind it to achieve more accurate correction of supplier data.
[0005] Therefore, there is an urgent need for a method that can accurately identify supplier names in different business scenarios, thereby achieving effective correction and improving the accuracy of supplier data. Summary of the Invention
[0006] In view of this, the present invention proposes a method and system for correcting supplier names, which can collect data from multiple sources, construct a knowledge graph covering all aspects of suppliers, and through a supplier name correction model, accurately identify supplier names and corresponding data problems in different business scenarios, effectively correct them, improve the accuracy of supplier data, and ensure the smooth and efficient operation of enterprise supply chain management, decision-making and other operation links.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A method for correcting supplier names, comprising:
[0009] Collect data from multiple data sources, parse the data to obtain the supplier identifier corresponding to each data, and determine each supplier s i corresponding supplier data;
[0010] Based on the supplier data, construct a knowledge graph for each supplier s i The knowledge graph takes the supplier s i as the center and radiates multiple supplier-related entities outward. The supplier-related entities include supplier name entities, supplier basic information entities, supplier business association entities, supplier qualification and certification entities, supplier financial entities, supplier credit entities, and supplier dynamic event entities. The supplier name entities include the supplier s i in the supplier names corresponding to different business scenarios s y ;
[0011] Construct a supplier name correction model, which includes an input layer, an extraction layer, a business scenario recognition layer, a first matching layer, a name update layer, a second matching layer, a correction strategy generation layer, an update layer, and an output layer;
[0012] Receive supplier data through the input layer, extract supplier name data from the supplier data through the extraction layer, determine the business scenario corresponding to each supplier name through the business scenario recognition layer, search for matching items in the knowledge graph guided by the business scenario through the first matching layer. If no matching item is found, correct the supplier name based on the knowledge graph through the name update layer. After the correction is completed, locate the error type of the supplier data based on the knowledge graph through the second matching layer, generate a correction strategy based on the error type through the correction strategy generation layer, execute the correction strategy through the update layer, and output the correction result through the output layer;
[0013] Obtain the supplier data to be corrected, and input the supplier data to be corrected into the supplier name correction model to obtain the correction result.
[0014] Based on the above technical solutions, the present invention can also be improved as follows:
[0015] Optionally, the method for correcting supplier names further includes:
[0016] Collect the latest data, parse the latest data to obtain the supplier identifier corresponding to each piece of the latest data, and determine each supplier s based on the supplier identifier i corresponding latest supplier data;
[0017] Update the corresponding knowledge graph based on the latest supplier data.
[0018] Optionally, determining the business scenario corresponding to each supplier name through the business scenario recognition layer includes:
[0019] Calculate the business scenario matching value through formula (1);
[0020] M(s y , sc j ) = W(s y , sc j ) × F cos (V(s y ), V(sc j )) + B(s y , sc j ) Formula (1);
[0021] In the formula, M(s y , sc j ) is the matching value between the supplier name s y and the business scenario sc j , W is the weight function, F cos is the cosine similarity function, V(s y ) is the vector obtained after feature extraction and vectorization of the data corresponding to the supplier name s y , V(sc j ) is the vector obtained after feature extraction and vectorization of the data of the business scenario sc j , and B is the bias term;
[0022] Select the business scenario sc y with the highest M(s j ) as the business scenario corresponding to the supplier name s j . y
[0023] Optionally, searching for matching items in the knowledge graph by the first matching layer guided by the business scenario includes:
[0024] Calculate the matching value between the business scenario sc j and the entity e j through formula (2);
[0025]
[0026] Wherein, F(sc j , e j ) is the matching value between the business scenario sc j and the entity e j . KG is the knowledge graph, r is the relationship in the knowledge graph KG, ω r is the weight related to the relationship r, and I(sc j , e j , r) is the indicator function;
[0027] When the value of I(sc j , e j , r) is 1, there is a connection between the business scenario sc j , the entity e j and the relationship r. When the value of I(sc j , e j , r) is 0, there is no connection between the business scenario sc j , the entity e j and the relationship r;
[0028] When F(sc j , e j ) is greater than or equal to the preset threshold of the matching value, it is determined that the business scenario sc j matches the entity e j . When F(sc j , e j ) is less than the preset threshold of the matching value, it is determined that the business scenario sc j does not match the entity e j .
[0029] Optionally, the method for locating the error type of the supplier data based on the knowledge graph through the second matching layer includes:
[0030] Calculating the comprehensive error score of the supplier data T through formula (3);
[0031]
[0032] Wherein, S(T) is the comprehensive error score of the supplier data T, l is the total number of supplier data error types, j is the index variable, α j is the weight of the error type et j , and S(et j , T) is the jth supplier data error type;
[0033] When S(T) is greater than or equal to the preset threshold of the error score, it is determined that there is an error in the supplier data. Analyze the proportion of the score of each error type et j in the total score S(T) to determine the main error type, and locate the error type of the supplier data based on the main misalignment type.
[0034] Optionally, the method for correcting the supplier name further includes:
[0035] Comparing the correction result output by the supplier name correction model with the actual correction data to obtain an error value of the correction result;
[0036] Determining whether the error value exceeds a preset error threshold. If it exceeds, determining a parameter to be adjusted based on the error value, and adjusting the parameter to be adjusted until the error value of the text supplier name correction model does not exceed the preset error threshold, so as to obtain an optimal supplier name correction model.
[0037] Optionally, determining the parameter to be adjusted based on the error value, and adjusting the parameter to be adjusted until the error value of the text supplier name correction model does not exceed the preset error threshold, so as to obtain an optimal supplier name correction model, includes:
[0038] Dividing the parameters of the supplier name correction model into parameter groups including input layer parameters, extraction layer parameters, business scenario recognition layer parameters, first matching layer parameters, name update layer parameters, second matching layer parameters, correction strategy generation layer parameters, update layer parameters, and output layer parameters;
[0039] Conducting small-range perturbation experiments on each parameter group one by one, and determining sensitive parameters sensitive to the error value based on the error value between the correction result output by the supplier name correction model after each perturbation and the actual correction data, and determining the sensitive parameters as the parameters to be adjusted;
[0040] Calculating new parameter values corresponding to the parameters to be adjusted, adjusting the parameters to be adjusted based on the new parameter values, obtaining a new correction result output by the adjusted supplier name correction model, and calculating a new error value based on the correction result;
[0041] Comparing the new error value with the preset error threshold. If the new error value does not exceed the preset error threshold after continuous multiple iterations, terminating the parameter adjustment to obtain an optimal supplier name correction model.
[0042] A supplier name correction system includes:
[0043] A task allocation module, configured to construct a data collection task in response to a data collection request, and allocate the data collection task to each crawling node in a distributed node architecture;
[0044] A data collection module, configured to enable the crawling node to access different data sources based on the data collection task to obtain collected data;
[0045] An intelligent scheduling module, which is used to calculate the collection efficiency and progress of each crawling node through an intelligent scheduling algorithm after the crawling nodes complete a round of data collection, and adjust the task allocation based on the collection progress and efficiency;
[0046] A data quality evaluation module, which is used to perform the first data processing on the collected data. The first data processing includes data deduplication, data format unification, and missing value processing of the data, and perform data quality evaluation on the collected data after the first data processing;
[0047] An invalid information removal and text correction module, which is used to remove invalid information and correct text from the collected data after data quality evaluation;
[0048] A multi-data source integration module, which is used to integrate the collected data that has completed the first data processing from different data sources to obtain integrated data, and perform the second data processing on the integrated data to obtain LLM data. The second data processing includes data deduplication and data format unification.
[0049] An electronic device includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the steps of the method are implemented.
[0050] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.
[0051] The present invention has the following advantages:
[0052] In the supplier name correction method of the present invention, data is collected from multiple data sources, which can collect comprehensive information related to suppliers to the greatest extent, avoid information loss caused by limited data sources, parse the data and determine the supplier identifier to lock the corresponding supplier data, and can effectively sort out a large amount of messy data, so that subsequent processing for each supplier has an accurate and complete data basis, improving the comprehensiveness and accuracy of the entire correction work, and laying a foundation for accurately identifying different supplier individuals and performing targeted name correction.
[0053] In the supplier name correction method of the present invention, a knowledge graph is constructed with the supplier as the core, connecting numerous related entities, which helps to explore the internal relationships between information in different dimensions. For example, it can clearly show the connections between the changes in the supplier name in different business scenarios and the basic information of the supplier, business transactions, etc. This way of integrating multi-dimensional information breaks through the limitation of traditionally viewing supplier data in isolation, provides strong support for more in-depth and accurate understanding of suppliers and subsequent name correction based on multiple clues, enabling the correction decision to comprehensively consider multiple factors and be more scientific and reasonable.
[0054] In the supplier name correction method of the present invention, each layer of the supplier name correction model has clear division of labor and close cooperation. The input layer can widely accept supplier data in different formats to ensure the smooth inflow of data. The extraction layer can accurately extract the key supplier name data, focus on the core content, and reduce the interference of irrelevant information. The business scenario recognition layer further clarifies the scenario corresponding to the supplier name, which helps to carry out the correction work in line with the actual business situation. The first matching layer searches for matching items according to the business scenario with the help of the knowledge graph, and can efficiently utilize the associated information in the graph to achieve accurate matching. If the matching fails, the name is updated. Based on the knowledge graph, the error type of the supplier data can be located, enabling in-depth analysis of the supplier data. The knowledge graph contains comprehensive information about the supplier, such as multiple dimensions including name entity, basic information, business association, qualification certification, etc. By comparing the supplier data after the supplier name is corrected with the standard information in the knowledge graph, the error can be accurately identified.
[0055] The supplier name correction method of the present invention ensures the orderly and efficient development of the correction work. Receiving data ensures the input source of the data, and extracting name data makes the subsequent operations more targeted. Accurately identifying the business scenario makes the correction conform to the actual application scenario and avoids blind operations. Searching for matching items with the help of the knowledge graph makes full use of the associative advantages of the knowledge graph, improving the matching success rate and accuracy. When the search is fruitless, updating the name based on the knowledge graph makes the most reasonable correction based on the existing comprehensive information. Further locating the error type can accurately grasp the problem, providing a basis for generating appropriate correction strategies. Executing the strategy and outputting the result, the operation method is simple and convenient, enabling enterprises to easily connect the supplier data with name problems to the already constructed supplier name correction model for processing without complex additional operation processes. Through the automatic processing of the supplier name correction model, the correction result is quickly and accurately output, timely updating the enterprise's supplier data, ensuring that in many business links related to suppliers such as procurement, finance, and logistics, the enterprise can carry out work based on accurate and consistent supplier names, improving business efficiency and reducing communication costs, management costs, and potential operation risks caused by name confusion. Brief Description of the Drawings
[0056] For purposes of illustration and not limitation, the present invention will now be described in connection with embodiments and the accompanying drawings of the present invention, wherein:
[0057] Figure 1 is a schematic flowchart of a method for correcting a supplier name in an embodiment of the present invention;
[0058] Figure 2 is a schematic diagram of the main components of a supplier name correction system in an embodiment of the present invention;
[0059] Figure 3 is a schematic diagram of the physical structure of an electronic device provided by the present invention. Specific Embodiments
[0060] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0061] It should be noted that the terms "first", "second", etc. in the specification and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present invention here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0062] It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0063] Figure 1 is a schematic flowchart of a method for correcting a supplier name in an embodiment of the present invention. As Figure 1 shown, the method for correcting a supplier name provided by the embodiments of the present invention includes the following steps S101 to S104.
[0064] S101, collect data from multiple data sources, parse the data to obtain a supplier identifier corresponding to each data, and determine supplier data corresponding to each supplier based on the supplier identifier.
[0065] Specifically, first, a variety of data sources cover various business systems within an enterprise. For example, the procurement management system records detailed information on contracts signed with suppliers; the warehousing management system stores supplier-related data involved in the inbound and outbound of goods; the financial management system contains financial transaction data such as payments, receipts, and invoices of suppliers. At the same time, it also includes external data sources, such as industry authoritative databases, which provide basic information and reputation evaluations of suppliers within the entire industry; commercial credit platforms, from which credit ratings and historical credit records of suppliers can be obtained; and various qualification certification documents and enterprise introduction materials provided by the suppliers themselves.
[0066] For the complex and variously formatted data collected, data parsing technology is used to parse and process it according to preset data structures and rules. For example, by identifying specific field formats, key fields, etc., information that can uniquely identify a supplier is extracted as a supplier identifier, such as the unified social credit code of the enterprise, the registration number, or the unique supplier number assigned within the enterprise's system. Based on these supplier identifiers, the data scattered in different data sources is integrated and classified, so as to accurately determine the complete supplier data corresponding to each supplier, laying a solid data foundation for subsequent construction of knowledge graphs and name correction.
[0067] S102, construct a knowledge graph for each supplier based on the supplier data.
[0068] Specifically, the knowledge graph takes supplier s i as the center and radiates a variety of supplier-related entities outward. The supplier-related entities include supplier name entities, supplier basic information entities, supplier business association entities, supplier qualification and certification entities, supplier financial entities, supplier credit entities, and supplier dynamic event entities. The supplier name entities include supplier s i in the supplier names s corresponding to different business scenarios y .
[0069] When constructing the knowledge graph, taking the supplier as the core node, for the supplier name entity, all name expression forms in different business scenarios, such as the official full name used by the supplier in the procurement link, the commonly used abbreviated name in daily communication or logistics distribution scenarios, and the foreign name corresponding to the cross-border business scenario, will be collected and sorted out, and associated and stored as different attributes of this entity.
[0070] The supplier basic information entity covers basic information such as the enterprise's registered address, legal representative, business scope, contact information, etc. Through the relationship construction of the knowledge graph, the association between these basic information and the supplier entity can be clearly presented.
[0071] The supplier business association entity records the business transactions between the supplier and its upstream and downstream enterprises, such as cooperation projects, cooperation time, cooperation models, etc. For example, which enterprises it supplies goods to and which enterprises it purchases raw materials from, thus reflecting the position of the supplier in the industrial chain and its business network relationship.
[0072] The supplier qualification and certification entity stores various industry qualification certificate information obtained by the supplier, such as quality management system certification, environmental management system certification, production license qualifications for specific industries, etc., to clarify its capabilities and compliance standards in the professional field.
[0073] The supplier financial entity includes important financial indicators such as balance sheet data, income statement data, and cash flow situation, directly reflecting the financial health status and operating strength of the supplier.
[0074] The supplier credit entity integrates credit-related information such as credit ratings given by credit rating agencies, default records, and overdue payments, to help enterprises evaluate the risk level of cooperation with this supplier.
[0075] The supplier dynamic event entity records major events related to the supplier in real time, such as enterprise mergers and acquisitions, new product launches, major quality accidents, or legal disputes involved, etc., facilitating enterprises to timely grasp the latest dynamic changes of the supplier. Then, based on these rich and interrelated entities and relationships, a complete knowledge graph is constructed, providing comprehensive and in-depth information support for subsequent name correction work.
[0076] The described supplier name correction method further includes:
[0077] Collect the latest data, parse the latest data to obtain the supplier identifier corresponding to each latest data, and determine the latest supplier data corresponding to each supplier s i Based on the latest supplier data, update the corresponding knowledge graph. Collecting the latest data and updating the knowledge graph based on this can ensure that the supplier information covered in the graph is the latest and accurate. For example, the name of the supplier may change due to enterprise restructuring, renaming, etc. Timely updating can avoid problems such as matching errors and data association mistakes caused by the continued use of the old name, enabling subsequent operations such as supplier name correction and error type positioning based on the knowledge graph to be carried out on the basis of accurate data, improving the accuracy of the overall correction work.
[0078] S103, construct a supplier name correction model. The supplier name correction model includes an input layer, an extraction layer, a business scenario recognition layer, a first matching layer, a name update layer, a second matching layer, a correction strategy generation layer, an update layer, and an output layer.
[0079] Specifically, the supplier data is received through the input layer. The input layer has good data compatibility and can accept supplier data from different formats (such as structured database table data, semi-structured XML files, and unstructured text files, etc.) and different channels (including various internal business systems of the enterprise and various external data sources), ensuring that the data can flow smoothly into the model for subsequent processing.
[0080] Through the extraction layer, the supplier name data in the supplier data is extracted. The extraction layer uses natural language processing techniques (such as named entity recognition algorithms, etc.) and data extraction rules to accurately extract the data part related to the supplier name from the received supplier data, focus on the key name information, and eliminate other irrelevant redundant data, providing a clear and definite object for subsequent more targeted operations.
[0081] Through the business scenario recognition layer, the business scenario corresponding to each supplier name is determined. The business scenario recognition layer will analyze the extracted supplier name data, combine the pre-constructed business scenario knowledge base and relevant business rules, and accurately judge the specific business scenario corresponding to each supplier name through means such as feature matching and semantic analysis, for example, whether it belongs to the procurement business scenario, logistics transportation scenario, after-sales service scenario, or other business scenarios.
[0082] Calculate the business scenario matching value through formula (1).
[0083] M(s y ,sc j )=W(s y ,sc j )×F cos (V(s y ),V(sc j ))+B(s y ,sc j ) Formula (1);
[0084] In the formula, M(s y ,sc j ) is the matching value between the supplier name s y and the business scenario sc j , W is the weight function, F cos is the cosine similarity function, V(s y ) is the vector obtained after feature extraction and vectorization of the data corresponding to the supplier name s y , V(sc j ) is the vector obtained after feature extraction and vectorization of the data of the business scenario sc j , and B is the bias term;
[0085] Weight function For adjusting the supplier name s y and the business scenario sc j The matching weight between them. Different supplier names and business scenarios may have different importance or relevance, and the weight function can be adjusted according to specific circumstances to reflect this difference.
[0086] Cosine similarity function Is used to measure the similarity between two vectors. The value range of cosine similarity is between -1 and 1, and the larger the value, the more similar the two vectors are.
[0087] Bias term Is used to adjust the matching value. It can compensate for systematic biases that may occur during feature extraction and vectorization, or be adjusted according to specific business rules to ensure the accuracy of the matching value.
[0088] Select the business scenario sc that maximizes M(s y , sc j ) as the business scenario corresponding to the supplier name s j y
[0089] Search for matching items in the knowledge graph guided by the business scenario through the first matching layer. The first matching layer conducts in-depth search and matching in the already constructed knowledge graph guided by the identified business scenario. For example, if the current business scenario is a procurement scenario, this layer will search in the knowledge graph for all entity information related to the supplier in the procurement business, including its commonly used correct name in the procurement scenario, the procurement departments it cooperates with, and other relevant matching items, and make efficient and accurate matching judgments by making full use of the relevance between entities in the knowledge graph;
[0090] Calculate the matching value between the business scenario sc j and the entity e j through formula (2);
[0091]
[0092] In the formula, F(sc j , e j ) is the matching value between the business scenario sc j and the entity e j , KG is the knowledge graph, r is the relationship in the knowledge graph KG, ω r is the weight related to the relationship r. Different relationships may have different importance when judging the matching between the business scenario and the entity. The weight ω r is used to reflect this importance, and I(sc j , e j , r) is the indicator function;
[0093] When the value of I(sc j , e j , r) is 1, there is a connection between the business scenario sc j , entity e j and relationship r. When the value of I(sc j , e j , r) is 0, there is no connection between the business scenario sc j , entity e j and relationship r;
[0094] ∑ r∈KG ω r ×I(sc j , e j , r) is actually the sum of the weights ω r of all related relationships r.
[0095] ∑ r∈KG ω r is the sum of the weights ω r of all relationships r in the knowledge graph. The purpose of this design is to normalize the weighted sum of the numerator so that the matching value F(sc j , e j ) is within a reasonable range.
[0096] When F(sc j , e j ) is greater than or equal to the preset threshold of the matching value, it is determined that the business scenario sc j matches the entity e j . When F(sc j , e j ) is less than the preset threshold of the matching value, it is determined that the business scenario sc j does not match the entity e j .
[0097] If no matching item is found, the supplier name is corrected based on the knowledge graph through the name update layer. If no matching item is found in the knowledge graph, it is determined that the supplier name is incorrect, and the most suitable supplier name for the current business scenario is selected from the knowledge graph to reasonably correct the incorrect supplier name.
[0098] After the correction is completed, the second matching layer locates the error types of the supplier data based on the knowledge graph. Although all the supplier names in the supplier data tend to be correct after the supplier name correction, the supplier data contains multiple dimensions of information, such as contact information, business scope, qualification status, etc., which are equally important. By further checking whether the entire supplier data is correct in the knowledge graph based on the correct supplier name, it is possible to comprehensively control the quality of the data, avoid situations where other data errors besides the name affect subsequent business processes related to suppliers, such as incorrect purchase order issuance and inconsistent cooperation negotiation object information, and ensure the integrity and accuracy of the entire supplier data, thereby improving the overall quality level of the data. Incorrect supplier data may pose many risks in actual business operations. For example, if the qualification validity period data of a supplier is incorrect, it may lead to the enterprise choosing an unqualified supplier for cooperation, facing compliance risks; if key information such as the financial account is incorrect, it will cause problems in payment, affecting the security of funds. Through this checking and further correction mechanism, these potential data errors can be discovered and corrected in advance, minimizing various risks brought to the enterprise's business development due to data problems and ensuring the smooth and compliant progress of business processes. During the business execution process, if relying on incorrect supplier data, situations such as repeated communication and confirmation, process interruption and rework often occur. For example, in the logistics and distribution link, if the supplier address data is incorrect, the goods may not be delivered on time and accurately, and additional time and effort will be required to re-verify and adjust later. By meticulous checking and correction to ensure the accuracy of the supplier data, it can smoothly promote various business processes involving suppliers, reduce unnecessary time waste and resource consumption, and improve the operation efficiency of the overall business process.
[0099] Calculate the comprehensive error score of the supplier data T through formula (3);
[0100]
[0101] In the formula, S(T) is the comprehensive error score of the supplier data T, l is the total number of error types of the supplier data, j is the index variable, and α j is the weight of the error type et j . Different error types may have different degrees of impact on the quality of the supplier data, and the weight α j is used to reflect this difference. For example, errors in certain key data (such as incorrect supplier names) may have a higher weight than errors in other secondary data (such as non-critical information errors in contact information), and S(et j , T) is the j-th error type of the supplier data, and this score reflects the specific error degree of the j-th error type in the supplier data T;
[0102] Indicates the summation of all supplier data error types. Here, l is the total number of supplier data error types, and j is an index variable used to traverse all error types.
[0103] When S(T) is greater than or equal to the preset error score threshold, it is determined that there are errors in the supplier data, and each error type et is analyzed. j The proportion of the score of et in the total score S(T) is used to determine the main error type, and the error type of the supplier data is located based on the main misalignment type.
[0104] The correction strategy generation layer generates a correction strategy based on the error type, selects a suitable strategy template from the preset correction strategy template library, and combines the specific situation of the current supplier (such as business importance, cooperation history, etc.) to generate a practical correction strategy by adjusting and optimizing parameters, etc., to ensure that the discovered error problems can be effectively solved.
[0105] The correction strategy is executed by the update layer, and corresponding parts in the supplier data are modified, replaced, supplemented, etc. according to the strategy requirements, so that the supplier data is more accurate and standardized.
[0106] The correction result is output through the output layer, and the output result can be conveniently fed back to various relevant business systems of the enterprise for data update, so as to ensure the accuracy and consistency of the supplier data in the entire enterprise operation.
[0107] S104. Obtain the supplier data to be corrected, and input the supplier data to be corrected into the supplier name correction model to obtain the correction result.
[0108] Specifically, the step of obtaining the supplier data to be corrected involves actively collecting data marked as possibly having name problems from various business links within the enterprise. For example, data with non-standard name formats found during data integration, data with obvious differences for the same supplier in different business records, or data suspected of name errors feedback by business personnel, etc. At the same time, a regular data screening mechanism can also be set up to scan all supplier data in the enterprise system and screen out data that meets the preset name error characteristics as the supplier data to be corrected.
[0109] After organizing these supplier data to be corrected according to the format and specifications required by the model, they are input into the already constructed supplier name correction model. The model will automatically perform a series of operations in sequence according to the internal set processing logic of each layer, including data reception, name data extraction, business scenario recognition, matching item search, name correction, error type positioning, correction strategy generation, and execution strategy. The entire process requires little manual intervention and can efficiently and accurately output the correction results. The output correction results can be directly fed back to the corresponding business system to update and replace the original supplier data, thus quickly correcting the chaos, errors, etc. in the supplier names within the enterprise system, ensuring that various subsequent supplier-related operations of the enterprise, such as procurement, supply chain management, and financial settlement, can be smoothly carried out based on accurate name information, improving the overall operation efficiency, and reducing the adverse impacts such as increased communication costs and elevated business risks caused by inaccurate names.
[0110] The supplier name correction method further includes:
[0111] Compare the correction results output by the supplier name correction model with the actual correction data to obtain the error value of the correction results;
[0112] Determine whether the error value exceeds the preset error threshold. If it exceeds, determine the parameter to be adjusted based on the error value, and adjust the parameter to be adjusted until the error value of the text supplier name correction model does not exceed the preset error threshold to obtain the optimal supplier name correction model.
[0113] Divide the parameters of the supplier name correction model into parameter groups including input layer parameters, extraction layer parameters, business scenario recognition layer parameters, first matching layer parameters, name update layer parameters, second matching layer parameters, correction strategy generation layer parameters, update layer parameters, and output layer parameters;
[0114] Conduct a small-scale perturbation experiment on each parameter group one by one. Based on the error value between the correction results output by the supplier name correction model after each perturbation and the actual correction data, determine the sensitive parameters sensitive to the error value, and determine the sensitive parameters as the parameters to be adjusted;
[0115] Calculate the new parameter values corresponding to the parameters to be adjusted, adjust the parameters to be adjusted based on the new parameter values, obtain the new correction results output by the adjusted supplier name correction model, and calculate the new error value based on the correction results;
[0116] Compare the new error value with the preset error threshold. If the new error value does not exceed the preset error threshold after continuous multiple iterations, terminate the parameter adjustment to obtain the optimal supplier name correction model.
[0117] An example:
[0118] Suppose there is a large manufacturing enterprise that has long cooperated with numerous suppliers, covering multiple business areas such as raw material supply, component processing, and packaging material provision. The enterprise has built a supplier name correction model to ensure the accuracy of supplier data for better conducting business processes such as procurement and cooperation evaluation;
[0119] The enterprise's procurement department recently compiled a new supplier directory, which was collected through different channels and includes some newly expanded potential suppliers and updated information of some old suppliers; in this supplier directory to be processed (i.e., the supplier data to be corrected), there are approximately 200 supplier records, and each record contains multiple field information such as supplier name, location, contact information, main products, and past cooperation situation. As shown in Table 1, one of the records is as follows:
[0120] Table 1
[0121]
[0122] Due to possible manual input errors, untimely information updates, etc. during the collection process, there are some inaccurate or non-compliant-with-the-enterprise's-internal-uniform-specification situations in the supplier names in this directory, so it is necessary to process them using the supplier name correction model.
[0123] Import these 200 supplier records as a whole into the supplier name correction model. The supplier name correction model first receives this data through the input layer, and then the extraction layer extracts the supplier name data from each record. For example, the "XX Hardware Factory (formerly XX Hardware Products Factory)" in the above record will be extracted separately.
[0124] Next, the business scenario recognition layer will determine the corresponding business scenario for each supplier based on relevant information such as past cooperation data and main products. For example, the business scenario corresponding to this hardware factory is to provide hardware production raw materials, which are mainly used in the product assembly link of the enterprise.
[0125] After that, the first matching layer, guided by the identified business scenario, searches for matching items in the knowledge graph pre-constructed by the enterprise. This knowledge graph covers the detailed information of all suppliers the enterprise has cooperated with in the past and various industry-standard supplier naming specifications, etc.
[0126] Suppose during the search process, it is found that the name "XX Hardware Factory (formerly XX Hardware Products Factory)" should be "XX City XX Hardware Co., Ltd." according to the standard name corresponding to the current business scenario in the knowledge graph. Then the name update layer will correct this name based on the knowledge graph.
[0127] After the correction is completed, the second matching layer will re-check other data (location, contact information, main products, etc.) of the supplier based on the knowledge graph. If it is found that one digit is missing from the contact information, the correction strategy generation layer of the model will generate a corresponding correction strategy according to this error type. For example, it will prompt to supplement the complete and correct phone number digits. After the update layer executes this correction strategy, the corrected result will be finally output through the output layer. The output corrected result will be displayed on the system interface and presented as an updated and correct supplier directory, as shown in Table 2. After the above supplier record is corrected, it becomes:
[0128] Table 2
[0129]
[0130] In this way, the enterprise has completed the processing of the data of this batch of suppliers to be corrected by using the supplier name correction model, ensuring the accuracy of the supplier data and laying a good data foundation for the subsequent smooth development of business activities such as procurement. You can adjust and improve the relevant content (such as enterprise type, supplier data details, etc.) in this example according to actual needs to make it more suitable for specific application scenarios.
[0131] Figure 2 It is a schematic diagram of the main components of the supplier name correction system in the embodiment of the present invention. As Figure 2 shown, the supplier name correction system 1 provided by the embodiment of the present invention includes a data division module 10, a knowledge graph construction module 20, a model construction module 30, and a correction result acquisition module 40.
[0132] The data division module 10 is used to collect data from multiple data sources, parse the data to obtain the supplier identifier corresponding to each data, and determine the supplier data corresponding to each supplier s i based on the supplier identifier;
[0133] The knowledge graph construction module 20 is used to construct a knowledge graph for each supplier s i based on the supplier data. The knowledge graph takes the supplier s i as the center and radiates multiple supplier-related entities outward. The supplier-related entities include a supplier name entity, a supplier basic information entity, a supplier business association entity, a supplier qualification and certification entity, a supplier finance entity, a supplier credit entity, and a supplier dynamic event entity. The supplier name entity includes the supplier name s i corresponding to the supplier s y in different business scenarios;
[0134] The model construction module 30 is used to construct a supplier name correction model, and the supplier name correction model includes an input layer, an extraction layer, a business scenario recognition layer, a first matching layer, a name update layer, a second matching layer, a correction strategy generation layer, an update layer, and an output layer; the input layer receives supplier data, the extraction layer extracts the supplier name data from the supplier data, the business scenario recognition layer determines the business scenario corresponding to each supplier name, the first matching layer searches for matching items in the knowledge graph guided by the business scenario, if no matching item is found, the name update layer corrects the supplier name based on the knowledge graph, after the correction is completed, the second matching layer locates the error type of the supplier data based on the knowledge graph, the correction strategy generation layer generates a correction strategy based on the error type, the update layer executes the correction strategy, and the output layer outputs the correction result;
[0135] The correction result acquisition module 40 is used to obtain the supplier data to be corrected, and input the supplier data to be corrected into the supplier name correction model to obtain the correction result.
[0136] Figure 3 The following is a schematic structural diagram of the electronic device entity provided by the embodiment of the present invention, as Figure 3 shown, the electronic device 50 includes: a processor 501 (processor), a memory 502 (memory), and a bus 503;
[0137] Among them, the processor 501 and the memory 502 communicate with each other through the bus 503;
[0138] The processor 501 is used to call the program instructions in the memory 502 to execute the methods provided by the above method embodiments to execute the methods provided by the embodiments of the present invention.
[0139] This embodiment provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the methods provided by the embodiments of the present invention.
[0140] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various storage media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0141] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for modifying a supplier name, characterized in that: include: Collect data from multiple data sources, parse the data, obtain the supplier ID corresponding to each data, and determine each supplier based on the supplier ID. i Corresponding supplier data; Based on the supplier data for each supplier i Build a knowledge graph based on supplier i As the center, it radiates outwards a variety of supplier-related entities, including supplier name entity, supplier basic information entity, supplier business association entity, supplier qualification and certification entity, supplier financial entity, supplier credit entity and supplier dynamic event entity. The supplier name entity includes supplier s i Supplier names corresponding to different business scenarios y ; Constructing a supplier name correction model, the supplier name correction model comprising an input layer, an extraction layer, a business scenario identification layer, a first matching layer, a name update layer, a second matching layer, a correction strategy generation layer, an update layer and an output layer; Receive supplier data through the input layer, extract supplier name data from the supplier data through the extraction layer, determine the business scenario corresponding to each supplier name through the business scenario identification layer, search for matching items in the knowledge graph with the business scenario as a guide through the first matching layer, and if no matching items are found, correct the supplier name based on the knowledge graph through the name update layer. After the correction is completed, locate the error type of the supplier data based on the knowledge graph through the second matching layer, generate a correction strategy based on the error type through the correction strategy generation layer, execute the correction strategy through the update layer, and output the correction result through the output layer; The supplier data to be revised is obtained, and the supplier data to be revised is input into the supplier name revision model to obtain a revision result.
2. The method for modifying a supplier name according to claim 1, characterized in that: The supplier name correction method further includes: Collect the latest data, parse the latest data, obtain the supplier ID corresponding to each latest data, and determine each supplier based on the supplier ID. i Corresponding latest supplier data; The corresponding knowledge graph is updated based on the latest supplier data.
3. The method for modifying a supplier name according to claim 1, characterized in that: Determining the business scenario corresponding to each supplier name through the business scenario identification layer includes: Calculate the business scenario matching value through formula (1); M(s y , sc j ) = W(s y , sc j ) × F cos (V(s y ), V(sc j )) + B(s y , sc j ) Equation (1); In the formula, M(s y , sc j ) is the supplier name y With business scenario sc j The matching value between them, W is the weight function, F cos is the cosine similarity function, V(s y ) is the supplier name y The vector obtained after feature extraction and vectorization of the corresponding data, V(sc j ) is the business scenario sc j The vector obtained after feature extraction and vectorization of the data, B is the bias term; Choose to make M(s y , sc j )The highest business scenario sc j As supplier name y Corresponding business scenarios.
4. The method for modifying a supplier name according to claim 3, characterized in that: The searching for matching items in the knowledge graph using the first matching layer and guided by the business scenario includes: Calculate the business scenario sc by formula (2) j With entity e j The matching value between In the formula, F(sc j , e j ) is the business scenario sc j With entity e j The matching value between them, KG is the knowledge graph, r is the relationship in the knowledge graph KG, ω r is the weight associated with relation r, I(sc j , e j , r) is the indicator function; When I(sc j , e j , r) is 1, the business scenario sc j Entity e j There is a connection between I(sc j , e j , r) is 0, the business scenario sc j Entity e j There is no connection between r and relation r; When F(sc j , e j ) is greater than or equal to the preset matching value threshold, the business scenario sc is determined j With entity e j Match, when F(sc j , e j ) is less than the preset matching value threshold, the business scenario sc is determined j With entity e j No match.
5. The method for modifying a supplier name according to claim 1, characterized in that: The locating the error type of the supplier data based on the knowledge graph through the second matching layer includes: The comprehensive error score of supplier data T is calculated by formula (3); Where S(T) is the comprehensive error score of supplier data T, l is the total number of supplier data error types, j is the index variable, and α j For error type et j The weight of S(et j , T) is the jth supplier data error type; When S(T) is greater than or equal to the preset error score threshold, it is determined that there is an error in the supplier data, and each error type is analyzed. j The proportion of the score to the total score S(T) is used to determine the main error type, and the error type of the supplier data is located based on the main error type.
6. The method for modifying a supplier name according to claim 1, characterized in that: The supplier name correction method further includes: Comparing the correction result output by the supplier name correction model with the actual correction data to obtain an error value of the correction result; Determine whether the error value exceeds a preset error threshold. If so, determine the parameter to be adjusted based on the error value, and adjust the parameter to be adjusted until the error value of the text supplier name correction model does not exceed the preset error threshold, so as to obtain the optimal supplier name correction model.
7. The method for modifying a supplier name according to claim 6, characterized in that: The step of determining the parameter to be adjusted based on the error value and adjusting the parameter to be adjusted until the error value of the text supplier name correction model does not exceed a preset error threshold to obtain an optimal supplier name correction model includes: Dividing the parameters of the supplier name correction model into a parameter group of input layer parameters, extraction layer parameters, business scenario identification layer parameters, first matching layer parameters, name update layer parameters, second matching layer parameters, correction strategy generation layer parameters, update layer parameters and output layer parameters; Conduct a small-scale disturbance experiment on each parameter group one by one, determine the sensitive parameters that are sensitive to the error value based on the error value between the correction result output by the supplier name correction model and the actual correction data after each disturbance, and determine the sensitive parameters as the parameters to be adjusted; Calculate a new parameter value corresponding to the parameter to be adjusted, adjust the parameter to be adjusted based on the new parameter value, obtain a new correction result output by the supplier name correction model after adjustment, and calculate a new error value based on the correction result; The new error value is compared with the preset error threshold. If the new error value does not exceed the preset error threshold after multiple consecutive iterations, the parameter adjustment is terminated to obtain the optimal supplier name correction model.
8. A system for modifying supplier names, characterized in that: include: The data partitioning module is used to collect data from multiple data sources, parse the data, obtain the supplier identification corresponding to each data, and determine each supplier based on the supplier identification. i Corresponding supplier data; A knowledge graph building module is used to construct a knowledge graph for each supplier based on the supplier data. i Build a knowledge graph based on supplier i As the center, it radiates outwards a variety of supplier-related entities, including supplier name entity, supplier basic information entity, supplier business association entity, supplier qualification and certification entity, supplier financial entity, supplier credit entity and supplier dynamic event entity. The supplier name entity includes supplier s i Supplier names corresponding to different business scenarios y ; A model building module is used to build a supplier name correction model, the supplier name correction model includes an input layer, an extraction layer, a business scenario identification layer, a first matching layer, a name update layer, a second matching layer, a correction strategy generation layer, an update layer and an output layer; the input layer receives supplier data, the extraction layer extracts supplier name data from the supplier data, the business scenario identification layer determines the business scenario corresponding to each supplier name, the first matching layer searches for matching items in the knowledge graph with the business scenario as a guide, if no matching items are found, the name update layer corrects the supplier name based on the knowledge graph, after the correction is completed, the second matching layer locates the error type of the supplier data based on the knowledge graph, the correction strategy generation layer generates a correction strategy based on the error type, the update layer executes the correction strategy, and the output layer outputs the correction result; The correction result acquisition module is used to obtain the supplier data to be corrected, and input the supplier data to be corrected into the supplier name correction model to obtain the correction result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Entity disambiguation method and device, computer device and computer storage medium
CN109635297A
Knowledge graph engineering construction method and device, computer equipment and storage medium
CN112650855A
Knowledge graph ontology evaluation method, device and equipment and storage medium
CN113946692A
Supplier knowledge graph-based surrounding bidding prediction method
CN115630169A
Artificial intelligence-based entity identification method, apparatus and device, and medium
CN116992879A