Methods for updating risk information of certified suppliers
By employing multi-source cross-validation and intelligent inference, supplier information is obtained from multiple data sources, resolving the data bias issue on commercial data platforms and improving the accuracy and reliability of supplier risk information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA GENERAL CERTIFICATION CENT
- Filing Date
- 2026-02-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing business data platforms suffer from data bias, leading to inaccurate supplier registration information and affecting the accuracy of supplier qualification verification.
The defense system employs multi-source cross-validation and intelligent reasoning to obtain supplier data from multiple independent data sources. By constructing retrieval and language models, it extracts related information, performs logical reasoning and comparison, and generates more accurate supplier risk information.
It effectively identifies and corrects data deviations, improves the accuracy and reliability of supplier risk information, avoids being misled by a single data source, and enhances the precision of qualification review.
Smart Images

Figure CN121684968B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more particularly to a method for updating risk information of certified suppliers. Background Technology
[0002] As the industry continues its steady development, more and more companies are recognizing the importance of various certifications and applying for them. Consequently, the workload of third-party certification bodies is also increasing. Certification bodies need to review the information of certified suppliers, but current mainstream business data platforms suffer from data bias, resulting in inaccurate supplier registration information. Summary of the Invention
[0003] At least one aspect and advantage of the invention will be set forth in part in the description which follows, or may be apparent from the description, or may be obtained by practicing the subject matter of this disclosure.
[0004] According to a first aspect of the present invention, a method for updating risk information of certified suppliers includes:
[0005] First and second data associated with the supplier to be audited and certified are obtained based on the first and second data sources, respectively.
[0006] A search query is constructed based on the first and second data to retrieve historical data and related documents;
[0007] A first-dimensional information sequence is obtained based on the first data and the second data, and a second-dimensional information sequence is obtained from historical data and related documents based on the first-dimensional information sequence.
[0008] Reasoning is performed based on the first-dimensional information sequence and the second-dimensional information sequence to obtain the score of the dimension associated with the first data and the second data;
[0009] The scores of the first and second data are combined based on the dimensions associated with the first and second data to obtain the credibility of the third data and its association.
[0010] In response to the overall credibility exceeding the first threshold, the supplier information is updated based on third-party data;
[0011] Historical data is obtained from a third data source based on a retrieval formula constructed from the first dimension information sequence, and related documents are obtained from a fourth data source based on a retrieval formula constructed from the first dimension information sequence. The historical data and data in the related documents are extracted based on entities associated with the dimension information.
[0012] The process of obtaining the second-dimensional information sequence includes:
[0013] Based on the dimensional information corresponding to the first dimensional information sequence, slices are obtained from historical data or related documents. A language model is used to extract related information from the slices in the slice sequence, and a second dimensional information sequence is generated based on the related information.
[0014] The overall credibility is calculated based on the credibility of third-party data association.
[0015] According to one embodiment of the present invention, the process of constructing the search query includes:
[0016] Obtain the name of the first data contained in the data based on the first data or the second data;
[0017] Determine the second data name sequence associated with the information to be updated;
[0018] The search query is constructed based on preset query suggestions, second data names, first data and / or second data.
[0019] According to one embodiment of the present invention, the fourth data source is documents submitted by the supplier to be audited and certified, publicly available information, web search APIs, or information acquisition agents.
[0020] According to one embodiment of the present invention, the process of obtaining the slice sequence includes:
[0021] Retrieve the name of the dimension information based on the dimension information;
[0022] Document structure is obtained based on document structure analysis of historical data and related documents;
[0023] The slice labels used for slices of a document are determined based on the names of the slices in the document, which are based on the document structure and dimensional information.
[0024] Use slice tags to slice historical data and related documents to obtain slice sequences.
[0025] According to an embodiment of the present invention, the process of obtaining the associated information includes:
[0026] The second prompt word is obtained based on the first prompt word, the sequence value, and the first major language model;
[0027] Based on the second cue word and the slice sequence, the associated information is obtained.
[0028] According to one embodiment of the present invention, the score of the dimension associated with the first data and the second data is obtained based on the following method:
[0029] Obtain the sequence of dimension information names based on the first dimension information sequence;
[0030] Iterate through the dimension information names in the sequence of dimension information names, retrieve the first dimension data value associated with the first data, retrieve the second dimension data value associated with the second data, and retrieve the third dimension data value associated with the second dimension information sequence.
[0031] In response to the fact that the intersection of the third dimension data value, the first dimension data value, and the second dimension data value is empty, the score of the dimension associated with the first and second dimension data is determined based on the first dimension data value and the second dimension data value;
[0032] In response to the fact that the intersection of the third dimension data value with the first dimension data value and the second dimension data value is not simultaneously empty, the master data is determined based on the size of the intersection of the first dimension data value, the second dimension data value and the third dimension data value, and the score of the dimension associated with the first data and the second data is determined based on the master data.
[0033] According to one embodiment of the present invention, the update process of the third data and credibility includes:
[0034] Traverse the first dimension sequence, obtain the master data and slave data for each dimension based on the scores of the dimensions associated with the first and second data, use the score of the master data as the confidence level, and merge the slave data and master data in response to the slave data score being greater than the second threshold.
[0035] According to one embodiment of the present invention, the overall credibility is obtained by clustering the credibility of the third data association.
[0036] The beneficial effects of this invention are as follows: This invention adopts a defense system of multi-source cross-validation and intelligent reasoning, does not blindly trust any single data source, especially commercial data platforms that may have mutual crawling and pollution, actively seeks evidence from more original and diverse information sources, and effectively discovers and corrects data deviation problems through logical reasoning and comparison, greatly improving the accuracy and reliability of supplier risk information. Attached Figure Description
[0037] Figure 1 This is a flowchart of the method for updating risk information of certified suppliers in an embodiment of the present invention. Detailed Implementation
[0038] The present disclosure will now be discussed with reference to several exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and thus implement the present disclosure, and are not intended to imply any limitation on the scope of the disclosure.
[0039] like Figure 1 As shown in the figure, this embodiment introduces a method for updating risk information of certified suppliers, including steps 1100-1600.
[0040] Step 1100: Obtain the first data and the second data associated with the supplier to be audited and certified, respectively, based on the first data source and the second data source.
[0041] The system does not rely on a single data source; it retrieves data from the same supplier from two independent data sources. The first data source retrieves data on the supplier awaiting verification and certification, while the second data source retrieves data on the same supplier from the second data source. This data includes basic supplier information such as company name, address, and business scope.
[0042] Step 1200: Construct a search query based on the first and second data to obtain historical data and related documents.
[0043] Step 1300: Obtain the first dimension information sequence based on the first data and the second data, and obtain the second dimension information sequence from historical data and related documents based on the first dimension information sequence.
[0044] Historical data is obtained from a third data source based on a retrieval formula constructed from the first dimension information sequence, and related documents are obtained from a fourth data source based on a retrieval formula constructed from the first dimension information sequence. The data in the historical data and related documents are extracted based on entities associated with the dimension information.
[0045] The process of obtaining the second-dimensional information sequence includes:
[0046] Based on the dimensional information corresponding to the first dimensional information sequence, slices are obtained from historical data or related documents. A language model is used to extract related information from the slices in the slice sequence, and a second dimensional information sequence is generated based on the related information.
[0047] The third data source can be an internal database of supplier cooperation history, providing information over time. The fourth data source can be bidding websites, etc., used to obtain unstructured documents such as official documents, reports, and news related to the supplier.
[0048] The first dimension information sequence refers to the dimensions that the system explicitly needs to audit, such as equity structure, key management personnel, intellectual property, and external investments. These dimensions are organized into a sequence, namely the first dimension information sequence.
[0049] The system searches and segments large blocks of text, including historical data and related documents, based on the first dimension of information sequence. For example, for the equity structure dimension, the system extracts all paragraphs mentioning shareholders and equity from a company's annual report.
[0050] Then, these slices are read using a language model. The advantage of a language model lies in its ability to understand natural language and accurately extract key information from a text. For example, it can extract "the company obtained an invention patent authorization in 2023" from an announcement text and categorize it under the intellectual property dimension. All the information extracted from the original document constitutes the second-dimensional information sequence. The second-dimensional information sequence comes from more diverse information sources, effectively supplementing and verifying the shortcomings of the first and second data sources.
[0051] Step 1400: Based on the first dimension information sequence and the second dimension information sequence, perform inference to obtain the score of the dimension associated with the first data and the second data.
[0052] The reasoning process can include consistency checks, such as determining whether the registered capital in the first and second data sets is consistent. The reasoning process can also include assessing the strength of evidence, such as determining whether the related information in the second-dimensional information sequence supports the first or second data set, or neither. Furthermore, it can be based on historical data of the second-dimensional information sequence to determine whether the information in a certain dimension of the first or second data set is reasonable.
[0053] Ultimately, the system will give a credibility score for the first and second data points in each dimension.
[0054] Step 1500: Combine the first and second data based on the scores of the dimensions associated with the first and second data to obtain the credibility of the third data and its association.
[0055] For each dimension, the system determines whether to use the first or second set of data based on the score. Ultimately, the system generates a completely new third set of data. Each dimension in this third set of data is associated with a credibility level.
[0056] Step 1600: In response to the overall credibility being greater than the first threshold, update the supplier information based on third data.
[0057] The overall credibility is calculated based on the credibility of third-party data association. For example, the overall credibility can be calculated by weighted averaging of the credibility of all dimensions, or the lowest value can be directly used as the overall credibility.
[0058] The system will only update the supplier database with the merged third-party data when the overall credibility of a supplier exceeds a preset security threshold.
[0059] This embodiment employs a defense system of multi-source cross-validation and intelligent reasoning. It does not blindly trust any single data source, especially commercial data platforms that may be subject to mutual crawling and contamination. It proactively seeks evidence from more original and diverse information sources, and through logical reasoning and comparison, it can effectively discover and correct data deviations, greatly improving the accuracy and reliability of supplier risk information.
[0060] According to one embodiment of the present invention, the process of constructing the search query includes: obtaining a first data name contained in the data based on the first data or the second data; determining a second data name sequence associated with the information to be updated; and constructing a search query based on preset query prompts, the second data name, the first data and / or the second data.
[0061] The primary data name is the entity name that uniquely identifies the supplier. This could be the supplier's official full name or its unified social credit code. This serves as the anchor for all subsequent search operations. Even if other information on the commercial data platform is incorrect, the entity name is correct in the vast majority of cases; otherwise, it would be impossible to find the supplier.
[0062] The second data name is the specific information dimension or attribute that needs to be verified or filled. These names constitute a list to be checked, such as "foreign investment", "change of senior management", "bidding records", etc.
[0063] The second data name sequence can be determined based on different data sources. For example, for a certain supplier, if the information on "outward investment" from the first data source is inconsistent with the information on the same dimension from the second data source, then this dimension "outward investment" will be added to the second data name sequence.
[0064] The second data name sequence can also include dimensions that the system pre-sets as requiring special attention. For example, for high-tech suppliers, information in the "intellectual property" dimension requires special attention.
[0065] In addition, if a key dimension is missing in both data sources, it will also be added to the second data name sequence.
[0066] Preset query suggestions serve as task instruction templates to guide the language model in generating retrieval-oriented queries. For example, inputting "You are a professional business analyst. Please generate precise web search terms for company {company name} regarding {query dimensions} to find relevant official announcements or authoritative news reports" into the language model.
[0067] The system sends the query suggestion, the second data name, the first data and / or the second data combination to the language model, and the language model will output a highly customized search query.
[0068] In this embodiment, the system does not directly trust the first and second data sources. Instead, it uses this list of questionable sources to search for evidence from other, more authoritative data sources, thus avoiding the cyclical pollution caused by the mutual scraping between the first and second data sources. Through the generated intelligent search queries, the system no longer searches aimlessly but targets precisely, directly seeking original evidence that can confirm or refute a specific fact. For information missing from the data sources, the system can proactively generate search queries to discover it. For example, if the data sources do not contain information on "outward investment," the system will proactively generate search queries related to "outward investment" to investigate risks.
[0069] According to one embodiment of the present invention, the fourth data source is documents submitted by the supplier to be audited and certified, publicly available information, web search APIs, or information acquisition agents.
[0070] According to one embodiment of the present invention, the process of obtaining the slice sequence includes: obtaining the name of the dimension information based on the dimension information; obtaining the document structure based on the document structure analysis of historical data and related documents; determining the slice tags used for slicing the document based on the document structure and the name of the dimension information; and slicing the historical data and related documents using the slice tags to obtain the slice sequence.
[0071] The system extracts the specific name of the dimension to be processed from the first dimension information sequence. The dimension information name is the item to be searched and verified, such as bidding records, foreign investment, etc.
[0072] For a regular document, the document structure includes a table of contents, chapters, headings, paragraphs, etc. For a web page document, the document structure can be specific tags, CSS class names, etc.
[0073] The original document is often complex, such as a 100-page corporate annual report. Directly feeding the original document to a language model is inefficient and prone to missing key information. Parsing the document structure is like obtaining a map of the document, clearly showing the location of information in each dimension within the document.
[0074] Based on the dimension name, the system infers identifiers for characteristic chapters or regions within the document structure that may contain relevant information; these identifiers are called slice tags. Slice tags are keywords used to locate content within the document structure. For example, for the "outward investment" dimension, the slice tag might be the title of the "investment projects" chapter within the document.
[0075] The system uses the slice tags determined in the previous step to search and match within the parsed document structure. Once a structural unit corresponding to a tag is found, such as a chapter or a table, that part of the content is extracted completely to form a slice. The collection of all slices is the slice sequence.
[0076] For example, for a 100-page PDF corporate annual report, output slice 1 and slice 2. Slice 1 is the three paragraphs under the "Significant Matters" section about an external investment, and slice 2 is the table in the "Notes to the Financial Statements" section that lists all subsidiaries and their shareholding percentages.
[0077] This embodiment slices historical data and related documents, feeding the subsequent language model not the entire document, but several highly relevant and essential segments. This reduces the model's processing burden, lowers the possibility of interference from irrelevant information, and significantly improves the accuracy of information extraction. Unlike simple keyword matching, this structure-based slicing preserves complete sentences, paragraphs, and logical relationships, facilitating language model comprehension.
[0078] According to an embodiment of the present invention, the process of obtaining the associated information includes: obtaining a second prompt word based on a first prompt word, a sequence value, and a first large language model; and obtaining associated information based on the second prompt word and the slice sequence.
[0079] The first prompt is pre-configured; it's a generic task instruction template. The sequence value is the specific dimension name and the specific data in question obtained from the initial data source. For example, the dimension name could be "outward investment," and the corresponding specific data could be a list of invested companies.
[0080] The system combines the first prompt word and the sequence value and sends it to the primary language model. This primary language model's task is not to directly extract information, but rather to generate a more targeted question, i.e., the second prompt word. This approach incorporates questionable data as potential suspects into the question, dynamically creating a highly customized query tailored to the specific dimension and data, thus solving the problem of generic first prompt words being insufficiently accurate.
[0081] The system combines the optimized second cue word with the previously obtained slice sequence and sends this combination to the large language model. The large language model's task is to strictly follow the instructions of the second cue word, read the text in the slice sequence, and output the final association information. This output is no longer the original natural language paragraph, but rather organized, summarized, and structured information, such as a list or a clear summary text.
[0082] According to one embodiment of the present invention, the score of the dimension associated with the first data and the second data is obtained based on the following method:
[0083] Obtain the sequence of dimension information names based on the first dimension information sequence;
[0084] Iterate through the dimension information names in the sequence of dimension information names, retrieve the first dimension data value associated with the first data, retrieve the second dimension data value associated with the second data, and retrieve the third dimension data value associated with the second dimension information sequence.
[0085] In response to the fact that the intersection of the third dimension data value, the first dimension data value, and the second dimension data value is empty, the score of the dimension associated with the first and second dimension data is determined based on the first dimension data value and the second dimension data value;
[0086] In response to the fact that the intersection of the third dimension data value with the first dimension data value and the second dimension data value is not simultaneously empty, the master data is determined based on the size of the intersection of the first dimension data value, the second dimension data value and the third dimension data value, and the score of the dimension associated with the first data and the second data is determined based on the master data.
[0087] For each dimension requiring scoring, the system retrieves three data values corresponding to that dimension: a first-dimensional data value obtained from the first data set, a second-dimensional data value obtained from the second data set, and a third-dimensional data value obtained from the second-dimensional information sequence. The first and second-dimensional data values are obtained directly from the original data source, while the third-dimensional data value is extracted as evidence from historical data and related documents.
[0088] Then, it is determined whether the third-dimensional data value, which serves as external evidence, is valid, that is, whether it has anything in common with any initial data source.
[0089] If the intersection of the third-dimensional data value with the first and second-dimensional data values is empty, it indicates that the external evidence may point to a completely different aspect, or that the current dimension is not applicable to the evidence. Therefore, the third-dimensional data value is invalid in this round of scoring. In this case, the system compares the two initial data sources. If the first-dimensional and second-dimensional data values are consistent, a higher score is obtained; otherwise, a lower score is obtained. In this way, when external evidence is lacking, the system assigns a lower score, avoiding blindly trusting a particular data source.
[0090] For example, because the first and second data sources crawled each other, they mistakenly recorded a non-existent change record in the "Executive Changes" dimension. Meanwhile, the evidence extracted by the system from the official announcement (i.e., the third dimension data value) is empty, indicating no change. At this point, the intersection of the third dimension data value and the two initial data values is empty. Although the two unreliable data sources are consistent, they lack supporting external evidence, so they will both receive low scores in this dimension. The system thus determines that this data is unreliable.
[0091] If the intersection of the third-dimensional data value and the first-dimensional and second-dimensional data values is not simultaneously empty, that is, if the third-dimensional data value has a common item with one of the initial data values, then the third-dimensional data value can be used as a scoring basis.
[0092] The master data is the intersection of the three dimensions of data values. Data with a higher degree of overlap in the master data has a higher score.
[0093] For example, in the "outward investment" dimension, the first dimension data value records [A, B], indicating that the company has invested in companies A and B. The second dimension data value records [A, C], indicating that the company has invested in companies A and C. The evidence extracted from the annual report (the third dimension data value) records [A, B, D], indicating that the company has invested in companies A, B, and D. In this case, the master data is [A]. Both the first and second dimension data values record the master data. Because the first dimension data value also includes data supported by the third dimension data value, its score is higher in the "outward investment" dimension.
[0094] According to one embodiment of the present invention, the update process of the third data and credibility includes:
[0095] Traverse the first dimension sequence, obtain the master data and slave data for each dimension based on the scores of the dimensions associated with the first and second data, use the score of the master data as the confidence level, and merge the slave data and master data in response to the slave data score being greater than the second threshold.
[0096] For the dimension being processed, the system determines which data source provides more reliable information based on the dimension's score. The master data is the data value provided by the data source with the higher score on that dimension. For example, on the "outward investment" dimension, if the first data has a score of 0.9 and the second data has a score of 0.5, then the first data is determined to be the master data, and the second data is determined to be the slave data.
[0097] The scores of the master data in this dimension are then used as the credibility of the information in this dimension in the third data. The generated third data is a set of information with credibility.
[0098] The second threshold is a higher standard. If the data score is greater than the second threshold, it indicates that the data source also provides good quality information in this dimension. The system will perform a deduplication and merging operation, adding information present in the data source but not in the main data to the final third data. If the data score is less than the second threshold, the system considers the data unreliable, and its additional information may be noise or errors, therefore it is discarded.
[0099] This embodiment fundamentally avoids using erroneous information from low-scoring data sources by adopting master data for each dimension. Through intelligent fusion with high thresholds, the system can supplement correct information missed by mainstream data sources but contained in another high-quality data source. Assigning credibility to each data point makes the entire system's decision-making process no longer a black box. Business personnel can trust high-credibility data and remain vigilant or manually intervene in low-credibility data, which greatly improves the sophistication of supplier risk management.
[0100] According to one embodiment of the present invention, the overall credibility is obtained by clustering the credibility of the third data association.
[0101] While specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of the invention.
[0102] Those skilled in the art will recognize that the modules and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and equipment can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0104] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0105] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.
[0106] In addition, the functional modules in the embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0107] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0108] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
[0109] It should be understood that the sequence numbers of the steps in the invention's content and embodiments do not absolutely imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. The foregoing description of embodiments of this disclosure has been provided for illustrative and descriptive purposes. The foregoing description is not exhaustive and is not intended to limit this disclosure to the exact form disclosed. Various modifications and variations may exist based on the foregoing teachings, or various modifications and variations may be derived from the practice of this disclosure. These embodiments were chosen and described to illustrate the principles of this disclosure and its practical application, so that those skilled in the art can utilize this disclosure in various implementations and modifications suitable for the specific purpose of the concept.
Claims
1. A method for updating risk information of certified suppliers, characterized in that, include: First data and second data associated with the supplier to be audited and certified are obtained based on the first data source and the second data source, respectively. The first data and second data include company name, address and business scope. A search query is constructed based on the first and second data to retrieve historical data and related documents; The first dimension information sequence is obtained based on the first data and the second data, and the second dimension information sequence is obtained from historical data and related documents based on the first dimension information sequence. The first dimension information sequence includes equity structure, key management personnel, intellectual property rights, and external investments. Reasoning is performed based on the first-dimensional information sequence and the second-dimensional information sequence to obtain the score of the dimension associated with the first data and the second data; Based on the rating, the first and second data are combined to obtain the third data and its associated credibility. In response to the overall credibility exceeding the first threshold, the supplier information is updated based on third-party data; Historical data is obtained from a third data source based on a retrieval formula constructed from the first dimension information sequence, and related documents are obtained from a fourth data source based on a retrieval formula constructed from the first dimension information sequence. The historical data and data in the related documents are extracted based on entities associated with the dimension information. The process of obtaining the second-dimensional information sequence includes: Based on the dimensional information corresponding to the first dimensional information sequence, slices are obtained from historical data or related documents. A language model is used to extract related information from the slices in the slice sequence, and a second dimensional information sequence is generated based on the related information. The overall credibility is calculated based on the credibility of the third-party data association; The scores for the dimensions relating the first and second data are obtained based on the following method: Obtain the sequence of dimension information names based on the first dimension information sequence; Iterate through the dimension information names in the sequence of dimension information names, retrieve the first dimension data value associated with the first data, retrieve the second dimension data value associated with the second data, and retrieve the third dimension data value associated with the second dimension information sequence. In response to the fact that the intersection of the third dimension data value and the first dimension data value and the second dimension data value is empty, the score is determined based on the first dimension data value and the second dimension data value; In response to the fact that the intersection of the third dimension data value and the first dimension data value and the second dimension data value is not simultaneously empty, the master data is determined based on the size of the intersection of the first dimension data value, the second dimension data value and the third dimension data value, and the score is determined based on the master data.
2. The method for updating risk information of certified suppliers as described in claim 1, characterized in that, The process of constructing the search query includes: Obtain the name of the first data contained in the data based on the first data or the second data; Determine the second data name sequence associated with the information to be updated; The search query is constructed based on preset query suggestions, second data names, first data and / or second data.
3. The method for updating risk information of certified suppliers as described in claim 1, characterized in that, The fourth data source is the documents, public information, web search APIs, or information acquisition agents submitted by the suppliers to be audited and certified.
4. The method for updating risk information of certified suppliers as described in claim 1, characterized in that, The process of obtaining the slice sequence includes: Retrieve the name of the dimension information based on the dimension information; Document structure is obtained based on document structure analysis of historical data and related documents; The slice labels used for slices of a document are determined based on the names of the slices in the document, which are based on the document structure and dimensional information. Use slice tags to slice historical data and related documents to obtain slice sequences.
5. The method for updating risk information of certified suppliers as described in claim 1, characterized in that, The process of obtaining the associated information includes: The second prompt word is obtained based on the first prompt word, the sequence value, and the first major language model; Based on the second cue word and the slice sequence, the associated information is obtained.
6. The method for updating risk information of certified suppliers as described in claim 1, characterized in that, The process of updating the third data and credibility includes: Traverse the first dimension sequence, obtain the master data and slave data for each dimension based on the scores of the dimensions associated with the first and second data, use the score of the master data as the confidence level, and merge the slave data and master data in response to the slave data score being greater than the second threshold.
7. The method for updating risk information of certified suppliers as described in claim 1, characterized in that, The overall credibility is obtained by clustering the credibility of the third data association.