Self-learning data acquisition system for supplier evaluation
By leveraging the multi-level collaborative operation of the self-learning data acquisition system, the problem of inaccurate supply data processing was solved, the accuracy and reliability of supplier evaluation were improved, the pass rate of data processing was increased, and the objectivity of the company's procurement plan was ensured.
Patent Information
- Application Number
- CN202510332058.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Existing technologies cannot perform targeted verification and processing of the collected supply data, resulting in the inability to obtain supply data samples that match the actual situation, affecting the accuracy of supplier evaluation and the objective guidance of enterprise production plans.
A self-learning data acquisition system was designed, comprising an acquisition layer, an instruction layer, an acquisition layer, a processing layer, a cleaning layer, an analysis layer, and a control layer. Through the collaborative work of these layers, the system acquires, compresses, extracts, cleans, and integrates the supply data to generate complete data. Based on the data ratio and the distribution of abnormal data, the system judges the qualification of the data processing process and adjusts the system operating parameters to improve the data processing capability.
It improved the pass rate and accuracy of supply data processing, enhanced the reliability and accuracy of supplier evaluation, and provided more objective guidance for enterprises' procurement plans.
Smart Images

Figure CN120316165B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of supplier data management technology, and in particular to a self-learning data acquisition system for supplier evaluation. Background Technology
[0002] With the rapid development of computer technology, most companies now use computer-collected supply data to evaluate suppliers as a whole. Through data collection, integration, and analysis, various supply data of each supplier are summarized and evaluated to determine the individual supplier's assessment score within the company, providing guidance and reference for the company's subsequent procurement plans.
[0003] However, due to flaws in the data processing during the collection and analysis of supply data generated by suppliers, existing technologies are unable to perform targeted verification and generate corresponding processing methods for the collected supply data. This results in the inability to obtain supply data samples that match the actual situation, thereby affecting the final evaluation results of suppliers and failing to provide objective guidance for the company's production plan. Summary of the Invention
[0004] To address this, the present invention provides a self-learning data acquisition system for supplier evaluation, which solves the problem that the existing technology cannot perform targeted verification and processing on the collected supply data, thus failing to obtain supply data samples that match the actual situation and accurately evaluate suppliers based on the sample data.
[0005] To achieve the above objectives, the present invention provides a self-learning data acquisition system for supplier evaluation, comprising:
[0006] The data acquisition layer includes several data acquisition terminals, each of which is connected to the server of the corresponding supplier and acquires the supply data generated by the supplier within a preset period.
[0007] The instruction layer, which is connected to the acquisition layer, is used to send acquisition instructions to the acquisition layer so that each acquisition terminal can obtain the corresponding supply data.
[0008] An acquisition layer, which is connected to the acquisition layer, is used to acquire and compress several supply data output by each acquisition terminal.
[0009] A processing layer, which is connected to the acquisition layer, is used to receive and decompress each of the supply data, and to extract tags from each supply data to obtain several strings, wherein the categories of the strings include product information, logistics information, price information and time nodes.
[0010] A cleaning layer, which is connected to the processing layer, is used to clean each of the strings to filter out abnormal strings, wherein abnormal strings include garbled strings and discrete strings;
[0011] An analysis layer, which is connected to the cleaning layer and the acquisition layer respectively, is used to integrate and process the cleaned strings to obtain several complete data, and generate corresponding processing methods based on the data ratio, including: outputting complete data, or generating corresponding instructions based on the determined reasons for non-compliance, wherein the data ratio is the ratio between the total number of the complete data and the total number of the supplied data;
[0012] A control layer, which is connected to the acquisition layer, the cleaning layer, and the analysis layer respectively, is used to adjust the operating parameters of the corresponding layers based on the instructions. The parameters include: the bandwidth of the acquisition layer, the acquisition time interval of the acquisition layer, the compression ratio of the acquisition layer, and the cleaning standard range of the cleaning layer.
[0013] An evaluation layer, connected to the analysis layer, is used to receive several sets of complete data and evaluate the corresponding suppliers based on the complete data.
[0014] Furthermore, the analysis layer is also used to determine whether the data processing process is qualified based on the comparison result between the data ratio and the preset data ratio stored in the analysis layer, or based on the total number of collected data, and to determine the reason for the unqualified process based on the distribution of the abnormal data when the data processing process is determined to be unqualified, wherein the total number of collected data is the number of supply data collected by the collection layer within a preset period.
[0015] Furthermore, the analysis layer is also used to re-determine whether the data processing process is unqualified based on the comparison result between the total number of collected data and the preset total number of collected data stored in the analysis layer, or to increase the compression ratio of the acquisition layer based on the total difference, wherein the total difference is the difference between the total number of collected data and the preset total number of collected data.
[0016] Furthermore, the analysis layer is also used to increase the compression ratio of the acquisition layer based on the comparison result between the total difference and the preset total difference stored in the analysis layer, and the increase in compression ratio is proportional to the total difference.
[0017] Furthermore, the analysis layer is also used to determine the reasons for the data processing process being unqualified based on the comparison result between the average value and the preset average value of the analysis layer, and to generate corresponding correction methods based on the determined reasons, including: increasing the bandwidth of the acquisition layer, expanding the cleaning standard range of the cleaning layer, or re-determining the reasons for unqualification based on the string types in the abnormal data, wherein the time node of the abnormal data within the preset period is obtained, and the interval between adjacent abnormal data is obtained based on the time node, and the average value is calculated based on the interval.
[0018] Furthermore, the analysis layer is also used to increase the bandwidth of the acquisition layer based on the comparison result between the average traffic value and the preset average traffic value stored in the analysis layer, and the increase in bandwidth is inversely proportional to the average traffic value. In this case, the traffic value within the time interval of each abnormal data group is obtained, and the average traffic value is calculated based on each traffic value.
[0019] Furthermore, the analysis layer is also used to re-determine the reasons for the data processing failure based on the comparison result of the abnormal data ratio and the preset abnormal data ratio, and to generate corresponding correction methods based on the determined reasons, including: increasing the bandwidth of the acquisition layer, or performing market trend analysis on the abnormal data of a single type, wherein the abnormal data ratio is the ratio of abnormal data containing the garbled string to the total abnormal data.
[0020] Furthermore, the analysis layer is also used to re-determine the reasons for the data processing failure based on the comparison results between the average slope and the preset average slope stored in the analysis layer, and to generate corresponding correction methods based on the determined reasons, including: expanding the collection time interval of the collection layer, or increasing the bandwidth of the collection layer, wherein several historical price data within the corresponding historical period are obtained, and a curve is drawn based on the several historical price data and the current price data and the average slope is calculated.
[0021] Furthermore, the analysis layer is also used to expand the acquisition time interval of the acquisition layer based on the comparison result between the slope mean difference and the preset slope mean difference stored in the analysis layer, and the expansion of the acquisition time interval is inversely proportional to the slope mean difference, wherein the slope mean difference is the difference between the preset slope mean difference and the slope mean difference.
[0022] Furthermore, the analysis layer is also used to expand the cleaning standard range of the cleaning layer based on the comparison result between the data ratio difference and the preset data ratio difference stored in the analysis layer, and the expansion of the cleaning standard is proportional to the data ratio difference, wherein the data ratio difference is the difference between the preset data ratio and the data ratio.
[0023] Compared with existing technologies, the self-learning data acquisition system for supplier evaluation of the present invention has the following advantages: The system collects several supply data generated by each supplier through the acquisition layer, receives and compresses the supply data through the acquisition layer before transmitting it to the processing layer, extracts tags from each supply data to obtain several strings, identifies and removes abnormal data containing abnormal strings through the cleaning layer, and then re-integrates the supply data after removing abnormal data through the analysis layer to obtain complete data. The system also determines whether the data processing process is qualified based on the proportion of complete data in the supply data. If the process is deemed unqualified, the cause is identified, and a corresponding instruction is generated based on the cause and sent to the control layer. The control layer adjusts the operating parameters of the corresponding layer based on the instruction, thereby improving the subsequent data processing capability of the system, increasing the data processing qualification rate, and thus increasing the proportion of complete data. This provides more complete data samples for the subsequent evaluation process, thereby improving the accuracy and reliability of supplier evaluation.
[0024] Furthermore, the present invention can determine whether the data processing process is qualified by comparing the data ratio with the pre-stored preset data ratio, and when it is determined to be unqualified, it can also determine the reason for the unqualification based on the distribution of abnormal data.
[0025] Furthermore, in addition to the data ratio-based judgment, the present invention adds a second judgment based on the total number of collected data and the pre-stored preset total number of collected data to re-determine whether the data processing process is qualified, thereby improving the accuracy of the judgment on the data processing process.
[0026] Furthermore, the present invention also increases the compression ratio of the acquisition layer based on the total difference, so that when the amount of acquired supply data is too large, the data can be completely transmitted to the processing layer. The increase in the compression ratio of the acquisition layer is proportional to the total difference, and thus the sample size of the supply data can be increased according to the total difference, thereby improving the accuracy of the analysis results.
[0027] Furthermore, the present invention also determines the reasons for the unqualified quantity processing based on the comparison results between the average value and the pre-stored preset average value, and generates corresponding correction methods based on the determined reasons. By increasing the bandwidth of the acquisition layer or expanding the cleaning standard range of the cleaning layer, the overall data processing capability of the system is improved, thereby improving the pass rate of the system for data processing.
[0028] Furthermore, the present invention also re-determines the reasons for the failure of the data processing process based on the comparison results between the average flow rate and the pre-stored preset average flow rate, and increases the bandwidth of the acquisition layer based on the comparison results. The increase in the bandwidth of the acquisition layer is inversely proportional to the average flow rate, thereby increasing the number of samples of supplied data according to the average flow rate, and further improving the pass rate of data processing by improving the data processing capability of the system.
[0029] Furthermore, the present invention also re-determines the reasons for the failure of the data processing process based on the comparison results of the abnormal data ratio and the pre-stored preset abnormal data ratio, further deepening the judgment of data processing. By increasing the bandwidth of the acquisition layer or performing market trend analysis on individual abnormal data, the data processing capability of the system can be improved.
[0030] Furthermore, the present invention also re-determines the reasons for the failure of the data processing process based on the comparison results between the average slope and the pre-stored average slope, thereby further improving the judgment accuracy of data processing. By increasing the bandwidth of the acquisition layer or expanding the acquisition time interval of the acquisition layer, the data processing capability of the system can be improved.
[0031] Furthermore, the present invention also determines the expansion range of the cleaning standard range of the cleaning layer based on the difference between the data ratio and the pre-stored preset data ratio. The expansion range is proportional to the difference between the data ratios. By expanding the cleaning standard range, the proportion of complete data is increased, thereby improving the pass rate of data processing and thus improving the accuracy of supplier evaluation. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the self-learning data acquisition system for supplier evaluation according to the present invention;
[0033] Figure 2 This is a flowchart illustrating the self-learning data acquisition system for supplier evaluation according to the present invention.
[0034] Figure 3 This is a schematic diagram of the process for determining whether a data processing procedure is qualified based on a data ratio, according to the present invention.
[0035] Figure 4 This is a schematic diagram illustrating the process of determining the causes of data processing defects and correcting them based on the average value, according to the present invention.
[0036] Figure 5 This is a schematic diagram illustrating the process of determining the causes of data processing defects and making corrections based on the ratio of abnormal data, as per the present invention. Detailed Implementation
[0037] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0038] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0039] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0040] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0041] This embodiment provides a self-learning data acquisition system for supplier evaluation. This system collects supply data generated by each supplier during delivery. Then, based on the proportion of clearly identifiable and complete supply data, it determines the adequacy of the data processing. The string categories in the supply data include, but are not limited to, product information, logistics information, price information, and time points. Product information includes the product name, logistics information includes the expected delivery time, and price information includes the product price range. By determining the results, the system's operating parameters are adjusted, enabling the system to have a self-learning function. This improves the system's subsequent data processing capabilities, making subsequent supplier evaluations based on the filtered supply data more reasonable and accurate.
[0042] Please see Figure 1The diagram shown illustrates the modules of the self-learning data acquisition system for supplier evaluation in this embodiment. The system includes: an acquisition layer, an instruction layer, an acquisition layer, a processing layer, a cleaning layer, an analysis layer, a control layer, and an evaluation layer. The acquisition layer includes several acquisition terminals, each connected to the server of a corresponding supplier and acquiring supply data generated by the supplier within a preset period. The instruction layer, connected to the acquisition layer, sends acquisition instructions to the acquisition layer to enable each acquisition terminal to acquire the corresponding supply data. The acquisition layer, connected to the acquisition layer, acquires the supply data output by each acquisition terminal and compresses it. The processing layer, connected to the acquisition layer, receives and decompresses the supply data, and extracts tags from each supply data to obtain several strings, where the string categories include product information, logistics information, price information, and time nodes. The cleaning layer, connected to the processing layer, cleans the strings to filter out abnormal strings, where abnormal strings include random strings. The system comprises: a code string and a discrete string; an analysis layer, connected to the cleaning layer and the acquisition layer, for integrating the cleaned strings to obtain several complete data sets, and generating corresponding processing methods based on the data ratio, including: outputting complete data, or generating corresponding instructions based on determined reasons for non-compliance, wherein the data ratio is the ratio between the total number of complete data sets and the total number of supply data sets; a control layer, connected to the acquisition layer, the cleaning layer, and the analysis layer, for adjusting the operating parameters of the corresponding layers based on the instructions, including: the bandwidth of the acquisition layer, the acquisition time interval of the acquisition layer, the compression ratio of the acquisition layer, and the cleaning criteria of the cleaning layer; and an evaluation layer, connected to the analysis layer, for receiving several complete data sets and evaluating the corresponding suppliers based on the complete data sets.
[0043] Specifically, the purchasing party connects to the servers of various suppliers through several acquisition terminals in the acquisition layer, collecting supply data generated by suppliers when providing goods or services to the purchasing party within a preset period. Supply data for a single item contains several different types of strings, forming the supply data. The instruction layer connects to the acquisition layer and can issue acquisition instructions to the acquisition layer, enabling each acquisition terminal to obtain supply data from the corresponding supplier. The acquisition layer connects to both the acquisition and processing layers. The acquisition layer obtains the collected supply data from the acquisition layer, compresses and processes it, and then transmits it to the processing layer. The acquisition layer compresses and transmits the supply data in batches according to the time sequence when transmitting the data to the processing layer. The processing layer receives and decompresses the compressed supply data to obtain several supply data sets. It then extracts tags from each supply data set to obtain several strings specific to each supplier. The strings in a single supply data set are categorized into product information, logistics information, price information, and the time the data was uploaded to the corresponding server. These strings constitute the specific content of the supply data. At this point, the individual strings are arranged into a set of strings corresponding to the extraction time points. Each set of strings corresponds to a set of supply data. Subsequent supplier evaluations can be based on a comprehensive assessment of the string categories. The cleaning layer connects to the processing layer and obtains several strings. The cleaning layer can identify abnormal strings among these strings and marks supply data containing abnormal strings as abnormal data, removing them from the supply data set. Abnormal strings include garbled strings or discrete strings. For example, strings that deviate significantly from the guidance price in the existing market price information are marked as discrete strings, and data containing discrete strings are marked as abnormal data. Since abnormal data will affect subsequent supplier evaluations, it needs to be removed before any judgment is made during the data processing process. The analysis layer connects with the cleaning layer to obtain supply data after removing abnormal data. It then reassembles the data based on the corresponding time points during extraction for each type of string to obtain complete data. The ratio of the total number of complete data points to the total number of collected supply data points is calculated. This ratio is compared with a preset ratio stored in the analysis unit to determine whether the data acquisition process of the acquisition layer is qualified. Based on the determination result, the corresponding processing method is determined, including: directly outputting complete data and continuing data acquisition based on the current acquisition layer if the current acquisition process is qualified; or generating corresponding instructions based on the determined reasons for non-compliance and sending these instructions to the control layer connected to the analysis layer. The control layer adjusts the operating parameters of the corresponding layer according to the instructions. These parameters include: the bandwidth of the acquisition layer, the acquisition time interval of the acquisition layer, the compression ratio of the acquisition layer, and the cleaning standard of the cleaning layer.The evaluation layer is connected to the analysis layer. It acquires a number of complete data and evaluates the corresponding suppliers based on the complete data. During the evaluation, it generates several evaluation options based on various types of strings in the complete data, and generates corresponding scoring weights based on each evaluation option. Finally, the suppliers are reviewed based on each scoring weight.
[0044] Please see Figure 2 The diagram shown illustrates the flowchart of the self-learning data acquisition system for supplier evaluation in this embodiment. The flowchart is as follows:
[0045] S1: Collect supply data generated by suppliers within a preset period through the acquisition layer;
[0046] S2: Send acquisition commands to the acquisition layer through the command layer;
[0047] S3: Obtain several supply data through the acquisition layer and compress them;
[0048] S4: Receive and decompress each supply data through the processing layer to obtain various supply data, and extract tags from each supply data to obtain several strings;
[0049] S5: The cleaning layer cleans each string to filter out abnormal strings;
[0050] S6: The cleaned strings are integrated and processed by the analysis layer to obtain several complete data. The processing method is generated based on the data ratio, including: outputting complete data, or generating corresponding instructions based on the determined reasons for non-compliance.
[0051] S7: The control layer adjusts the operating parameters of the corresponding layer based on instructions. The parameters include: the bandwidth of the acquisition layer, the acquisition time interval of the acquisition layer, the compression ratio of the acquisition layer, and the cleaning standard of the cleaning layer.
[0052] S8: Receive several complete data sets through the evaluation layer and evaluate the corresponding suppliers based on the complete data.
[0053] Please see Figure 3 The diagram illustrates the process of determining whether a data processing procedure is qualified based on a data ratio in this embodiment. The analysis layer is further used to re-determine whether the data processing procedure is qualified based on the comparison result between the data ratio and a preset data ratio stored in the analysis layer, or based on the total number of collected data, and, if the data processing procedure is determined to be unqualified, to determine the reason for the unqualification based on the distribution of the abnormal data. The total number of collected data is the amount of supply data collected by the collection layer within a preset period.
[0054] Specifically, in this embodiment, the preset data ratio includes a first preset data ratio and a second preset data ratio that is greater than the first preset data ratio. The comparison process based on the data ratio and the preset data ratio is as follows:
[0055] If the data ratio is less than or equal to the first preset data ratio, it indicates that the total number of complete data is relatively small, meaning the sample size for subsequent supplier evaluation is insufficient. In this case, an effective and objective evaluation cannot be conducted based on the current amount of complete data, and the data processing process is deemed unqualified. The reason for the unqualification is determined based on the distribution of abnormal data. If the data ratio is greater than the first preset data ratio but less than or equal to the second preset data ratio, it indicates that the total number of complete data is between the sample size that can be evaluated and the sample size that cannot. A second judgment is required, re-evaluating the data processing process based on the total number of data collected by the acquisition layer for the supply data, to improve accuracy and avoid misjudgments. If the data ratio is greater than the second preset data ratio, it indicates that the total number of complete data meets the sample size requirement for evaluation, and therefore, the data processing process can be deemed qualified. Specifically, in this embodiment, the first preset data ratio H1 and the second preset data ratio H2 are set to 0.9*H and 1.1*H, respectively, where H = 0.95, rounded up when decimals are encountered. It should be noted that in other embodiments, H is determined based on the total number of supply data collected, and the multipliers of H1 and H2 are assigned values according to demand, and are not specifically limited here.
[0056] Furthermore, the analysis layer is also used to re-determine whether the data processing process is unqualified based on the comparison result between the total number of collected data and the preset total number of collected data stored in the analysis layer, or to increase the compression ratio of the acquisition layer based on the total difference, wherein the total difference is the difference between the total number of collected data and the preset total number of collected data.
[0057] Specifically, in this embodiment, the comparison process between the total number of collected data and the preset total number of collected data is as follows:
[0058] If the total number of collected data is less than or equal to the preset total number of collected data, it indicates that the total number of supplied data collected is relatively small. In this case, it can be determined that the data ratio is between the first and second preset data ratios because the total number of complete data is small. Therefore, the data processing process can be directly judged as unqualified, and the reason for the unqualification can be determined based on the distribution of abnormal data. If the total number of collected data is greater than the preset total number of collected data, it indicates that the total number of supplied data collected by the acquisition layer in the current preset period is relatively large. Therefore, it can be determined that the data ratio is between the first and second preset data ratios because the total number of supplied data is large. In this case, the data processing process can be determined as qualified, but it is necessary to increase the compression ratio of the acquisition layer through the control layer so that all the supplied data can be transmitted to the processing layer within a unit time. The compression ratio for the acquisition layer can be determined based on the difference in the total number. Specifically, in this embodiment, the preset total number of collected data D = 1000. In other embodiments, D can be assigned a value as needed, and is not specifically limited here.
[0059] Furthermore, the analysis layer is also used to increase the compression ratio of the acquisition layer based on the comparison result between the total difference and the preset total difference stored in the analysis layer, and the increase in compression ratio is proportional to the total difference.
[0060] Specifically, in this embodiment, the preset total difference includes a first preset total difference and a second preset total difference. The first preset total difference is less than the second preset total difference. The comparison process based on the total difference and the preset total difference is as follows:
[0061] If the total difference is less than or equal to the first preset total difference, the compression ratio is increased to 1.15 times the initial value using the first compression ratio adjustment coefficient by the control layer. If the total difference is greater than the first preset total difference and less than or equal to the second preset total difference, the compression ratio is increased to 1.35 times the initial value using the second compression ratio adjustment coefficient by the control layer. If the total difference is greater than the second preset total difference, the compression ratio is increased to 1.55 times the initial value using the third compression ratio adjustment coefficient by the control layer. Specifically, in this embodiment, the first preset total difference N1 and the second preset total difference N2 are assigned values of 1.1*N and 1.3*N, respectively, where N=100 records. It should be noted that the increase factor of the compression ratio is also related to the total amount of supply data. The larger the total amount of supply data collected, the greater the corresponding increase should be made on the basis of the original adjustment.
[0062] Please see Figure 4The diagram illustrates the process of determining and correcting data processing defects based on average values in this embodiment. The analysis layer is further used to determine the reasons for data processing defects based on a comparison between the average value and a preset average value of the analysis layer, and to generate corresponding correction methods based on the determined reasons. These methods include: increasing the bandwidth of the acquisition layer, expanding the cleaning standard range of the cleaning layer, or re-determining the reasons for defects based on the types of strings in the abnormal data. Specifically, this involves obtaining the time node within a preset period from the time node where the abnormal data was identified, obtaining the interval between adjacent abnormal data based on the time node, and calculating the average value based on the interval.
[0063] Specifically, in this embodiment, when determining the cause of data processing failure, since the supply data is transmitted according to time nodes during the acquisition layer, the time node corresponding to when the cleaning layer identifies abnormal data within a preset period can be obtained first. The interval between each pair of abnormal data points is calculated according to each time node. Then, the interval duration and the number of abnormal data points are accumulated, and an average value is calculated based on this. By comparing the average value with a preset average value, the distribution of abnormal data at the time node level is determined, thereby identifying the cause. The preset average value includes a first preset average value and a second preset average value that is greater than the first preset average value. The comparison process based on the average value and the preset average value is as follows:
[0064] If the average value is less than or equal to the first preset average value, it indicates that the interval between abnormal data detected by the cleaning layer is relatively short, and the abnormal data is concentrated. In this case, the bandwidth of the acquisition layer can be increased by the control layer, i.e., increasing the acquisition capability of each acquisition terminal, so as to increase the amount of supply data transmitted within the preset period, thereby expanding the interval and increasing the proportion of complete data in the supply data within the preset period. This can improve the data processing pass rate by adjusting the system. If the average value is greater than the first preset average value but less than or equal to the second preset average value, the distribution of abnormal data cannot be accurately determined. In this case, the specific reasons for non-compliance can be further determined based on the string type in the abnormal data, improving the accuracy of determining the reasons for non-compliance. If the average value is greater than the second preset average value, it indicates that the interval between abnormal data detected by the cleaning layer is relatively long, and the abnormal data is relatively scattered. In this case, the reason for non-compliance is due to the inappropriate removal standard for abnormal data. It is necessary to expand the removal standard of the cleaning layer based on the data ratio difference so that the abnormal data that should be removed is identified as normal supply data, increasing the amount of complete data. This can improve the data processing pass rate by adjusting the system. Specifically, the data ratio difference is the difference between the first preset data ratio and the data ratio. Specifically, in this embodiment, the first preset average value T1 and the second preset average value T2 are set to 4*T and 5*T, respectively, where T = 10 milliseconds. In other embodiments, the settings of T1 and T2 and the assignment of T can be reset as needed, and are not specifically limited here.
[0065] Furthermore, the analysis layer is also used to increase the bandwidth of the acquisition layer based on the comparison result between the average traffic value and the preset average traffic value stored in the analysis layer, and the increase in bandwidth is inversely proportional to the average traffic value. In this case, the traffic value within the time interval of each abnormal data group is obtained, and the average traffic value is calculated based on each traffic value.
[0066] Specifically, in this embodiment, several consecutive occurrences of abnormal data are recorded as an abnormal data group. The flow rate value of a single abnormal data group within the corresponding time interval is obtained. Based on this flow rate value and the number of abnormal data groups, the average flow rate is calculated. The magnitude of the average flow rate reflects the amount of supply data that the acquisition layer can collect per unit time. The smaller the average flow rate, the less supply data the acquisition layer can collect per unit time. The preset average flow rate includes a first preset average flow rate and a second preset average flow rate that is greater than the first preset average flow rate. The comparison process between the average flow rate and the preset average flow rate is as follows:
[0067] If the average traffic is less than or equal to the first preset average traffic, the bandwidth is increased to 3.2 times the initial value using a first bandwidth adjustment coefficient by the control layer. If the average traffic is greater than the first preset average traffic and less than or equal to the second preset average traffic, the bandwidth is increased to 2.5 times the initial value using a second bandwidth adjustment coefficient by the control layer. If the average traffic is greater than the second preset average traffic, the bandwidth is increased to 1.6 times the initial value using a third bandwidth adjustment coefficient by the control layer. Specifically, in this embodiment, the first preset average traffic V1 and the second preset average traffic V2 are set to 1.5*V and 1.9*V, respectively, where V = 10 Mbps. In other embodiments, the settings of V1 and V2 and the assignment of V are reset according to requirements and are not specifically limited here.
[0068] Please see Figure 5 The diagram illustrates the process of determining and correcting data processing failures based on the anomaly ratio in this embodiment. The analysis layer is further used to re-determine the reasons for data processing failures based on a comparison between the anomaly ratio and a preset anomaly ratio, and to generate corresponding correction methods based on the determined reasons, including: increasing the bandwidth of the acquisition layer, or performing market trend analysis on data of a single type of anomaly, wherein the anomaly ratio is the ratio of anomaly data containing the garbled string to the total anomaly data.
[0069] Specifically, in this embodiment, the comparison process based on the abnormal data ratio and the preset abnormal data ratio is as follows:
[0070] If the abnormal data ratio is greater than the preset abnormal data ratio, it indicates that the number of abnormal data generated by string garbled characters is greater than the number of abnormal data generated by string discretization. In this case, the main reason for the abnormal supply data is the presence of too many garbled strings. Therefore, it can be determined that the reason for the non-compliance is a problem in the data collection process, requiring an increase in the bandwidth of the collection layer to improve collection efficiency. If the abnormal data ratio is less than or equal to the preset abnormal data ratio, it indicates that the number of abnormal data generated by string garbled characters is less than the number of abnormal data generated by string discretization. In this case, the main reason for the abnormal supply data is the presence of too many discrete strings. In this embodiment, taking the discrete occurrence of price information strings as an example, several historical price data points within the corresponding historical period are obtained for the price information. By plotting the historical price data and the abnormal data within the current preset period into a curve, the reason for the non-compliance of data processing is determined based on the curve changes. For example, if the historical price data fluctuates within the range of 80-100 yuan, while the current price data is 121, exceeding 100 by 20%, or 63, falling below 80 by 20%, then the string of the current price data is determined to be a discrete string. Specifically, in this embodiment, the preset abnormal data ratio G is assigned a value of 0.6. In other embodiments, G is assigned a value based on the determined number of abnormal data, which is not limited here.
[0071] Furthermore, the analysis layer is also used to re-determine the reasons for the data processing failure based on the comparison results between the average slope and the preset average slope stored in the analysis layer, and to generate corresponding correction methods based on the determined reasons, including: expanding the collection time interval of the collection layer, or increasing the bandwidth of the collection layer, wherein several historical price data within the corresponding historical period are obtained, and a curve is drawn based on the several historical price data and the current price data and the average slope is calculated.
[0072] Specifically, in this embodiment, taking the string in the supply data as price information as an example, a time-price curve is drawn based on the time node of acquiring the price data, and the price slope corresponding to each time node is calculated, and the average slope is calculated. The comparison process based on the average slope and the preset average slope is as follows:
[0073] If the average slope is less than or equal to the preset average slope, it indicates that the historical price data in the market has changed relatively steadily. The current abnormal data generated by discrete strings is actually normal supply data, which is based on market changes. Therefore, by expanding the collection time interval of the collection layer, the current collection time range of the supply data can be extended, thereby redetermining new discrete string abnormal data based on the new comparison benchmark. Since the market changes are relatively stable, the number of redetermined discrete strings will be lower than the previous discrete strings, thus solving the problem of a large number of discrete string abnormal data. If the average slope is greater than the preset average slope, it indicates that the price data of the current abnormal data deviates significantly from the range of historical price data in the market. Therefore, it indicates that the current collection time interval of the collection layer is not problematic, and it is determined that the abnormal data is concentrated within the current collection time. Therefore, the bandwidth of the collection layer can be increased by the control layer to increase the proportion of complete data in the supply data within the preset period, thereby improving the data processing qualification rate by adjusting the system. Specifically, in this embodiment, the time node interval is set to 1 hour, and the average slope K = 20%. In other embodiments, the setting value of K is determined according to market price changes.
[0074] Furthermore, the analysis layer is also used to expand the acquisition time interval of the acquisition layer based on the comparison result between the slope mean difference and the preset slope mean difference stored in the analysis layer, and the expansion of the acquisition time interval is inversely proportional to the slope mean difference, wherein the slope mean difference is the difference between the preset slope mean difference and the slope mean difference.
[0075] Specifically, in this embodiment, the slope mean difference includes a first slope mean difference and a second slope mean difference that is greater than the first slope mean difference. The comparison process based on the slope mean difference and the preset slope mean difference is as follows:
[0076] If the slope mean difference is less than or equal to the first slope mean difference, the control layer uses a first time interval adjustment coefficient to increase the collection time interval to 3.5 times the initial interval. If the slope mean difference is greater than the first slope mean difference and less than or equal to the second slope mean difference, the control layer uses a second time adjustment coefficient to increase the collection time interval to 2.5 times the initial interval. If the slope mean difference is greater than the second slope mean difference, the control layer uses a third time adjustment coefficient to increase the collection time interval to 1.5 times the initial interval. The larger the slope mean difference, the smaller the slope mean, indicating a more stable market change and thus less need for diverse supply data samples, resulting in a smaller increase in the collection time interval. Specifically, in this embodiment, the slope mean difference ΔK = 5%, the first slope mean difference ΔK1 = 0.8 * ΔK, and the second slope mean difference ΔK2 = 1.2 * ΔK. In other embodiments, ΔK1 and ΔK2, as well as the amplification factor of the collection time interval, are determined based on historical market price data.
[0077] Furthermore, the analysis layer is also used to expand the cleaning standard range of the cleaning layer based on the comparison result between the data ratio difference and the preset data ratio difference stored in the analysis layer, and the expansion of the cleaning standard is proportional to the data ratio difference, wherein the data ratio difference is the difference between the preset data ratio and the data ratio.
[0078] Specifically, in this embodiment, since the abnormal data is relatively scattered, the scope of the removal criteria can be expanded to include the originally removed abnormal data into the normal data, thereby increasing the proportion of complete data and improving the data processing pass rate. The preset data ratio difference includes a first preset data ratio difference and a second preset data ratio difference. The first preset data ratio difference is smaller than the second preset data ratio difference. The comparison process based on the data ratio difference and the preset data ratio difference is as follows:
[0079] If the data ratio difference is less than or equal to the first preset data ratio difference, the control layer uses the first clearing standard adjustment coefficient to expand the clearing standard range to 1.35 times the initial value. If the data ratio difference is greater than the first preset data ratio difference and less than or equal to the second preset data ratio difference, the control layer uses the second clearing standard adjustment coefficient to expand the clearing standard range to 1.65 times the initial value. If the data ratio difference is greater than the second preset data ratio difference, the control layer uses the third clearing standard adjustment coefficient to expand the clearing standard to 1.86 times the initial value.
[0080] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0081] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A self-learning data acquisition system for supplier evaluation, characterized in that, The application relates to a data processing system, comprising: a collection layer comprising a plurality of collection terminals, each of which is connected to a server of a corresponding supplier and collects supply data generated by the supplier within a preset period; an instruction layer connected to the collection layer, used to send collection instructions to the collection layer to enable each collection terminal to obtain the corresponding supply data; an acquisition layer connected to the collection layer, used to acquire a plurality of supply data output by each collection terminal and perform compression processing; a processing layer connected to the acquisition layer, used to receive and decompress each supply data, and perform label extraction on each supply data to obtain a plurality of strings, wherein the categories of the strings include product information, logistics information, price information and time nodes; a cleaning layer connected to the processing layer, used to clean each string to filter out abnormal strings, wherein the abnormal strings include garbled code strings and discrete strings; an analysis layer connected to the cleaning layer and the collection layer, used to integrate the cleaned strings to obtain a plurality of complete data, generate a corresponding processing mode based on a data ratio, including outputting the complete data, or generating a corresponding instruction based on a determined unqualified reason, wherein the data ratio is the ratio between the total number of complete data and the total number of supply data; a control layer connected to the collection layer, the acquisition layer, the cleaning layer and the analysis layer, used to adjust the operating parameters of the corresponding layers based on the instruction generated by the determined unqualified reason, the parameters including the bandwidth of the collection layer, the collection time interval of the collection layer, the compression ratio of the acquisition layer and the cleaning standard range of the cleaning layer; an evaluation layer connected to the analysis layer, used to receive a plurality of complete data and evaluate the corresponding suppliers based on the complete data.
2. The self-learning data collection system for supplier evaluation of claim 1, wherein, The analysis layer is also used to compare the data ratio with a preset data ratio pre-stored in the analysis layer, or re-determine whether the data processing process is qualified based on the total number of collected data, and determine the unqualified reason based on the distribution of abnormal data in the case that the data processing process is determined to be unqualified, wherein the total number of collected data is the number of supply data collected by the collection layer within a preset period.
3. The self-learning data collection system for supplier evaluation of claim 2, wherein, The analysis layer is also used to re-determine that the data processing process is unqualified based on the comparison result of the total number of collected data and a preset total number of collected data pre-stored in the analysis layer, or increase the compression ratio of the acquisition layer based on the total number difference, wherein the total number difference is the difference between the total number of collected data and the preset total number of collected data.
4. The self-learning data collection system for supplier evaluation of claim 3, wherein, The analysis layer is also used to increase the compression ratio of the acquisition layer based on the comparison result of the total number difference and a preset total number difference pre-stored in the analysis layer, and the increase amplitude of the compression ratio is proportional to the total number difference.
5. The self-learning data collection system for supplier evaluation of claim 2, wherein, The analysis layer is further configured to determine the reason for the unqualified data processing process based on a comparison result of the average value and a preset average value of the analysis layer, and generate a corresponding correction method based on the determined reason, including: increasing the bandwidth of the collection layer, expanding the removal standard range of the cleaning layer, or re-determining the unqualified reason based on the string type in the abnormal data, wherein the average value is calculated based on the interval duration between adjacent abnormal data at a time node when the abnormal data is identified in a preset period.
6. The self-learning data collection system for supplier evaluation of claim 5, wherein, The analysis layer is further configured to increase the bandwidth of the collection layer based on a comparison result of the flow average value and a preset flow average value stored in the analysis layer, and the increase amplitude of the bandwidth is inversely proportional to the flow average value, wherein the flow value in the time interval where each abnormal data group is located is obtained, and the flow average value is calculated based on each flow value.
7. The self-learning data collection system for supplier evaluation of claim 6, wherein, The analysis layer is further configured to re-determine the reason for the unqualified data processing process based on a comparison result of the abnormal data ratio and a preset abnormal data ratio, and generate a corresponding correction method based on the determined reason, including: increasing the bandwidth of the collection layer, or performing market trend analysis on the single type of abnormal data, wherein the abnormal data ratio is the ratio of the abnormal data containing the garbled string to the total abnormal data.
8. The self-learning data collection system for supplier evaluation of claim 7, wherein, The analysis layer is further configured to re-determine the reason for the unqualified data processing process based on a comparison result of the slope average value and a preset slope average value stored in the analysis layer, and generate a corresponding correction method based on the determined reason, including: expanding the collection time interval of the collection layer, or increasing the bandwidth of the collection layer, wherein a plurality of historical price data in a corresponding historical period is obtained, and a curve is drawn based on the plurality of historical price data and the current price data, and the slope average value is calculated.
9. The self-learning data collection system for supplier evaluation of claim 8, wherein, The analysis layer is further configured to expand the collection time interval of the collection layer based on a comparison result of the slope average value difference and a preset slope average value difference stored in the analysis layer, and the expansion amplitude of the collection time interval is inversely proportional to the slope average value difference, wherein the slope average value difference is the difference between the preset slope average value difference and the slope average value difference.
10. The self-learning data collection system for supplier evaluation of claim 5, wherein, The analysis layer is further configured to expand the removal standard range of the cleaning layer based on a comparison result of the data ratio difference and a preset data ratio difference stored in the analysis layer, and the expansion amplitude of the removal standard is proportional to the data ratio difference, wherein the data ratio difference is the difference between the preset data ratio and the data ratio.
Citation Information
Patent Citations
Supplier portrait modeling method and system
CN113345080A
Supplier intelligent evaluation and grading system based on multi-dimensional material data
CN113706009A