A data standardization processing method and system for an industrial supply chain

CN122509864APending Publication Date: 2026-08-04XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN UNIV OF TECH
Filing Date
2026-06-09
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0005]本发明提供了一种用于产业供应链的数据标准化处理方法及系统,以解决信息整合效率低下的问题

Benefits of technology

[0023] (1) This invention generates time adjustment markers by parsing supplier data update records, maps the delay duration to production sequence displacement vectors and prioritizes downstream production data, and generates a priority push list by combining resource occupation fluctuation values ​​and supplier real-time load. This enables real-time perception of delivery anomalies and dynamic adaptation of production scheduling, solving the problems of delayed response to delivery time changes and disconnection between production scheduling and material supply in traditional technologies. It effectively avoids production line shutdowns due to material shortages and improves the utilization efficiency of production capacity and logistics resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122509864A_ABST
    Figure CN122509864A_ABST
Patent Text Reader

Abstract

This invention relates to the field of supply chain data processing technology, and discloses a data standardization processing method and system for industrial supply chains. The method includes acquiring supply chain supplier, logistics, and production data and integrating them into heterogeneous data; obtaining a format classification set through K-means clustering; parsing semantics and matching them with a standard dictionary to obtain a set of field differences; filling in missing values ​​exceeding a threshold and establishing a field mapping table; extracting supplier update records based on the mapping table, and generating time adjustment markers for delivery changes; obtaining a priority push list through hierarchical matching of downstream production data, sending scheduling notifications, and obtaining response times; generating a dynamic latency threshold by combining node load and latency; retrying pushes and correcting for exceeding the threshold to obtain collaboration parameters; updating logs; and outputting a business collaboration adjustment plan after the data flow meets the requirements. This method can improve the efficiency of supply chain data integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of supply chain data processing technology, and in particular to a data standardization processing method and system for industrial supply chains. Background Technology

[0002] Currently, industrial supply chain management is a crucial field, involving the entire process from raw material procurement to product delivery. The efficient integration and sharing of data has become the key to ensuring the smooth operation of each link.

[0003] In one existing technology, heterogeneous data is manually processed and transmitted in batches via fixed electronic data interchange, relying on preset mappings and manual anomaly detection. However, data generated by different enterprises and at different stages often have inconsistent formats and update frequencies. Due to a lack of big data management capabilities, existing methods often result in low efficiency in supply chain data integration when processing cross-organizational and cross-system data.

[0004] In summary, existing technologies suffer from low efficiency in supply chain data integration. Summary of the Invention

[0005] This invention provides a data standardization processing method and system for industrial supply chains to solve the problem of low information integration efficiency.

[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a data standardization processing method for industrial supply chains, comprising:

[0007] The raw heterogeneous data of suppliers, logistics nodes and production links in the supply chain are obtained, and the raw heterogeneous data are grouped by K-means clustering algorithm to obtain a set of format classifications.

[0008] The format classification set is completed by using a preset standard dictionary to obtain a high-quality data sequence. The mapping relationship of fields is established based on the high-quality data sequence to obtain a format mapping table.

[0009] Based on the formatted mapping table, data update records are extracted from the original heterogeneous data. If the data update record involves a change in delivery time, a time adjustment flag is generated.

[0010] Acquire downstream production data, prioritize the downstream production data using the time adjustment marker, and obtain a priority push list;

[0011] Extract material codes from the priority push list, map the material codes to supplier IDs, send customized scheduling notifications to suppliers based on the supplier IDs, and obtain the time it takes for suppliers to respond to the customized scheduling notifications to obtain response time data.

[0012] Obtain the node load status of the supplier and combine it with the real-time time consumption value in the response time data to generate a dynamic latency threshold; if the real-time time consumption value is greater than the dynamic latency threshold, send the customized scheduling notification content to the supplier again, and correct the dynamic latency threshold based on the feedback of the second push to determine the collaborative response parameters.

[0013] Based on the collaboration response parameters, update the push log records of the supply chain data stream, extract the updated supply chain data stream segments based on the push log records, and if the supply chain data stream segments meet the preset grouping requirements, output the business collaboration adjustment plan.

[0014] Secondly, the present invention provides a data standardization processing apparatus for an industrial supply chain, comprising:

[0015] The heterogeneous data acquisition module is used to acquire raw heterogeneous data from suppliers, logistics nodes, and production processes in the supply chain, and uses the K-means clustering algorithm to group the raw heterogeneous data to obtain a set of format classifications.

[0016] The data standardization mapping module is used to complete the format classification set by using a preset standard dictionary to obtain a high-quality data sequence, and to establish the field mapping relationship based on the high-quality data sequence to obtain a formatted mapping table;

[0017] The delivery change marking module is used to extract data update records from the original heterogeneous data according to the formatted mapping table, and generate a time adjustment mark if the data update record involves a change in delivery time.

[0018] The production priority classification module is used to acquire data from downstream production processes, and to classify the data from downstream production processes into priority categories based on the time adjustment markers, thereby obtaining a priority push list.

[0019] The collaborative scheduling and push module is used to extract material codes from the priority push list, map the material codes to supplier IDs, send customized scheduling notification content to suppliers based on the supplier IDs, and obtain the time when the supplier responds to the customized scheduling notification content to obtain response time data.

[0020] The dynamic latency optimization module is used to obtain the node load status of the supplier and generate a dynamic latency threshold by combining the real-time time consumption value in the response time data; if the real-time time consumption value is greater than the dynamic latency threshold, the customized scheduling notification content is sent to the supplier again, and the dynamic latency threshold is corrected according to the feedback of the second push to determine the collaborative response parameters.

[0021] The closed-loop collaboration verification module is used to update the push log records of the supply chain data flow according to the collaboration response parameters, extract the updated supply chain data flow segments according to the push log records, and output a business collaboration adjustment plan if the supply chain data flow segments meet the preset grouping requirements.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] (1) This invention generates time adjustment markers by parsing supplier data update records, maps the delay duration to production sequence displacement vectors and prioritizes downstream production data, and generates a priority push list by combining resource occupation fluctuation values ​​and supplier real-time load. This enables real-time perception of delivery anomalies and dynamic adaptation of production scheduling, solving the problems of delayed response to delivery time changes and disconnection between production scheduling and material supply in traditional technologies. It effectively avoids production line shutdowns due to material shortages and improves the utilization efficiency of production capacity and logistics resources.

[0024] (2) This invention integrates heterogeneous data from multiple links in the supply chain, uses K-means clustering algorithm combined with cosine similarity to complete field standardization classification, and uses time series linear interpolation to complete missing data and establish a formatted mapping table. This effectively solves the information interaction barrier problem of heterogeneous supply chain data format and inconsistent field naming in traditional technology, and improves the interoperability and flow accuracy of cross-node data.

[0025] (3) This invention generates a dynamic delay threshold by combining the load of supplier nodes, performs a second push on the scheduling notification for response timeout, and determines the collaborative response parameters by correcting the threshold based on the push feedback. Then, it updates the log based on the parameters and outputs an adapted business collaboration adjustment scheme, realizing closed-loop feedback and adaptive optimization of supply chain collaborative scheduling. This solves the problems of fixed scheduling judgment criteria and lack of flexibility in response strategies in traditional technologies, and significantly improves the robustness and intelligence level of supply chain collaboration. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the data standardization processing method for industrial supply chains provided in the first embodiment of the present invention;

[0027] Figure 2 This is a schematic diagram of the data standardization processing system for industrial supply chains provided in the second embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] Reference Figure 1 The first embodiment of the present invention provides a data standardization process for industrial supply chains, including the following steps:

[0030] S11, Obtain the original heterogeneous data of suppliers, logistics nodes and production links in the supply chain, and use the K-means clustering algorithm to group the original heterogeneous data to obtain a format classification set;

[0031] S12, the format classification set is completed by using a preset standard dictionary to obtain a high-quality data sequence, and the field mapping relationship is established based on the high-quality data sequence to obtain a format mapping table;

[0032] S13, according to the formatted mapping table, extract data update records from the original heterogeneous data; if the data update records involve changes in delivery time, generate a time adjustment flag.

[0033] S14, Obtain downstream production process data, and prioritize the downstream production process data according to the time adjustment marker to obtain a priority push list;

[0034] S15, extract material codes from the priority push list, map the material codes to supplier IDs, send customized scheduling notification content to suppliers based on the supplier IDs, and obtain the time when the supplier responds to the customized scheduling notification content to obtain response time data.

[0035] S16, obtain the node load status of the supplier, and generate a dynamic latency threshold by combining the real-time time consumption value in the response time data; if the real-time time consumption value is greater than the dynamic latency threshold, send the customized scheduling notification content to the supplier again, and correct the dynamic latency threshold according to the feedback of the second push to determine the collaborative response parameters.

[0036] S17. Update the push log of the supply chain data stream according to the collaboration response parameters, extract the updated supply chain data stream fragments according to the push log records, and output the business collaboration adjustment plan if the supply chain data stream fragments meet the preset grouping requirements.

[0037] In step S11, raw heterogeneous data from suppliers, logistics nodes, and production processes in the supply chain are obtained. The raw heterogeneous data is then grouped using the K-means clustering algorithm to obtain a set of format classifications, including:

[0038] Data is collected from suppliers, logistics nodes, and production links in the supply chain to obtain the first heterogeneous data, which includes text, tables, and timestamps.

[0039] The first heterogeneous data is subjected to mean interpolation to obtain the second heterogeneous data;

[0040] Obtain the field names of the second heterogeneous data, vectorize the field names using a preset bag-of-words model to obtain field name vectors, calculate the cosine similarity of the field name vectors, and obtain the field name difference metric.

[0041] When the field naming difference metric exceeds a preset difference threshold, feature extraction is performed on the second heterogeneous data to obtain the original heterogeneous data.

[0042] The original heterogeneous data was grouped using the K-means clustering algorithm to obtain a format classification set.

[0043] In this invention, dedicated data collection points are first deployed at the supplier, logistics, and production stages of the supply chain. Each collection point establishes a stable data interaction link with the corresponding business system, extracting full business data from each system to obtain a first heterogeneous data set encompassing text, tables, and timestamps. Specifically, the supplier-side data includes textual data such as material specifications, the logistics node-side data includes tabular data such as transportation status statistics, and the production stage-side data includes timestamped data such as process execution records. All data types retain their original format and field attributes.

[0044] The first heterogeneous data undergoes integrity verification by traversing each data field and checking for missing values. For numerical fields with missing values, all valid data for that field are extracted, and their arithmetic mean is calculated. This arithmetic mean is then used to fill in the missing value positions, resulting in the second heterogeneous data. For example, in an automotive manufacturing supply chain scenario, the temperature monitoring field for a certain batch of tires in a tire supplier's logistics transportation data contains three consecutive missing values. Ten valid data points before and after this field are extracted, and their arithmetic mean is calculated to be 25°C. This value is then used to fill in the missing positions, completing the interpolation. This yields the second heterogeneous data. Due to the independence of the systems of different participants, inconsistencies in field naming exist in the second heterogeneous data. The second heterogeneous data undergoes full field parsing to extract the original naming text of each data field, forming a set of field naming information.

[0045] It should be noted that the preset bag-of-words model is used to perform vectorization conversion on the field naming information set, converting each named text into a multi-dimensional numerical vector of the same dimension. The vector dimension is determined by the total number of words after removing duplicates from the word segmentation results of all field naming texts. Each component in the vector represents the frequency of occurrence of the corresponding word in the named text. The cosine similarity between any two field naming vectors is calculated and used as the field naming difference metric value.

[0046] The preset bag-of-words model is used to convert field naming texts into numerical vectors for similarity calculation. The construction of this model includes two stages: vocabulary library construction and vectorization rule definition.

[0047] In the vocabulary library construction stage, first collect all the field naming texts that appear in the business systems of suppliers, logistics providers, and manufacturers already connected to the supply chain collaboration platform to form an original corpus set. Perform word segmentation on the original corpus using a word segmenter based on the supply chain industry dictionary. This word segmenter has built-in business term libraries for major industries such as automobile manufacturing and electronics manufacturing. When segmenting words, retain complete business vocabulary and remove stop words such as "of", "了", "in", "and" that have no practical meaning. Count the total frequency of each word in the original corpus, set the low-frequency word filtering threshold to 3 times, and剔除 words with a frequency of occurrence lower than 3 times to reduce the vector dimension and the impact of noise. Use the filtered word list as candidate roots, which are reviewed by supply chain business experts to剔除 words irrelevant to the enterprise's core business and merge synonyms. Finally, form the supply chain core business vocabulary library. The vocabulary library is stored in the supply chain collaboration database in the form of a structured data table, containing fields such as vocabulary identifier, vocabulary text, business domain, creation time, and status. The vocabulary library is updated quarterly. When updating, collect the newly added field naming texts in the past three months, repeat the above word segmentation, statistics, filtering, and review processes, append the newly added words to the vocabulary library, and mark words that have not appeared in two consecutive update cycles as abandoned status, but retain their historical records to maintain the traceability of the generated vectors. The total number of words in the vocabulary library is the dimension of the vector. For example, if the initial version of the vocabulary library includes 800 core business words, the vector dimension is 800.

[0048] In the vectorization rule definition phase, the mapping method from field-named text to numerical vectors is determined. Term frequency vectorization is adopted. For a given field-named text, it is first segmented, with segmentation rules consistent with those in the vocabulary construction phase. A zero vector of length equal to the total number of words in the vocabulary is created, where each position in the vector corresponds one-to-one with a word in the vocabulary, arranged in ascending order of word identifiers. Each word in the segmentation results is traversed, and its corresponding position index in the vocabulary is found, with the component value of that position incremented by 1. After traversal, the term frequency vector corresponding to the field-named text is obtained, where each component value represents the number of times the corresponding word appears in the text. For example, after segmenting the field-named text "shipping date" into two words, "shipping" and "date," if the index of "shipping" in the vocabulary is 120 and the index of "date" is 450, then the generated vector will have component values ​​of 1 in the 120th and 450th dimensions, and 0 in the remaining dimensions. TF-IDF weights are not used in the vectorization process; the original term frequency method is used to maintain computational simplicity and suitability for short text scenarios.

[0049] When calculating similarity, the word frequency vectors corresponding to the named texts of the two fields are taken, and the cosine similarity is calculated. The formula for cosine similarity is the dot product of the two vectors divided by the product of their magnitudes. The result ranges from 0 to 1, with a value closer to 1 indicating a higher semantic similarity between the two named texts. A similarity matching threshold of 0.8 is set. This threshold is determined by testing 100 sets of manually labeled synonymous field name pairs and 100 sets of manually labeled non-synonymous field name pairs, selecting the threshold corresponding to the balance between recall and precision. When the cosine similarity of the named texts of the two fields is greater than or equal to 0.8, they are considered semantically matched; when the similarity is less than 0.8, field differences are considered to exist.

[0050] It should be noted that the cosine similarity calculation result ranges from [-1, 1]. The closer the result is to 1, the higher the similarity of the two field names and the smaller the corresponding field naming difference metric. The further the result is from 1, the lower the similarity and the larger the difference metric. For example, if the supplier names the delivery time field as "shipping date" and the production team names the field with the same business meaning as "warehousing time," the calculated cosine similarity after vectorization is 0.98, which exceeds the preset similarity threshold of 0.8. Therefore, the two are considered synonymous fields. Since the field names in each link of the supply chain are short texts with highly overlapping core business terms, setting the similarity threshold to 0.8 can accommodate slight differences in expression while avoiding the erroneous merging of semantically unrelated fields.

[0051] In this invention, synonymous fields with a cosine similarity exceeding a preset threshold of 0.8 are uniformly mapped to standard field names in the supply chain industry. For example, the "shipping date" from the supplier, the "delivery time" from the logistics node, and the "warehousing time" from the production stage are uniformly mapped to the standard field "material delivery time". Based on the mapped fields, key features, such as material delivery time and material turnover cycle, are extracted to form the original data set.

[0052] In this invention, the K-means clustering algorithm is used to cluster the original heterogeneous data. Based on the business type of the supply chain data, the number of clusters K=3, corresponding to the supplier material characteristic cluster, logistics timeliness characteristic cluster, and production assembly characteristic cluster, respectively. Three feature vectors possessing the aforementioned typical characteristics are randomly selected from the original heterogeneous data as initial cluster centers. The Euclidean distance from each feature vector to each cluster center is calculated, and the feature vector is assigned to the nearest cluster. Then, the mean vector of each cluster is recalculated as the new cluster center. This process is repeated until the change in the cluster centers is less than a preset convergence threshold (e.g., 0.001), resulting in a trained clustering model.

[0053] The original heterogeneous data is input into the trained clustering model. Based on the clustering results, the corresponding second heterogeneous data is divided into supplier material feature group, logistics timeliness feature group, and production assembly feature group, which together constitute a format classification set.

[0054] For example, in the automotive manufacturing supply chain scenario, feature vectors such as material model and material specifications are assigned to the supplier material feature cluster, feature vectors such as transportation time and loading and unloading time are assigned to the logistics timeliness feature cluster, and feature vectors such as assembly time and process completion rate are assigned to the production assembly feature cluster. Based on this clustering result, the second heterogeneous data is classified to obtain a format classification set containing the above three groups of data.

[0055] In step S12, the format classification set is completed using a preset standard dictionary to obtain a high-quality data sequence. Based on this high-quality data sequence, a field mapping relationship is established to obtain a format mapping table, including:

[0056] The business semantics of the format classification set are analyzed, the matching degree between the business semantics and the preset standard dictionary is calculated, and the set of field differences is determined.

[0057] Based on a preset conversion protocol, the set of field differences is renamed to obtain the data block to be verified;

[0058] The percentage of missing data blocks to be verified is calculated to obtain the data integrity verification result;

[0059] If the data integrity verification result shows that the missing percentage exceeds the preset missing threshold, the missing values ​​are filled with weighted data using time series linear interpolation to obtain a high-quality data sequence.

[0060] Based on the high-quality data sequence, a mapping relationship between the fields is established to obtain a formatted mapping table.

[0061] First, the field names of each data field are extracted from the format classification set to serve as the business semantics to be parsed. A pre-defined standard dictionary is used to provide a standardized reference for field naming. Its construction process begins by extracting all field names from the business systems of suppliers, logistics providers, and manufacturers already connected to the supply chain collaboration platform, forming an original field name set. The field names in this set are then segmented, removing punctuation marks and common suffixes such as "date," "time," and "number," to extract core business terms. The frequency of each core business term is counted, and business terms with a frequency higher than 5 times are selected as candidate standard word roots.

[0062] Supply chain business experts review candidate standard terminology, eliminating ambiguous terms and merging synonyms. For example, "delivery date," "arrival date," and "delivery time" are unified as "delivery time," and "part number," "material code," and "part number" are unified as "material code," forming a standard terminology table. This table is stored in the supply chain collaboration database as a structured data table, containing fields such as standard terminology identifier, standard terminology name, business domain, creation time, and status. The default standard dictionary is constructed from standard terminology through combination rules, using the format "business object_attribute_unit," such as "material_delivery time" or "logistics_transportation time_minutes." Each standard dictionary entry includes fields such as standard field name, standard field identifier, data type, length constraint, business domain, and effective version number. The initial version of the default standard dictionary contains 300 standard field definitions, covering core business objects such as supplier master data, bill of materials, purchase orders, transportation documents, and production work orders. Subsequent versions are iterated quarterly based on new business needs, and new entries must be reviewed and approved by business experts before being added to the dictionary table.

[0063] The matching degree is calculated using the cosine similarity method. Specifically, the original field name and the standard field name are segmented into words separately. A pre-defined bag-of-words model is used to vectorize the segmentation results, with the vector dimension matching the total number of words in the standard root word table. The cosine similarity between the two vectors is calculated, with similarity values ​​ranging from 0 to 1. A similarity matching threshold of 0.8 is set. When the similarity between the original field name and a certain standard field name is greater than or equal to 0.8, they are considered semantically matched; when the similarity is less than 0.8, a field difference is considered to exist. Fields with matching degrees below the threshold are marked as candidate difference fields, and their semantic correspondence is subsequently confirmed through manual review. Once confirmed, a set of field differences is formed.

[0064] For sets of field differences, naming rule adaptation is performed based on a preset conversion protocol. The preset conversion protocol maps original field names from heterogeneous data sources to standard field names in a preset standard dictionary. Its construction process involves creating an independent field mapping configuration table for each connected supplier, logistics provider, or manufacturer. This configuration table includes fields such as source system identifier, original field name, original field data type, standard field identifier, mapping rule type, creation time, and effective status. For newly connected data sources, the field naming list provided by their business system is parsed, and the similarity between each original field name and the standard field names in the preset standard dictionary is calculated. The similarity calculation method is consistent with the aforementioned matching degree calculation.

[0065] When the similarity between the original field name and a standard field name is greater than or equal to 0.8, a mapping candidate record for that field is automatically generated; when the similarity is less than 0.8, it is marked as requiring manual confirmation. Mapping rule types include direct mapping and transformation mapping. Direct mapping indicates that the original field name and the standard field name have the same semantics and data type, requiring no additional processing; transformation mapping indicates that the original field name and the standard field name have the same semantics but differ in data type or format, requiring the use of a transformation function. After the field mapping configuration table is established, a supply chain data governance specialist will manually review it to confirm the correctness of the automatically generated mapping candidate records and manually mark fields with similarity below the threshold.

[0066] After approval, the mapping record's status is set to enabled. If multiple candidate mapping records exist for the same original field name, the one with the highest similarity is selected first. If multiple candidate mapping records have the same similarity, the one with the most recent creation time is selected first. Based on the enabled mapping record, the fields in the field difference set are renamed to obtain the data block to be verified.

[0067] The pre-defined conversion protocol also includes a data format conversion rule base, stored in the supply chain collaboration database, used to handle format differences between raw field values ​​and standard fields. The rule base includes fields such as source format type, target format type, conversion function identifier, and parameter configuration. Source and target format types include strings, integers, floating-point numbers, date / time, timestamps, and enumerated values. The conversion function identifier corresponds to pre-defined conversion functions in the system, such as date format conversion functions, timestamp-to-date functions, enumerated value mapping functions, and unit conversion functions. Parameter configuration stores the specific parameters required by the conversion functions; for example, date format conversion functions require configuring source and target date formats, and enumerated value mapping functions require configuring a mapping table between source and target enumerated values. The initial entries in the data format conversion rule base are configured by the supply chain data governance specialist according to the data format specifications of the access system. Subsequently, for each new data source, the corresponding conversion rules are supplemented according to the interface documentation of that data source.

[0068] The data block to be validated undergoes quality inspection by traversing each data field, counting the number of missing values, and calculating the missing percentage. A preset missing threshold of 5% is set. This threshold is determined based on the integrity requirements of core business data in the supply chain, allowing for a small margin of error due to data fluctuations while ensuring the validity of the test data. If the missing percentage does not exceed 5%, a high-quality data sequence is obtained directly; if the missing percentage exceeds 5%, an imputation algorithm is initiated to complete the data. Time series linear interpolation is used to weighted imputation of missing values. The three adjacent valid values ​​before and after the missing value are extracted, and weights are assigned according to the principle that the closer to the missing position, the greater the weight. The weight of the first and third adjacent data is 0.3, the weight of the second and third adjacent data is 0.15, and the weight of the third and fourth adjacent data is 0.05. This set of weights is an example value set based on historical data distribution and can be adjusted according to data fluctuation characteristics in actual applications. The valid values ​​are multiplied by their corresponding weights and summed to obtain the imputed value, which is then inserted into the missing position to obtain a continuous and complete high-quality data sequence.

[0069] Based on the completed high-quality data sequence, a one-to-one correspondence is established between the original field names and the standard field names. The mapping relationship is stored in the form of a structured data table, resulting in a formatted mapping table. This mapping table includes fields such as the original field name, the system to which the original field belongs, the standard field identifier, the mapping rule type, the transformation function identifier, the creation time, and the effective status. It is used in subsequent steps to uniformly parse data from different data sources.

[0070] In step S13, supplier data is acquired, and data update records are extracted from the supplier data according to the formatted mapping table. If the data update record involves a change in delivery time, a time adjustment flag is generated, including:

[0071] Obtain supplier data, and extract data update records from the supplier data according to the formatted mapping table;

[0072] The logistics node identifier is obtained by parsing the business semantics contained in the data update record;

[0073] For the logistics node identifier, compare the current delivery time in the data update record with the preset delivery time to determine the time offset;

[0074] If the time offset exceeds the preset offset threshold, the data update record is adjusted according to the preset tag generation rule to obtain the time adjustment tag.

[0075] Specifically, the formatted mapping table built by S12 extracts updated records from the real-time data stream of automotive transmission suppliers. For example, the original field ETA_T is mapped to the standard field Arrival_Timestamp, corresponding to data such as 2026-03-25 14:00:00 and 2026-03-25 14:10:00; the original field Transport Temperature is mapped to the standard field Transport_Temperature, corresponding to data such as 25.05℃. This process ensures that supplier data from different underlying architectures can be parsed uniformly.

[0076] In this embodiment, in-depth business semantic analysis is performed on the update records to extract key logistics node identifiers. For example, taking the real-time update records of an automotive transmission supplier as an example, the logistics-related fields in the standardized update records are displayed as "Origin warehouse: Shanghai Jiading Auto Parts Warehouse, Transit hub: Suzhou Kunshan Auto Parts Distribution Center, Destination distribution warehouse: Nanjing Jiangning Vehicle Factory Supporting Warehouse". Through business semantic analysis, based on the standardized data structure, target fields strongly related to logistics nodes, such as origin warehouse, transit hub, destination distribution warehouse, and transportation route stations, are accurately selected from the update records. Subsequently, the text content within these fields is cleaned and denoised. The process involves removing modifiers such as "to," "via," and "located in," as well as punctuation marks and meaningless redundant characters, retaining only the core text of locations and logistics facilities. Then, a named entity recognition algorithm is used to extract administrative region names and logistics industry keywords, such as distribution centers, sorting warehouses, transit warehouses, and supporting logistics parks. Next, the extracted keyword combinations are precisely matched and semantically associated with a pre-set standardized logistics node identifier library in the supply chain to determine the geographical area, business type, and hierarchical affiliation of the node. Scattered small storage locations are filtered out, and related nodes belonging to the same core hub are merged. Finally, semantic classification and standardized mapping are completed to extract key logistics node information with unified identifiers.

[0077] For example, the core node keywords extracted from this field are Shanghai Jiading Auto Parts Warehouse, Suzhou Kunshan Auto Parts Distribution Center, and Nanjing Jiangning Vehicle Factory Supporting Warehouse. These keywords are then matched against a pre-defined logistics node identifier database. It is determined that the Suzhou Kunshan Auto Parts Distribution Center, Shanghai Jiading Auto Parts Warehouse, and Nanjing Jiangning Vehicle Factory Supporting Warehouse all belong to the core logistics node system of East China. Therefore, the key logistics node identifier corresponding to this update record is extracted as East China Auto Parts Distribution Center. Real-time monitoring of this node allows for the tracking of the flow status of specific batches of gear components.

[0078] For example, when the latest status of a batch of core components is obtained, the system automatically compares the current delivery time in the update record with the planned delivery time stored in the database. Suppose the original delivery time was 09:00 on November 12, but the estimated arrival time in real time is changed to 14:30 on November 12 due to weather conditions, the time offset is calculated to be 5.5 hours.

[0079] At this point, the offset is compared with the preset 2-hour offset threshold. Automobile manufacturers employ a just-in-time production model. As a core component, the transmission's time deviation within 2 hours during transportation is considered normal and reasonable fluctuation, encompassing scenarios such as loading / unloading delays, short-distance traffic congestion, and routine road condition adjustments. This deviation will not affect the production scheduling and operation of the vehicle assembly line and is considered a compatible and normal deviation. However, a time deviation exceeding 2 hours may lead to a disruption in the supply of core components, potentially causing a production line shutdown. Therefore, 2 hours is used as the offset threshold. Since 5.5 hours far exceeds the threshold range, this record is immediately determined to involve a significant change in delivery time at a logistics node.

[0080] In this embodiment, for records determined to have been changed, a time adjustment tag containing risk level and time dimension is automatically generated according to the built-in tag generation rules. The risk level is automatically generated according to standardized rules based on time offset and fixed-format splicing. The risk level threshold determination rules are as follows: with a preset 2-hour threshold as the base threshold, three risk levels are divided according to the time offset duration. The offset duration is the difference between the actual arrival time and the standard estimated arrival time. If the offset duration is ≤2 hours, no tag is generated. If it exceeds, the level is determined according to the following rules: Level 1 delay warning (2 hours < offset duration ≤ 4 hours) is a minor logistics delay, usually caused by short-term congestion, loading and unloading queues, and minor weather fluctuations. It will only slightly affect the pace of parts warehousing and will not directly threaten the normal production line schedule. Only the grassroots logistics personnel need to follow up and verify, and there is no need to initiate emergency adjustments. Level 2 delay warning (4 hours < offset duration ≤ 8 hours) is a moderate to major delay. This type of delay exceeds the normal logistics fluctuation range, will disrupt the parts preparation plan, and there is a risk of material shortage. It is necessary to notify the purchasing and logistics supervisors for coordination and handling. Level 3 delay warning (offset duration > 8 hours) is a serious high-risk delay. It is highly likely to cause the supply of core parts to be interrupted, directly causing the vehicle production line to stop. For example, a time adjustment flag is generated to indicate the 5.5-hour delay mentioned above.

[0081] Subsequently, the time adjustment marker is stored in the supply chain collaboration database to ensure data persistence and traceability, and is synchronized to the logistics scheduling end in real time via a message queue. Upon receiving the marker, the scheduling end immediately triggers the delivery time change response mechanism for that logistics node. Regarding the aforementioned delay in the transmission component, the response mechanism does not merely send a notification, but rather coordinates with the production schedule. The affected vehicle identification number (VIN) sequences, totaling 50 vehicles awaiting assembly, are retrieved. The transmission installation process for these vehicles, originally scheduled for 13:00 that day, is uniformly postponed to 15:00, and priority is given to scheduling alternative models with sufficient transmission inventory for production. This dynamic adjustment avoids the entire assembly line being shut down due to a single component delay, maximizing capacity utilization. Simultaneously, based on the changed arrival time of 14:30, the warehousing department releases the unloading platform resources originally scheduled for 09:00 for other suppliers, effectively improving the turnover efficiency of the logistics park.

[0082] In step S14, downstream production data is acquired, and the data is prioritized using the time adjustment marker to obtain a priority push list, including:

[0083] Acquire data from downstream production processes and adjust the markers based on the time;

[0084] The delay duration in the time adjustment mark is mapped to a production sequence displacement vector, and the downstream production data is prioritized according to the magnitude of the production sequence displacement vector to obtain a priority classification.

[0085] Based on the priority classification, the delivery window corresponding to the business plan is matched to obtain a dynamic plan association dataset;

[0086] Obtain the resource usage fluctuation value from the dynamic plan associated dataset;

[0087] If the resource usage fluctuation value exceeds the preset fluctuation value threshold, the current real-time load of the supplier is obtained. If the current real-time load meets the preset scheduling conditions, a priority push list is listed based on the dynamic plan association data.

[0088] In this embodiment, the actual delay duration recorded in the time adjustment marker is first obtained. This delay duration, in minutes, represents the amount of delay between the actual arrival time and the planned arrival time of the material. The original delay duration is processed according to a preset effective delay duration correction rule, which is used to eliminate tolerable delays within the normal fluctuation range of the supply chain. The preset normal fluctuation buffer duration is 30 minutes. This value is calculated based on the 90th percentile of delay events that did not cause production interruptions in historical transportation data. If the original delay duration is less than or equal to 30 minutes, it is determined to be a normal fluctuation, and the effective delay duration is set to 0, indicating that no production sequence adjustment needs to be triggered. If the original delay duration is greater than 30 minutes, 30 minutes are subtracted from the original delay duration to obtain the effective delay duration.

[0089] The system retrieves preset production cycle time parameters from downstream production processes. These parameters characterize the standard time interval between two adjacent production units on the production line. The production cycle time parameters are stored in a production cycle time parameter table, measured in minutes per unit. This table includes fields such as production line identifier, vehicle model identifier, standard cycle time, effective time, and expiration time. In the automotive manufacturing scenario, for model A on the final assembly line, the production cycle time parameter is set to 1 minute per unit. This value is derived from the production line design standards and comprehensively considers workstation operation time, equipment operating speed, and personnel operating efficiency.

[0090] The displacement step size is calculated by dividing the effective delay duration by the production cycle time parameter. This displacement step size represents the number of positions that the production sequence needs to move backward. When the effective delay duration and the production cycle time parameter are not divisible by each other, the rounding rule is applied, that is, the displacement step size is equal to the integer value of the quotient obtained by dividing the effective delay duration by the production cycle time parameter and rounding it up.

[0091] Retrieve the current production sequence status data, which is stored in the production sequence management database. The production sequence management database records the ranking information of all production units within the current production planning cycle in the form of structured data tables, including fields such as production sequence number, bill of materials identifier, planned start time, planned end time, current ranking order, associated supplier identifier, and associated logistics batch number. The current ranking order field uses integer encoding, starting from 1 and incrementing sequentially, indicating the actual position of the production unit in the final assembly queue. Simultaneously, retrieve the material-production sequence binding relationship table, which contains fields such as material code, production sequence number, material arrival status, and planned usage time.

[0092] Based on the displacement step size and the material-production sequence binding table, the affected production unit range is located. The minimum current ranking order of the production units associated with the delayed material is obtained and denoted as the initial ranking order. All production units in the current production sequence whose ranking order is greater than or equal to the initial ranking order are identified as the affected production unit set. The current ranking order of each production unit in this set is added to the displacement step size to obtain the adjusted ranking order. The difference between the original ranking order and the adjusted ranking order is determined as the production sequence displacement vector. This displacement vector is stored in a structured data format in a displacement mapping table, which includes fields such as affected production unit identifier, original ranking order, adjusted ranking order, displacement step size, and mapping timestamp.

[0093] Based on the production sequence displacement vector, downstream production data is prioritized. Delay duration is mapped to the timeliness level of material demand, assigning the highest priority to materials consumed immediately at the final assembly line and a lower priority to general-purpose materials with high safety stock.

[0094] Based on the priority ranking results, the priority levels are mapped to delivery window weights, with the highest priority corresponding to a weight coefficient of 1.0, medium priority to 0.6, and low priority to 0.3. Ordered by weight coefficient from highest to lowest, available delivery windows within the same time interval in the business plan are matched sequentially, prioritizing the allocation of production demands with the highest weight coefficients to delivery windows, thus generating a dynamic plan association dataset.

[0095] Obtain the planned usage and available quantity of each type of resource in the dynamic planning association dataset. Calculate the demand change rate for each type of resource, which is the difference between planned usage and available quantity divided by the available quantity. Multiply the demand change rate of each type of resource by its corresponding weight coefficient and sum them to obtain the overall resource utilization fluctuation value. The weight coefficients are configured as follows: Determine the key resource types involved in the production plan, including material resources, production capacity resources, warehousing resources, and transportation resources. In the automotive manufacturing scenario, the weight coefficient for core components in material resources is set to 0.5, the weight coefficient for key processes in production capacity resources is set to 0.3, the weight coefficient for line-side storage locations in warehousing resources is set to 0.1, and the weight coefficient for in-plant logistics distribution in transportation resources is set to 0.1. The sum of all weight coefficients is 1.

[0096] The preset fluctuation threshold is used to determine whether changes in resource demand exceed the normal adjustment capacity of the supply chain. It is set based on the steady-state fluctuation boundary in the historical operation data of the supply chain. The comprehensive resource usage fluctuation value generated by each plan fine-tuning in the past twelve months without triggering external collaborative adjustments is collected, and the 85th percentile is taken as the benchmark threshold. In the automotive manufacturing scenario, the preset fluctuation threshold is set to 0.25.

[0097] If the overall resource utilization fluctuation exceeds a preset fluctuation threshold, the supplier's current real-time load is obtained through a preset supplier node status monitoring interface. The load data includes the supplier's production line occupancy rate, available finished goods inventory, and schedulable capacity. Preset scheduling conditions are used to determine whether the supplier has the capacity to accept plan adjustments. Specifically, the rules are: when the supplier's production line occupancy rate is lower than the preset load threshold, the available finished goods inventory is greater than or equal to the incremental demand, and the remaining capacity calculated per hour multiplied by the remaining production time is greater than or equal to the incremental demand, the preset scheduling conditions are met. The preset load threshold is set to 85% based on the supplier's production line design redundancy capacity.

[0098] When the supplier's current real-time load meets the preset scheduling conditions, a priority delivery list is generated based on dynamic planning correlation data. This list includes the triggering event, priority level, material code, required quantity, original delivery order, and adjusted order.

[0099] In step S15, material codes are extracted from the priority push list and mapped to supplier IDs. Customized scheduling notifications are sent to suppliers based on the supplier IDs, and the supplier's response time to the customized scheduling notifications is obtained, resulting in response time data, including:

[0100] Extract material codes from the priority push list and map the material codes to supplier IDs;

[0101] The system detects the service load of the supplier. If the service load is lower than a preset service load threshold, it determines the push channel and matches the transmission protocol based on the push channel.

[0102] The supply chain data stream is obtained according to the priority push list, the supply chain data stream is encapsulated, and customized push content is generated by combining it with a preset content template.

[0103] The customized push content is sent to the supplier according to the transmission protocol, and the response time data after sending is tracked.

[0104] First, based on the material codes in the priority push list, the corresponding supplier ID is obtained by querying the material-supplier association mapping table. The material-supplier association mapping table is stored in the supply chain collaboration database and includes fields such as material code, supplier ID, supplier material code, and whether it is a primary supplier. Basic information such as the supplier name is obtained by querying the supplier master data table using the supplier ID. The supplier master data table includes fields such as supplier ID, supplier name, contact information, and address.

[0105] The supplier's notification preference configuration is retrieved based on their supplier ID. The notification preference configuration table is stored in the supply chain collaboration database and includes fields such as supplier ID, preferred notification channel, backup notification channel, notification content language, recipient contact person, recipient email address, recipient API interface address, and whether real-time push is enabled. The preferred notification channel includes three types: API interface, email, and SMS. The API interface is the default preferred channel for real-time scheduling scenarios; email is used for non-urgent batch notifications; and SMS is used for supplementary notifications in emergency situations. The supplier notification preference configuration table is configured by the supply chain collaboration administrator when the supplier connects to the supply chain collaboration platform. Subsequent adjustments can be made according to supplier needs, and configuration changes require confirmation from both parties to take effect.

[0106] Based on the preferred notification channel in the notification preference configuration table, the corresponding transport protocol is matched. The API interface channel uses Hypertext Transfer Protocol version 2 (HTTP / 2) and encapsulates data in JSON format; the email channel uses Simple Mail Transfer Protocol (SMTP) and encapsulates content in HTML format; the SMS channel uses Short Message Protocol for Point-to-Point Communication (SMPP) and encapsulates content in plain text format. If the preferred channel fails to send or times out, the matching and sending process is repeated using the alternative notification channel.

[0107] The generation of customized scheduling notification content is accomplished through a pre-defined notification content template library and dynamic data population rules. The notification content template library is stored in the supply chain collaboration database and includes fields such as template identifier, template name, applicable scenario, applicable notification channel, template content, template version, and effective time. Applicable scenarios include priority scheduling, material requirement changes, delivery time adjustments, and capacity warnings. Template content is written using placeholders, with the placeholder format being a variable name enclosed in double curly braces.

[0108] The initial version of the preset notification content template library contains 12 templates, covering various combinations of scheduling scenarios and notification channels. Among them, the priority scheduling template content for API interface channels is "Scheduling notification, material code {{material code}}, required quantity {{required quantity}}, latest delivery time {{latest delivery time}}, priority level {{priority level}}, please confirm the response".

[0109] The priority scheduling template for the email channel reads: "Dear {{Supplier Name}}, due to production plan adjustments, the required quantity of your supplied material {{Material Code}} has changed to {{Required Quantity}}, the latest delivery time is {{Latest Delivery Time}}, and the priority level is {{Priority Level}}. Please confirm your response within {{Confirmation Deadline}}. Thank you for your cooperation."

[0110] The priority scheduling template for the SMS channel is as follows: "{{Material Code}} Requirement {{Required Quantity}}, Latest {{Latest Delivery Time}}, Priority {{Priority Level}}, Please reply to confirm".

[0111] The template library is reviewed and updated quarterly by the supply chain collaboration administrator based on business feedback. New templates must be tested and verified before being launched.

[0112] Dynamic data population rules are used to map business data from the priority push list to template placeholders. The dynamic data population rule table is stored in the supply chain collaboration database and includes fields such as rule identifier, placeholder name, data source, data source field, and transformation function identifier. Data sources include the priority push list, material-supplier association mapping table, supplier master data table, and production plan table. When the placeholder data source is the supplier master data table, the corresponding data is retrieved through a relational query using the supplier ID.

[0113] For example, the placeholder "{{Supplier Name}}" originates from the supplier master data table, with the supplier name as the data source field; the placeholder "{{Material Code}}" originates from the priority delivery list, with the material code as the data source field; the placeholder "{{Required Quantity}}" originates from the priority delivery list, with the required quantity as the data source field; and the placeholder "{{Latest Delivery Time}}" originates from the production plan table, with the planned start time as the data source field. The conversion function is identified as "Timestamp to Date and Time," and the parameter configuration is "yyyy-MM-dd". The format is "HH:mm"; the data source of the placeholder "{{priority level}}" is the priority push list, and the data source field is the priority classification result. The priority classification result includes three enumeration values: high priority, medium priority, and low priority, which correspond to the text output of "high", "medium", and "low" respectively; the data source of the placeholder "{{confirmation time limit}}" is the currently effective dynamic latency threshold. This dynamic latency threshold is generated in the previous scheduling in step S16. The initial value is preset to 300 milliseconds. This initial value is set based on the average value of the actual idle state response time of the supply chain collaboration system in the production environment. The conversion function is identified as "millisecond to minute", and the parameter configuration is to divide the millisecond value by 60000 and then round it down.

[0114] When generating customized scheduling notification content, the first step is to select a corresponding template from the notification content template library based on the applicable scenario and notification channel. The applicable scenario is determined by the business context of the priority push list. For example, when the reason for generating the list is production plan replacement, the matching applicable scenario is "priority scheduling". After selecting a template, the placeholder list in the template content is parsed. Each placeholder is traversed, and the data source and source field corresponding to the placeholder are queried according to the dynamic data filling rule table. The current business data is extracted from the corresponding data table. If a conversion function identifier exists, the corresponding conversion function is called to convert the data format. The converted value replaces the placeholder in the template. After all placeholders are replaced, the customized push content is obtained.

[0115] For API interface channels, after generating customized push content, the push content is encapsulated into a JSON format data packet. The data packet structure includes fields such as request header identifier, business serial number, sending timestamp, and push content body. The business serial number is generated in the format of "year, month, day, hour, minute, second plus 6 random numbers", such as "20260331143025123456", and is used to uniquely identify this push for easy subsequent tracking and logging.

[0116] Customized push content is sent to the supplier according to the transmission protocol, and the sending timestamp is recorded. A response monitoring timer is started, and the timeout duration is set to the currently effective dynamic latency threshold. The response confirmation signal returned by the supplier is monitored. The response confirmation signal includes fields such as business serial number, supplier ID, confirmation status, and confirmation timestamp. The time interval from sending to receiving the confirmation signal is recorded as the response time and stored in the push response record table. This data table includes fields such as business serial number, supplier ID, sending timestamp, response timestamp, response time, and whether it has timed out. If no confirmation signal is received within the timeout period, the push is marked as timed out, the timeout field is set to yes, and the response time field is set to negative one, indicating that there was no response after the timeout, which is used by step S16 to trigger the second push process.

[0117] In step S16, the node load status of the supplier is obtained, and a dynamic latency threshold is generated by combining the real-time latency value in the response time data; if the real-time latency value is greater than the dynamic latency threshold, the customized scheduling notification content is sent to the supplier again, and the dynamic latency threshold is corrected based on the feedback from the second push, determining the collaborative response parameters, including:

[0118] Obtain the node load status of the supplier and combine it with the real-time time consumption value in the response time data to generate a dynamic latency threshold;

[0119] If the real-time consumption value is greater than the dynamic latency threshold, a secondary push process is triggered;

[0120] Obtain the network latency fluctuation frequency of the supplier, and determine the interval duration of the secondary push process based on the network latency fluctuation frequency;

[0121] Obtain the secondary push feedback generated by the secondary push process, correct the dynamic latency threshold based on the secondary push feedback, and determine the collaborative response parameters.

[0122] In this embodiment, to address the timeliness requirements in supply chain collaboration, a deep analysis is conducted on the nonlinear relationship between real-time latency and node load. Based on the real-time load of the receiving end, such as the production scheduling system of a seat supplier, a dynamic latency benchmark is generated through dynamic adjustment rules for latency thresholds.

[0123] First, multi-dimensional load data of the supplier node is obtained, including CPU utilization, memory utilization, average network latency, network latency fluctuation, historical average response time, and response timeout rate. CPU utilization and memory utilization are expressed as percentages, ranging from 0 to 100. Average network latency, in milliseconds, represents the average round-trip time for communication with the supplier node over the past minute. Network latency fluctuation, in milliseconds, represents the difference between the maximum and minimum network latency over the past minute. Historical average response time, in milliseconds, represents the average time taken for the supplier node to respond to scheduling notifications over the past five minutes. Response timeout rate, expressed as a percentage, represents the proportion of scheduling notifications that timed out of the total number of notifications sent over the past five minutes.

[0124] The collected multi-dimensional load data is smoothed and filtered using a moving average filtering method. The window size is set to 5 sampling points, and the step size is set to 1 sampling point. That is, after each new load data is collected, the current data is added to the previous 4 historical data and then divided by 5 to obtain the filtered load value. This filtering method is used to eliminate the impact of instantaneous network jitter or acquisition anomalies on load judgment and ensure the stability of the dynamic latency threshold generation.

[0125] The filtered load data for each dimension is normalized, mapping the values ​​of each dimension to a range of 0 to 1. CPU utilization and memory utilization are normalized by directly dividing by 100. The mean network latency is normalized using an upper limit truncation method, with a reference upper limit of 1000 milliseconds. If the original value is greater than 1000 milliseconds, the normalized value is 1; otherwise, the normalized value is equal to the original value divided by 1000. Network latency fluctuation is normalized using an upper limit truncation method, with a reference upper limit of 800 milliseconds. If the original value is greater than 800 milliseconds, the normalized value is 1; otherwise, the normalized value is equal to the original value divided by 800. The historical average response time is normalized using an upper limit truncation method, with a reference upper limit of 2000 milliseconds. If the original value is greater than 2000 milliseconds, the normalized value is 1; otherwise, the normalized value is equal to the original value divided by 2000. The response timeout rate is normalized by dividing by 100. The normalized value is equal to the original percentage value divided by 100. The upper limit reference values ​​for the above dimensions are set based on the load data distribution characteristics collected in the early stage of the supply chain collaboration system operation. Specifically, the value is the 95th percentile of the historical data. It is recalculated and updated every quarter after six months of operation.

[0126] Weighting coefficients are assigned to each dimension of load data, with the sum of the weighting coefficients being 1. These coefficients reflect the degree of impact of different load dimensions on latency thresholds. The weighting coefficients are set as follows: CPU utilization rate 0.25, memory utilization rate 0.15, average network latency 0.2, network latency fluctuation amplitude 0.2, historical average response time 0.1, and response timeout rate 0.1. The weighting coefficients are based on multi-factor regression analysis of historical timeout events. All events triggering secondary pushes due to response timeouts within the past three months are collected. Logistic regression is performed with load data of each dimension as independent variables and timeout status as the dependent variable. The regression coefficients are normalized and used as the initial values ​​for the weighting coefficients. Weights are calibrated quarterly based on newly added data.

[0127] The normalized load values ​​of each dimension are multiplied by their corresponding weight coefficients and then summed to obtain the comprehensive load index. The comprehensive load index also ranges from 0 to 1. The closer the value is to 1, the heavier the node load is, and the closer it is to 0, the lighter the node load is.

[0128] Obtain a preset base latency threshold, which represents the expected response time when the node is completely idle. This threshold is set to 300 milliseconds and is derived from actual test data of the supply chain collaboration system in a production environment, taking the average response time of all nodes during the early morning off-peak business period. Obtain a preset expansion coefficient, which is used to convert the overall load index into a latency increment. This coefficient is set to 800 milliseconds and is calculated by subtracting the base latency threshold from the maximum tolerable response time when the node is fully loaded. That is, the tolerable response time at full load is 1100 milliseconds, and subtracting the base latency threshold of 300 milliseconds yields the expansion coefficient of 800 milliseconds.

[0129] The dynamic latency threshold is obtained by adding the product of the base latency threshold and the comprehensive load index multiplied by the expansion factor. When the comprehensive load index is 0, the dynamic latency threshold is equal to the base latency threshold of 300 milliseconds; when the comprehensive load index is 0.5, the dynamic latency threshold is equal to 300 milliseconds plus 0.5 multiplied by 800 milliseconds, which is 700 milliseconds; when the comprehensive load index is 1, the dynamic latency threshold is equal to 300 milliseconds plus 1 multiplied by 800 milliseconds, which is 1100 milliseconds.

[0130] The above parameters are stored in a dynamic latency configuration table. This table includes fields such as supplier node identifier, basic latency threshold, expansion coefficient, load weight coefficients for each dimension, normalized upper limit reference values ​​for each dimension, filter window size, filter step size, effective time, and expiration time. Different supplier nodes can configure differentiated parameters according to their network environment characteristics. For example, supplier nodes deployed in cloud data centers have higher network stability, so the weight coefficient of network latency fluctuation can be appropriately reduced and the basic latency threshold can be narrowed; supplier nodes deployed in local factory server rooms have a relatively complex network environment, so the weight coefficient of network latency fluctuation can be appropriately increased and the expansion coefficient can be widened. The dynamic latency configuration table is maintained through the supply chain collaboration management interface. Configuration changes must be approved before taking effect, and all change records are written to the configuration audit log.

[0131] The weighting coefficients and normalization upper limit reference values ​​used in the comprehensive load index calculation are updated quarterly through an automated training process. The training process includes the following steps: Extracting all scheduling notification sending records from the log database over the past three months. Each record includes the sending time, supplier node identifier, raw load data for each dimension at the time of sending, response time, and a tag indicating whether a secondary push was triggered. Abnormal records with response times exceeding a preset reasonable range are removed; the upper limit of the reasonable range is set at 5000 milliseconds. The load data for each dimension is processed using the aforementioned normalization method to obtain a normalized feature vector. Using whether a secondary push was triggered as the dependent variable and the normalized features for each dimension as independent variables, a logistic regression algorithm is used to fit the model, obtaining regression coefficients for each feature. The absolute values ​​of each regression coefficient are taken and normalized so that the sum of the coefficients is 1, resulting in new weighting coefficients. Simultaneously, the 95th percentile of the load data for each dimension among records that did not trigger a secondary push is calculated and used as the new normalization upper limit reference value. After the training process is completed, the updated weighting coefficients and normalization upper limit reference values ​​are written to the dynamic latency configuration table and take effect from the next statistical period.

[0132] During real-time operation, after each customized scheduling notification is sent to the supplier, the real-time latency is monitored and recorded. This real-time latency is compared with the currently effective dynamic latency threshold. If the real-time latency is less than or equal to the dynamic latency threshold, the push is considered to have been completed successfully; if the real-time latency is greater than the dynamic latency threshold, the push is considered to have timed out, triggering a second push process. Understandably, after the second push is completed, feedback is obtained, including the transmission result and the actual response latency data. Based on this feedback, the collaborative response configuration is adjusted. For example, if it is found that a supplier's response is consistently slightly slower than expected within a certain time period (e.g., the average deviation between the actual latency and the preset latency for this type of supplier during periods of slow response is 10%–20%), the collaborative response parameters will be updated, increasing the preset waiting time within that time window by 15%. Through this feedback-based closed-loop optimization, communication strategies with different collaborating partners can be adaptively adjusted, thereby significantly improving the robustness and intelligence of supply chain collaboration while ensuring data delivery rates.

[0133] In step S17, the push log of the supply chain data stream is updated according to the collaboration response parameters. The updated supply chain data stream segments are extracted from the push log. If the supply chain data stream segments meet preset grouping requirements, a business collaboration adjustment plan is output, including:

[0134] Based on the aforementioned collaborative response parameters, a differentiated push strategy and push content customization rules are constructed;

[0135] Execute the differentiated push strategy and the push content customization rules, and write the feedback interaction data into the push log record;

[0136] The updated supply chain data stream segment is extracted based on the push log record. The supply chain data stream segment is then compared with the preset supplier requirements. If the supply chain data stream segment meets the supplier requirements, a business collaboration adjustment plan is output.

[0137] Obtain the optimized collaboration response parameters, and construct a differentiated push strategy and push content customization rules based on the collaboration response parameters;

[0138] Execute the differentiated push strategy and the push content customization rules, and write the interaction data generated by the feedback closed-loop mechanism into the push log record;

[0139] The push log records are parsed to extract updated supply chain data stream fragments, and the supply chain data stream fragments are compared with the preset collaborative object group requirements for feature comparison.

[0140] If the supply chain data stream segment meets the grouping requirements of the collaborating objects, a closed-loop verification result is generated, and the final business collaboration adjustment plan is output based on the closed-loop verification result.

[0141] Specifically, the system first reads four core quantitative indicators—network fluctuation amplitude, dynamic latency benchmark, and 1-minute average latency—encapsulated in the optimized S16 collaboration response parameters. These indicators, combined with preset thresholds, determine the network status of the supplier node. If the parameters show a network fluctuation amplitude ≥ 800 milliseconds, a dynamic latency benchmark ≥ 800 milliseconds, and a 1-minute average network latency ≥ 500 milliseconds, then the supplier node is considered to be in a high-latency fluctuation state. If these parameters indicate that a specific supplier node's network environment is in a high-latency fluctuation state, a differentiated push strategy is then constructed.

[0142] For example, for seat fabric suppliers marked as having high latency fluctuations, the default real-time full data synchronization mode is no longer used. Instead, a customized push content rule prioritizing key fields is adopted. For instance, the original complete bill of materials containing 50 fields is streamlined into a lightweight data packet containing only three core fields: material code, required quantity, and latest delivery time. The data size is compressed from 50KB to less than 2KB, thereby greatly improving the success rate of data delivery under limited network bandwidth. After implementing the above strategy, each handshake signal, data packet transmission status, and acknowledgment receipt from the receiving end are recorded as interaction data and written to the push log in real time.

[0143] Next, the log parsing engine is activated to extract updated supply chain data stream fragments from massive amounts of log data. For example, if the log shows that the supplier successfully received a lightweight data packet and returned an acknowledgment code at 14:05, this acknowledgment code and its corresponding timestamp are the extracted data stream fragments. These fragments are then compared against the requirements of predefined collaboration groups. For instance, if the supplier belongs to the JIT (Just-In-Time) production group, this group's requirements stipulate that all material changes must be confirmed within 30 minutes before the production schedule is locked.

[0144] If the extracted data stream segment shows that the confirmation time is earlier than the scheduling lock time, and the confirmation content is completely consistent with the core fields issued, it is determined that the supply chain data stream segment meets the grouping requirements, and thus a closed-loop verification result is generated.

[0145] When the closed-loop verification result is satisfactory, the final business collaboration plan will be officially implemented and output according to the previously proposed collaboration adjustment plan. Production execution instructions will be automatically issued to downstream processes, and the batch of materials will be included in the production schedule for the afternoon of the same day. Workstation binding, time allocation and resource locking will be completed, so that this change will officially enter the actual production execution stage.

[0146] In summary, this invention discloses a data standardization processing method for industrial supply chains, which improves the efficiency of supply chain data integration.

[0147] Reference Figure 2 The second embodiment of the present invention provides a data standardization processing system for industrial supply chains, comprising:

[0148] The heterogeneous data acquisition module is used to acquire raw heterogeneous data from suppliers, logistics nodes, and production processes in the supply chain, and uses the K-means clustering algorithm to group the raw heterogeneous data to obtain a set of format classifications.

[0149] The data standardization mapping module is used to complete the format classification set by using a preset standard dictionary to obtain a high-quality data sequence, and to establish the field mapping relationship based on the high-quality data sequence to obtain a formatted mapping table;

[0150] The delivery change marking module is used to extract data update records from the original heterogeneous data according to the formatted mapping table, and generate a time adjustment mark if the data update record involves a change in delivery time.

[0151] The production priority classification module is used to acquire data from downstream production processes, and to classify the data from downstream production processes into priority categories based on the time adjustment markers, thereby obtaining a priority push list.

[0152] The collaborative scheduling and push module is used to extract material codes from the priority push list, map the material codes to supplier IDs, send customized scheduling notification content to suppliers based on the supplier IDs, and obtain the time when the supplier responds to the customized scheduling notification content to obtain response time data.

[0153] The dynamic latency optimization module is used to obtain the node load status of the supplier and generate a dynamic latency threshold by combining the real-time time consumption value in the response time data; if the real-time time consumption value is greater than the dynamic latency threshold, the customized scheduling notification content is sent to the supplier again, and the dynamic latency threshold is corrected according to the feedback of the second push to determine the collaborative response parameters.

[0154] The closed-loop collaboration verification module is used to update the push log records of the supply chain data flow according to the collaboration response parameters, extract the updated supply chain data flow segments according to the push log records, and output a business collaboration adjustment plan if the supply chain data flow segments meet the preset grouping requirements.

[0155] It should be noted that the data standardization processing system for industrial supply chains provided in this embodiment of the invention is used to execute all the process steps of the data standardization processing method for industrial supply chains in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0156] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0157] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A data standardization processing method for industrial supply chains, characterized in that, include: The raw heterogeneous data of suppliers, logistics nodes and production links in the supply chain are obtained, and the raw heterogeneous data are grouped by K-means clustering algorithm to obtain a set of format classifications. The format classification set is completed by using a preset standard dictionary to obtain a high-quality data sequence. The mapping relationship of fields is established based on the high-quality data sequence to obtain a format mapping table. Based on the formatted mapping table, data update records are extracted from the original heterogeneous data. If the data update record involves a change in delivery time, a time adjustment flag is generated. Acquire downstream production data, prioritize the downstream production data using the time adjustment marker, and obtain a priority push list; Extract material codes from the priority push list, map the material codes to supplier IDs, send customized scheduling notifications to suppliers based on the supplier IDs, and obtain the time it takes for suppliers to respond to the customized scheduling notifications to obtain response time data. Obtain the node load status of the supplier and combine it with the real-time time consumption value in the response time data to generate a dynamic latency threshold; if the real-time time consumption value is greater than the dynamic latency threshold, send the customized scheduling notification content to the supplier again, and correct the dynamic latency threshold based on the feedback of the second push to determine the collaborative response parameters. Based on the collaboration response parameters, update the push log records of the supply chain data stream, extract the updated supply chain data stream segments based on the push log records, and if the supply chain data stream segments meet the preset grouping requirements, output the business collaboration adjustment plan.

2. The data standardization processing method for industrial supply chains according to claim 1, characterized in that, Obtain raw, heterogeneous data from suppliers, logistics nodes, and production processes within the supply chain, including: Data is collected from suppliers, logistics nodes, and production links in the supply chain to obtain the first heterogeneous data, which includes text, tables, and timestamps. The first heterogeneous data is subjected to mean interpolation to obtain the second heterogeneous data; Obtain the field names of the second heterogeneous data, vectorize the field names using a preset bag-of-words model to obtain field name vectors, calculate the cosine similarity of the field name vectors, and obtain the field name difference metric. When the field naming difference metric exceeds a preset difference threshold, feature extraction is performed on the second heterogeneous data to obtain the original heterogeneous data.

3. The data standardization processing method for industrial supply chains according to claim 1, characterized in that, The step of completing the format classification set using a preset standard dictionary to obtain a high-quality data sequence includes: The business semantics of the format classification set are analyzed, the matching degree between the business semantics and the preset standard dictionary is calculated, and the set of field differences is determined. Based on a preset conversion protocol, the set of field differences is renamed to obtain the data block to be verified; The percentage of missing data blocks to be verified is calculated to obtain the data integrity verification result; If the data integrity verification result shows that the proportion of missing values ​​exceeds the preset missing value threshold, the missing values ​​are filled with weighted data using time series linear interpolation to obtain a high-quality data sequence.

4. The data standardization processing method for industrial supply chains according to claim 1, characterized in that, If the data update record involves a change in delivery time, a time adjustment flag is generated, including: The logistics node identifier is obtained by parsing the business semantics contained in the data update record; For the logistics node identifier, compare the current delivery time in the data update record with the preset delivery time to determine the time offset; If the time offset exceeds the preset offset threshold, the data update record is adjusted according to the preset tag generation rule to obtain the time adjustment tag.

5. The data standardization processing method for industrial supply chains according to claim 1, characterized in that, The process of acquiring downstream production data, prioritizing the data using the time adjustment marker to obtain a priority push list, includes: The delay duration in the time adjustment mark is mapped to a production sequence displacement vector, and the downstream production data is prioritized according to the magnitude of the production sequence displacement vector to obtain a priority classification. Based on the downstream production data, a business plan is obtained; Based on the priority classification, the delivery window corresponding to the business plan is matched to obtain a dynamic plan association dataset; Obtain the resource usage fluctuation value from the dynamic plan associated dataset; If the resource usage fluctuation value exceeds the preset fluctuation value threshold, the current real-time load of the supplier is obtained. If the current real-time load meets the preset scheduling conditions, a priority push list is listed based on the dynamic plan association data.

6. The data standardization processing method for industrial supply chains according to claim 1, characterized in that, The step of sending customized scheduling notification content to the supplier and obtaining the supplier's response time to the customized scheduling notification content, resulting in response time data, includes: The system detects the service load of the supplier. If the service load is lower than a preset service load threshold, it determines the push channel and matches the transmission protocol based on the push channel. The supply chain data stream is obtained according to the priority push list, the supply chain data stream is encapsulated, and customized push content is generated by combining it with a preset content template. The customized push content is sent to the supplier according to the transmission protocol, and the response time data after sending is tracked.

7. The data standardization processing method for industrial supply chains according to claim 1, characterized in that, If the real-time latency value is greater than the dynamic latency threshold, the customized scheduling notification content is sent to the supplier again, and the dynamic latency threshold is adjusted based on the feedback from the second push to determine the collaborative response parameters, including: If the real-time consumption value is greater than the dynamic latency threshold, a secondary push process is triggered; Obtain the network latency fluctuation frequency of the supplier, and determine the interval duration of the secondary push process based on the network latency fluctuation frequency; Obtain the secondary push feedback generated by the secondary push process, correct the dynamic latency threshold based on the secondary push feedback, and determine the collaborative response parameters.

8. The data standardization processing method for industrial supply chains according to claim 1, characterized in that, The step of updating the push log records of the supply chain data stream according to the collaborative response parameters includes: Based on the aforementioned collaborative response parameters, a differentiated push strategy and push content customization rules are constructed; The differentiated push strategy and the push content customization rules are executed, and the feedback interaction data is written into the push log record.

9. A data standardization processing system for industrial supply chains, characterized in that, include: The heterogeneous data acquisition module is used to acquire raw heterogeneous data from suppliers, logistics nodes, and production processes in the supply chain, and uses the K-means clustering algorithm to group the raw heterogeneous data to obtain a set of format classifications. The data standardization mapping module is used to complete the format classification set by using a preset standard dictionary to obtain a high-quality data sequence, and to establish the field mapping relationship based on the high-quality data sequence to obtain a formatted mapping table; The delivery change marking module is used to extract data update records from the original heterogeneous data according to the formatted mapping table, and generate a time adjustment mark if the data update record involves a change in delivery time. The production priority classification module is used to acquire data from downstream production processes, and to classify the data from downstream production processes into priority categories based on the time adjustment markers, thereby obtaining a priority push list. The collaborative scheduling and push module is used to extract material codes from the priority push list, map the material codes to supplier IDs, send customized scheduling notification content to suppliers based on the supplier IDs, and obtain the time when the supplier responds to the customized scheduling notification content to obtain response time data. The dynamic latency optimization module is used to obtain the node load status of the supplier and generate a dynamic latency threshold by combining the real-time time consumption value in the response time data. If the real-time latency value is greater than the dynamic latency threshold, the customized scheduling notification content is sent to the supplier again, and the dynamic latency threshold is corrected based on the feedback from the second push to determine the collaborative response parameters. The closed-loop collaboration verification module is used to update the push log records of the supply chain data flow according to the collaboration response parameters, extract the updated supply chain data flow segments according to the push log records, and output a business collaboration adjustment plan if the supply chain data flow segments meet the preset grouping requirements.