A supply chain demand prediction method and system based on multi-source heterogeneous data fusion
By identifying and processing sudden noise in community group-buying orders, dynamically corrected sales data is generated and combined with regular sales data for forecasting. This solves the problem of demand forecast distortion caused by sudden order noise in traditional forecasting methods, and achieves efficient and accurate replenishment decisions in the supply chain.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENGTIAN BANZI GROUP CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional supply chain demand forecasting fails to effectively handle sudden order noise in community group buying scenarios, resulting in distorted demand forecasts and failing to provide a reliable basis for automatic replenishment decisions in the supply chain.
By collecting the location coordinates and sales value of community group buying orders in real time, we can identify sudden noise, calculate the order coverage radius level, match the noise intensity coefficient, generate dynamic attenuation coefficient and hysteresis fluctuation correction amount, and combine them with regular sales data to predict demand.
Accurately capture supply and demand trends, improve the accuracy of supply chain demand forecasting and the timeliness of replenishment decisions, avoid inventory backlog or demand gaps, and ensure the efficient operation of the supply chain.
Smart Images

Figure CN122133979A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of supply chain demand forecasting technology, and more specifically, to a supply chain demand forecasting method and system based on the fusion of multi-source heterogeneous data. Background Technology
[0002] Supply chain demand forecasting technology is a crucial technology, specifically applied in the supply chain demand forecasting stage of community group buying scenarios. Its core function is to improve forecast reliability by accurately handling sudden order noise, thus meeting the supply chain's core needs for accurate demand forecasting and timely replenishment decisions. In community group buying scenarios, local areas are prone to sudden surges in sales due to temporary promotions or emergency procurement. These orders vary in coverage and noise intensity, and sales naturally decline over time, exhibiting an instantaneous decay characteristic. Traditional supply chain demand forecasting does not specifically address this sudden order noise; it neither quantifies the decay pattern to weaken abnormal peaks nor corrects for time-related fluctuations. This leads to data distortion when sudden orders are integrated with regular sales data, e-commerce transaction data, and store inventory data, thus affecting the accuracy of demand forecasting results and failing to provide a reliable basis for automatic supply chain replenishment decisions. To address this technical problem, we provide a supply chain demand forecasting method and system based on multi-source heterogeneous data fusion. Summary of the Invention
[0003] The purpose of this invention is to provide a supply chain demand forecasting method and system based on multi-source heterogeneous data fusion to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, one of the objectives of this invention is to provide a supply chain demand forecasting method based on multi-source heterogeneous data fusion, comprising the following steps: S1. Real-time collection of location coordinates and sales volume of community group-buying orders; S2. By using a preset sales threshold, sudden noise is identified in community group buying orders. When a local sudden order event is identified, the order coverage radius is calculated and the order coverage radius level is divided based on the location coordinates of the community group buying order using geofencing technology. A preset noise intensity coefficient is matched according to the order coverage radius level. At the same time, the order coverage radius level and the preset noise intensity coefficient are combined to query a preset three-dimensional mapping table to obtain the hysteresis fluctuation correction amount. A dynamic attenuation coefficient is generated based on the instantaneous attenuation characteristics of the local sudden order. The sales value of the sudden order is synchronously applied with the dynamic attenuation coefficient and the hysteresis fluctuation correction amount to generate dynamically corrected sales data. S3. Integrate dynamically corrected sales data with regular sales data, input the forecasting model to output demand forecasting results, and execute the demand forecasting results to drive supply chain response.
[0005] The second objective of this invention is to provide a system for implementing a supply chain demand forecasting method based on multi-source heterogeneous data fusion, comprising: The data acquisition unit obtains the order location coordinates and sales value in real time through the community group buying platform interface, performs reverse geocoding on the text address, and generates a time-stamped order geographic dataset. The sudden noise processing unit identifies local sudden order events based on the rolling daily average sales benchmark and dynamic threshold multiple. It uses density clustering algorithm to calculate the order coverage radius and divides the discretized radius level identifier. It obtains the hysteresis fluctuation correction amount by mapping the radius level and noise intensity coefficient through a two-dimensional relationship matrix. It generates a dynamic attenuation coefficient by combining a negative exponential function. It applies the attenuation coefficient and time dimension offset to the sudden sales value simultaneously and outputs a dynamically corrected sales dataset. The prediction execution unit dynamically corrects the dataset and aligns it with e-commerce transaction data and store inventory data according to time stamps. After time window slicing and normalization, it is input into the recurrent neural network prediction model to generate regional demand forecasts for the next seven days. The response execution unit triggers automatic replenishment decisions based on demand forecasts and drives the supply chain to execute responses.
[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention accurately acquires the location and sales volume of community group-buying orders through a data acquisition unit. The data is then formatted using reverse geocoding and time-stamped to provide standardized data for subsequent processing. A sudden noise processing unit accurately identifies local sudden orders based on a rolling daily average sales benchmark. It calculates the coverage radius and levels the data using density clustering algorithms and the minimum inclusive circle, and matches the lag fluctuation correction amount using a two-dimensional relationship matrix. A dynamic decay coefficient is generated based on a negative exponential function to achieve numerical decay and time-axis correction of sudden sales, effectively eliminating noise interference. The prediction execution unit integrates the dynamically corrected data with e-commerce transaction and store inventory data. After time window slicing and normalization, the data is input into an LSTM model to generate regional demand forecasts for the next seven days, accurately capturing supply and demand trends. The response execution unit triggers automatic replenishment decisions, driving rapid supply chain response. This avoids inventory backlog and prevents demand gaps, improving the accuracy of supply chain demand forecasting and the timeliness of replenishment decisions, providing a reliable guarantee for the efficient operation of the supply chain in the community group-buying scenario. Attached Figure Description
[0007] Figure 1 This is a flowchart illustrating the overall workflow of the present invention; Figure 2 This is a schematic diagram of the overall structure of the present invention; The meanings of the labels in the diagram are as follows: 1. Data acquisition unit; 2. Sudden noise processing unit; 3. Prediction execution unit; 4. Response execution unit. Detailed Implementation
[0008] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0009] Please see Figure 1 As shown, one of the objectives of this embodiment is to provide a supply chain demand forecasting method based on multi-source heterogeneous data fusion, including the following steps: S1. Real-time collection of location coordinates and sales volume of community group-buying orders; S2. By using a preset sales threshold, sudden noise is identified in community group buying orders. When a local sudden order event is identified, the order coverage radius is calculated and the order coverage radius level is divided based on the location coordinates of the community group buying order using geofencing technology. A preset noise intensity coefficient is matched according to the order coverage radius level. At the same time, the order coverage radius level and the preset noise intensity coefficient are combined to query a preset three-dimensional mapping table to obtain the hysteresis fluctuation correction amount. A dynamic attenuation coefficient is generated based on the instantaneous attenuation characteristics of the local sudden order. The sales value of the sudden order is synchronously applied with the dynamic attenuation coefficient and the hysteresis fluctuation correction amount to generate dynamically corrected sales data. S3. Integrate dynamically corrected sales data with regular sales data, input the forecasting model to output demand forecasting results, and execute the demand forecasting results to drive supply chain response.
[0010] Real-time collection of location coordinates for community group-buying orders, specifically including: The platform uses the open data interface of the community group buying platform to obtain the original location information when the order is submitted in real time. When the original location information is latitude and longitude coordinates, it is recorded directly. When the original location information is a text address, it is converted into latitude and longitude coordinates through reverse geocoding service. The latitude and longitude coordinates are then bound to the sales value of the corresponding order to form a time-stamped order geographic dataset.
[0011] The system identifies sudden noise in community group-buying orders by setting a preset sales threshold, specifically including: The daily average sales benchmark is calculated based on historical community group buying data of the target area and is updated on a rolling basis. The flow of new orders is monitored in real time. When the cumulative sales of a single geographical unit in a continuous period exceeds the dynamic threshold multiple of the daily average sales benchmark, a local sudden order event is determined to have occurred, and the subsequent coverage radius calculation process is activated.
[0012] The location coordinates of community group-buying orders are used to calculate the order coverage radius through geofencing technology, specifically including: For all latitude and longitude coordinates associated with sudden order events, a density-based spatial clustering algorithm is used to identify the core coordinate point of the coordinate point clustering area. The minimum enclosing circle radius covering the coordinate points within the clustering area is calculated by expanding outward from the core coordinate point, and the minimum enclosing circle radius is used as the order coverage radius.
[0013] The order coverage radius is divided into levels, specifically including: The hierarchical interval boundary values of the order coverage radius are preset. The order coverage radius is compared with the hierarchical interval boundary values, and the corresponding discretized radius level identifier is output according to the specific interval range to which the order coverage radius value belongs.
[0014] Simultaneously, by combining the order coverage radius level and the preset noise intensity coefficient, a preset three-dimensional mapping table is queried to obtain the hysteresis fluctuation correction amount, specifically including: A two-dimensional relational matrix is established with the discretized radius level identifier as the row dimension and the preset noise intensity coefficient as the column dimension. Each cell in the two-dimensional relational matrix stores the corresponding hysteresis correction amount. The discretized radius level identifier is input into the two-dimensional relational matrix to determine the noise intensity coefficient. Then, based on the combination of the discretized radius level identifier and the noise intensity coefficient, the hysteresis correction amount value stored in the two-dimensional relational matrix is queried.
[0015] A dynamic attenuation coefficient is generated based on the instantaneous attenuation characteristics of localized burst orders, specifically including: The instantaneous decay characteristic is the nonlinear change law of sudden order sales rising and then falling back to the baseline level within a preset time. A dynamic decay coefficient calculation module based on a negative exponential function is constructed. The dynamic decay coefficient calculation module automatically outputs the decay weight coefficient according to the time difference between the order generation time and the current processing time.
[0016] The sales value of sudden orders is simultaneously applied with a dynamic decay coefficient and a lag fluctuation correction to generate dynamically corrected sales data, specifically including: The dynamic decay coefficient is multiplied by the original sales value of the sudden order to obtain the decay-adjusted sales value. At the same time, the lag fluctuation correction amount is converted into a time dimension offset. The time stamp of the decay-adjusted sales data is shifted forward and backward. Finally, the dynamically corrected sales dataset after sales value decay and time axis correction is output.
[0017] The dynamically corrected sales dataset is aligned with the time stamp sequence and integrated with real-time e-commerce transaction data and store inventory data into a multi-source data stream with a unified spatiotemporal dimension. The data stream is normalized through a time window slicing mechanism, input into a demand forecasting model based on a recurrent neural network architecture, and outputs regional supply chain demand forecasts for the next seven days, triggering automatic replenishment decisions.
[0018] It needs further explanation that the core of the data collection unit's work in collecting community group-buying order data is to obtain complete location and sales information through standardized interfaces, and to complete format unification and data binding. This lays the foundation for subsequent identification of sudden noise and demand forecasting. The specific implementation process is as follows: First, the original location information at the time of order submission is obtained in real time through the open data interface of the community group buying platform. The specific acquisition process is as follows: Configure interface request parameters in the data acquisition unit, including target area order filtering conditions and a list of data return fields, and send GET requests to the platform interface at a frequency of 10 seconds per request. After the interface responds, perform format validation on the returned data, remove abnormal orders with missing fields or incorrect formats, and retain complete and valid order data. The focus is on extracting the original location information and corresponding sales values to prepare for subsequent location processing. When processing the original location information, it is necessary to distinguish its data type and process it accordingly. The original location information is the location data uploaded by the user terminal when the order is submitted, and it usually has two formats: latitude and longitude coordinates and text address. Geographic coordinates are the geographic coordinates of a point on the Earth's surface, directly representing its geographical location. Text addresses are addresses entered by the user or automatically filled in by the terminal; they need to be converted to a unified format for spatial analysis. When the original location information is latitude and longitude coordinates, these coordinates are directly extracted and preserved with 6 decimal places, while the coordinate type is marked as original latitude and longitude. When the original location information is a text address, it is converted to latitude and longitude coordinates through a reverse geocoding service. The reverse geocoding service is an online service that maps text addresses to latitude and longitude coordinates, supporting the parsing of multi-level address information such as province, city, district, street, and house number, with conversion accuracy down to the house number level. The specific conversion process is as follows: the text address is concatenated into request parameters according to the service requirements and a request is sent to the reverse geocoding service interface; after the service responds, the latitude and longitude coordinates are extracted from the returned result, again preserving 6 decimal places. If the text address is incomplete or unrecognizable, the system will automatically mark it as an ambiguous address and use the latitude and longitude of the area's center as the substitute coordinates, while recording the degree of address ambiguity for subsequent data quality assessment. After standardizing the format of the location information, the latitude and longitude coordinates are bound to the sales value of the corresponding order. The sales value is the actual quantity of goods purchased in the order, which is extracted directly from the data returned by the interface to ensure a unique correspondence with the order ID. Binding means establishing the association between latitude and longitude coordinates, sales value, and order ID to form a structured data record. Each record contains three core fields: order ID, latitude and longitude coordinates (latitude, longitude), and sales value. Simultaneously, time stamps are added to the data records. These time stamps are precise timestamps of the order submission, extracted from the order submission time field returned by the interface, ensuring millisecond-level time accuracy. This is used for subsequent analysis of the time distribution characteristics of orders. Finally, all order data bound with latitude and longitude coordinates, sales values, and time stamps are sorted in timestamp order and integrated into a time-stamped order geographic dataset. The order geographic dataset is a structured database table that supports quick querying and filtering by time range, geographic region, sales interval, and other conditions. Each record in the dataset is a complete record of order geographic information, preserving the location and quantity characteristics of the order as well as its time characteristics. This provides complete and standardized basic data support for the subsequent sudden noise processing unit to calculate the order coverage radius and identify local sudden order events based on geofencing technology.
[0019] After the data acquisition unit generates a time-stamped geographical dataset of orders, the primary task of the burst noise processing unit is to construct a baseline value that adapts to recent sales trends. Then, through real-time monitoring and threshold determination, it identifies local burst order events to avoid abnormal data interfering with the accuracy of demand forecasting. The specific implementation process is as follows: First, a daily average sales benchmark is calculated based on historical community group-buying data for the target area. The target area is a pre-defined supply chain service coverage area, which can be divided by administrative regions or business district boundaries. Each area serves as an independent analysis unit to ensure accurate spatial judgment. The historical community group-buying data consists of valid order data from the past 90 days in the target area, after removing orders with fuzzy addresses and extreme anomalies such as single orders exceeding 100 items to ensure data authenticity. The daily average sales benchmark is dynamically calculated based on a fixed time window, which differs from a fixed historical average and can adapt to recent sales fluctuations in real time. The window length has been experimentally verified and set at 7 days to ensure sufficient data volume and rapid response to trend changes. The specific calculation process is as follows: Order data from the past 90 days for the target region is extracted from the order geographic dataset. Daily sales are grouped by date and summed. A calculation window is set up with 7 consecutive days, sliding from the earliest date (sliding step size of 1 day). Within each window, extreme days with sales exceeding the average daily sales of all days in that window by ±3 times the standard deviation are first removed. Then, the sales of the remaining days are calculated and divided by the number of valid days to obtain the daily average sales benchmark value for that window. For example, days 1-7 are the first window. After calculating the benchmark value, days 2-8 are the next window, and so on, sliding to the latest date to form a benchmark value sequence that is updated over time. The benchmark value of the latest window is the core basis for judgment. Subsequently, real-time monitoring of the new order data stream is conducted. This new order data stream consists of newly added order data retrieved by the data collection unit every 10 seconds. It is fully bound to latitude and longitude coordinates, sales volume, and millisecond-level timestamps to ensure real-time monitoring. A single geographic unit is the smallest spatial analysis unit within the target area, divided by latitude and longitude grids. Each grid has a side length of 1 kilometer, ensuring that it is neither too large (leading to dilution of data from sudden events) nor too small (leading to data sparsity). Each geographic unit is assigned a unique grid ID for easy sales statistics. The continuous time period is set to 1 hour, i.e., sales are accumulated based on 6 consecutive 10-minute time periods (matching the data collection frequency). This captures short-term sales surges while avoiding misjudgments due to excessively short time periods. The specific monitoring process is as follows: The system establishes a real-time sales statistics cache for each geographic unit. Upon a new order, its latitude and longitude coordinates are matched with the corresponding grid ID, and the sales value is added to the current hourly cumulative sales for that geographic unit. The cumulative sales are updated every 10 minutes, while historical data from the previous hour is cleared to ensure the cumulative period is always the latest consecutive hour, enabling real-time tracking of sales changes for each geographic unit. The judgment process incorporates a dynamic threshold multiplier. This dynamic threshold multiplier is a variable multiplier based on historical event statistics to avoid misjudgments caused by fixed multipliers. It is negatively correlated with the current rolling daily average sales benchmark: 3 times when the benchmark is ≤50 units, 2.5 times when it's 50-100 units, and 2 times when it's >100 units. This prevents missed judgments in high-sales areas due to excessively high multipliers and misjudgments in low-sales areas due to excessively low multipliers. Furthermore, this multiplier is dynamically adjusted according to the scenario: increased by 20% during holidays and 30% during promotional activities to adapt to normal sales growth under special circumstances. The specific judgment process is as follows: First, obtain the rolling daily average sales benchmark value corresponding to the current geographic unit, determine the dynamic threshold multiple according to the above rules, and multiply the two to obtain the judgment threshold. Extract the cumulative sales value of the current geographic unit in the past hour and compare it with the judgment threshold. If the cumulative value exceeds the threshold, it is determined that a local sudden order event has occurred. This type of event is an abnormal situation in which sales significantly exceed the normal level in a short period of time. It is mostly caused by temporary promotions, sudden events, etc., and belongs to noise in demand forecasting, which needs to be processed later. At this time, the system automatically extracts the latitude and longitude coordinates of all orders associated with the sudden event and activates the subsequent process of calculating the order coverage radius based on geofencing technology to prepare for further correction of sudden noise. If the cumulative value does not exceed the threshold, it is determined to be a normal order flow, and the real-time monitoring status continues to be maintained to ensure that every potential sudden event is identified without omission.
[0020] After determining that a localized sudden order event has occurred, it is necessary to accurately calculate the order coverage radius based on all latitude and longitude coordinates associated with the event. This quantifies the impact range of the sudden event and provides a spatial dimension basis for subsequent matching of noise intensity coefficients and hysteresis correction amounts. The specific implementation process is as follows: First, all latitude and longitude coordinates associated with the sudden order event are extracted. These coordinates are all from a time-stamped order geographic dataset and are the precise geographic location identifiers of all valid orders in the sudden event. Alternative coordinates corresponding to ambiguous addresses have been removed to ensure that each coordinate point can truly reflect the actual distribution of orders. After extraction, duplicates are removed by order ID to avoid redundancy of coordinate points due to data duplication for the same order. Finally, a set of points containing unique latitude and longitude coordinates is formed, which serves as the basis for subsequent cluster analysis. Subsequently, a density-based spatial clustering algorithm was used to identify the core coordinate points of the coordinate point clustering area. The DBSCAN algorithm (a classic density clustering algorithm) was selected here. Its core logic is to divide the coordinate points into clusters by judging whether the density of the coordinate points in a specific neighborhood meets the threshold. It does not require pre-specifying the number of clusters and can adapt to the natural clustering characteristics of sudden order coordinate points, effectively distinguishing dense areas from isolated points. The key parameters of the algorithm need to be calibrated by engineering: the neighborhood radius (Eps) is set to 500 meters, which is determined based on the regular delivery range of community group buying to ensure that the order coordinates in the same delivery area can be classified into the same cluster; the minimum number of points (MinPts) is set to 5, that is, when a coordinate point contains at least 5 other coordinate points in its 500-meter neighborhood, the area is determined to be a dense area, avoiding misjudging a small number of scattered orders as clusters. The specific clustering process is as follows: The set of latitude and longitude coordinates is input into the DBSCAN algorithm. The algorithm first traverses each coordinate point, calculates the number of other coordinate points within its 500-meter neighborhood, and marks the core points that satisfy the minimum number of points ≥ 5. Then, the core points and all coordinate points within their neighborhoods are grouped into the same cluster. This process is repeated until all core points form independent clusters, while removing isolated points with fewer than 5 points in their neighborhoods (these points are mostly random scattered orders and are not included in the cluster analysis). For each cluster, the latitude and longitude mean of all coordinate points within the cluster is calculated (latitude mean = sum of all latitudes / number of points in the cluster; longitude mean = sum of all longitudes / number of points in the cluster). The latitude and longitude coordinates corresponding to this mean are the core coordinate points of the cluster area, which accurately represent the geographical center of the entire cluster. Next, starting from the core coordinate point, we expand outwards to calculate the minimum enclosing circle radius that covers the coordinate points within the cluster. The minimum enclosing circle is the circle with the smallest area that can completely cover all points within a set. Its radius directly reflects the spatial span of the cluster area and is also the core definition of the order coverage radius. The outward expansion starts from the core coordinate point and gradually increases the radius of the circle until all coordinate points within the cluster are included in the circle, and ensures that the radius cannot be reduced further (otherwise, some coordinate points will exceed the radius).The specific calculation process is as follows: First, using the core coordinate point as the center, calculate the straight-line distance from each coordinate point in the cluster to the core point (converted to meter-level distance using the latitude and longitude distance calculation formula), and record the maximum distance as the initial radius; then verify whether the circle corresponding to the initial radius can cover all coordinate points. If there is a coordinate point whose distance to the core point is greater than the initial radius, update the initial radius with the maximum distance and verify again; repeat the verification process until the distance from all coordinate points to the core point is less than or equal to the current radius. The radius obtained at this time is the minimum enclosing circle radius; if the cluster contains only 1 coordinate point, set the minimum enclosing circle radius to 100 meters (based on the standard coverage range of a single self-pickup point in community group buying) to ensure that the radius value has practical significance. Finally, the minimum enclosing circle radius is directly used as the order coverage radius. The order coverage radius is a key parameter for quantifying the spatial impact range of local sudden order events. Its value directly reflects the radiation breadth of the sudden event. For example, a radius of 1.2 kilometers means that the impact of the sudden event is concentrated within a 1.2-kilometer radius around the core coordinate point. This parameter will be associated with the subsequent radius level classification and noise intensity coefficient matching to provide spatial dimension quantitative support for the accurate correction of sudden noise and ensure that the subsequent acquisition of the hysteresis fluctuation correction amount is more in line with the actual scenario.
[0021] After calculating the order coverage radius using density clustering algorithm and minimum enclosing circle, in order to transform the continuous radius values into discretized identifiers that facilitate subsequent matching of noise intensity coefficients and querying of hysteresis fluctuation corrections, it is necessary to pre-set the boundary values of the hierarchical intervals and complete the hierarchical division. The specific implementation process is as follows: First, preset the boundary values of the graded intervals for order coverage radius. The boundary values of the graded intervals are critical values for dividing different coverage levels. They are used to map the continuously changing order coverage radius (unit: meters) into a finite number of discrete levels. The setting of the boundary values must be based on the actual operation data of community group buying in the target area to ensure that the order scenarios corresponding to each level are representative. The specific pre-setting process is as follows: Extract the coverage radius data of all local sudden order events in the target area within the past 6 months. After removing extreme outliers, use the statistical quantile method to determine the boundary values. Finally, mark the boundary of 5 levels of intervals: 0 meters, 500 meters, 1000 meters, 1500 meters, 2000 meters, and 3000 meters, forming 5 continuous and non-overlapping graded intervals: 0 < coverage radius ≤ 500 meters, 500 meters < coverage radius ≤ 1000 meters, 1000 meters < coverage radius ≤ 1500 meters, 1500 meters < coverage radius ≤ 2000 meters, and 2000 meters < coverage radius ≤ 3000 meters. Each interval corresponds to a type of coverage scenario. For example, 0-500 meters corresponds to concentrated sudden orders at a single community self-pickup point, and 2000-3000 meters corresponds to large-scale sudden orders across multiple communities. The boundary values are updated every 3 months based on newly added historical data to ensure adaptation to changes in operational scenarios. The calculated order coverage radius is then compared with the preset tiered interval boundary values. This comparison follows a left-open, right-closed principle: if the coverage radius is exactly equal to a certain boundary value, it is assigned to a higher-level interval. For example, a coverage radius of 500 meters is assigned to the 500-1000 meter interval. This avoids inconsistencies in tier classification due to ambiguous boundary values. The specific comparison process is as follows: extract the specific value (e.g., 850 meters, 1600 meters) from the calculated order coverage radius and compare it sequentially with the tiered interval boundary values in ascending order. First, determine if it is greater than 0 meters and less than or equal to 500 meters. If so, directly lock the corresponding interval. If not, continue to determine if it is greater than 500 meters and less than or equal to 1000 meters, and so on, until a unique interval to which the radius value belongs is found. This ensures that each order coverage radius accurately matches a tiered interval without omissions or duplicate assignments. Finally, based on the specific range of the order coverage radius value, the corresponding discretized radius level identifier is output. The discretized radius level identifier is a simplified level identifier, using the format R + number (such as R1, R2, R3, R4, R5), which facilitates quick indexing and querying in the two-dimensional relation matrix. Each identifier corresponds one-to-one with the level range: R1 corresponds to 0 < coverage radius ≤ 500 meters, R2 corresponds to 500 meters < coverage radius ≤ 1000 meters, R3 corresponds to 1000 meters < coverage radius ≤ 1500 meters, R4 corresponds to 1500 meters < coverage radius ≤ 2000 meters, and R5 corresponds to 2000 meters < coverage radius ≤ 3000 meters.The specific output process is as follows: After determining the range to which the order's coverage radius belongs, the system automatically calls the preset range-identifier mapping relationship and outputs the corresponding discrete radius level identifier. For example, a coverage radius of 850 meters belongs to the range of 500-1000 meters, and the output identifier is R2; a coverage radius of 2200 meters belongs to the range of 2000-3000 meters, and the output identifier is R5. This identifier will serve as the core index parameter for subsequent matching of noise intensity coefficients and querying of hysteresis fluctuation correction amounts, realizing the standardized conversion from continuous radius values to discrete level identifiers, and providing a unified level basis for the accurate correction of sudden noise.
[0022] After obtaining the discretized radius level identifier, in order to accurately match the corresponding hysteresis correction amount, a structured two-dimensional relational matrix needs to be established first. This matrix uses the discretized radius level and the preset noise intensity coefficient as dual indexes to achieve fast query and matching of the correction amount. The specific implementation process is as follows: First, a two-dimensional relational matrix is established with the discretized radius level identifier as the row dimension and the preset noise intensity coefficient as the column dimension. The two-dimensional relational matrix is a structured data table used to store the correspondence between the dual indexes and the target parameters. The core function here is to bind the combination of radius level and noise intensity with the hysteresis correction amount one by one, which facilitates fast retrieval later. The row dimension is the row direction of the matrix, which directly uses the discretized radius level identifier defined above ( R1, R2, R3, R4, R5), each identifier corresponds to a row, ensuring that the row index is completely consistent with the radius level; the preset noise intensity coefficient is a parameter that quantifies the impact of sudden order noise on demand forecasting. The value range is calibrated to 0.1-0.9 based on historical data, and divided into 5 levels (0.1, 0.3, 0.5, 0.7, 0.9) in intervals of 0.2. The larger the value, the more significant the impact of noise on sales fluctuations. The level setting of this coefficient is based on the actual noise impact statistics of sudden orders at different radius levels. For example, the noise impact of small-scale sudden orders (R1) is relatively concentrated, and the coefficient level interval does not need to be too close, while the noise impact of large-scale sudden orders (R5) varies greatly. The 5 levels can accurately cover different scenarios. The specific process is as follows: First, construct a basic framework of 5 rows and 5 columns in the matrix, with row headings R1 to R5 and column headings 0.1, 0.3, 0.5, 0.7, and 0.9 respectively. Then, fill each cell with the corresponding lag fluctuation correction amount. The lag fluctuation correction amount is a parameter used to correct the fluctuation of the sales time axis of sudden orders, and the unit is minutes. Its core function is to shift the time marker corresponding to the sudden sales to a reasonable time period to offset the time dimension deviation caused by noise. Its value is obtained by statistically analyzing historical sudden order data: For each radius level-noise intensity combination, extract all sudden order events under this combination in the past 6 months, calculate the average lag time of its sales fluctuation (i.e., the time difference between the sales peak and the normal period), and use this average time as the lag fluctuation correction amount of the corresponding cell. For example, the average lag time corresponding to the combination of R2 (500-1000 meters) and noise intensity coefficient 0.3 is 15 minutes, so the cell is filled with 15. Finally, a complete two-dimensional relationship matrix is formed. The matrix is updated once a quarter based on the newly added historical data to ensure the timeliness of the correction amount. Next, the discretization radius level identifier is input into the two-dimensional relation matrix to determine the noise intensity coefficient. The core logic here is that there is a strong correlation between the noise intensity coefficient and the discretization radius level. The higher the radius level (the wider the coverage), the more dispersed the noise impact of sudden orders, but the higher the overall intensity. Therefore, it is necessary to match the corresponding default noise intensity coefficient according to the radius level to avoid the subjectivity of coefficient selection.The specific determination process is as follows: The system has a built-in mapping rule for radius level and default coefficient. This rule is set based on historical noise intensity statistics: R1 corresponds to a default coefficient of 0.3 (small-scale sudden noise has a concentrated impact and medium intensity), R2 corresponds to 0.5 (medium-scale sudden noise has a diffuse impact and medium to high intensity), R3 corresponds to 0.5 (medium to large-scale sudden noise has a stable impact and medium to high intensity), R4 corresponds to 0.7 (large-scale sudden noise has a wide impact and high intensity), and R5 corresponds to 0.9 (ultra-large-scale sudden noise has the greatest impact and the highest intensity). After inputting the currently obtained discretized radius level identifier (such as R3) into the matrix, the system automatically calls this mapping rule to determine the corresponding default noise intensity coefficient (such as 0.5). If the coefficient needs to be manually adjusted later (such as in special scenarios where the noise intensity is abnormal), it can also be modified through the system interface. However, by default, the automatically matched coefficient is used to ensure the convenience and accuracy of the operation. Finally, based on the combination of discretized radius level identifier and noise intensity coefficient, the hysteresis correction value stored in the two-dimensional relation matrix is queried. The query process is essentially a dual-index positioning, which finds the unique corresponding cell through the row index (radius level identifier) and column index (noise intensity coefficient) and extracts the correction value from it. The specific query process is as follows: First, determine the target row in the matrix based on the discretization radius level identifier (e.g., R3 corresponds to the 3rd row), then determine the target column based on the determined noise intensity coefficient (e.g., 0.5 corresponds to the 3rd column). The cell where the target row and target column intersect is the matching storage location. The system automatically extracts the value in that cell (e.g., 20 minutes) as the lag fluctuation correction amount for the current sudden order event. If, due to special circumstances (e.g., adding a radius level or coefficient level), there is no corresponding cell in the matrix, the system will automatically use the average value of adjacent cells as the transition correction amount (e.g., if there is no corresponding cell for R3 and coefficient 0.6, then the average of 20 minutes corresponding to 0.5 and 25 minutes corresponding to 0.7, 22.5 minutes, will be taken), ensuring no query omissions. This correction amount will subsequently be used to adjust the time stamp of sudden order sales, realize time axis correction, and provide key parameter support for generating dynamically corrected sales data.
[0023] After obtaining the lagged fluctuation correction amount, in order to accurately mitigate the abnormal interference of sudden order sales on demand forecasting, it is also necessary to generate a dynamic decay coefficient based on its instantaneous decay characteristics. By quantifying the time decay law of sales, the corrected sales are made to better reflect normal demand trends. The specific implementation process is as follows: First, let's clarify the core definition of instantaneous decay characteristics. This characteristic refers to the non-linear change in sales patterns of sudden, localized orders, which rapidly rise to a peak within a preset timeframe and then gradually decline back to the daily average sales level. The preset timeframe is defined as 24 hours based on historical sudden order data. This is because sudden orders in community group buying scenarios (such as temporary promotions, emergency procurement in the community, short-term supply-demand imbalances, etc.) typically have a time-sensitive impact and do not deviate from normal demand for long periods. Furthermore, the decay process exhibits a non-linear characteristic of rapid initial decline followed by slower decline, rather than a uniform rate of decline. For example, in a community, emergency supply group buying orders triggered by a sudden power outage will peak within 2-4 hours. Subsequently, as power is restored and demand is met, sales rapidly decline, approaching normal levels after 12 hours, and completely declining after 24 hours. This natural decay pattern is the basis for calculating the dynamic decay coefficient. Based on this characteristic, a dynamic attenuation coefficient calculation module based on a negative exponential function is constructed. The negative exponential function accurately simulates the trend of rapid attenuation followed by stabilization, which highly matches the attenuation characteristics of sudden order sales. Its core logic is to make the attenuation coefficient decrease non-linearly with the increase of the time difference, ensuring both reasonable sales weighting in the initial stage of a sudden event and quickly weakening the abnormal impact in the later stages. The dynamic attenuation coefficient calculation module is an algorithm module integrated into the sudden noise processing unit, possessing functions such as timestamp reading, time difference calculation, and automatic coefficient output, requiring no manual intervention and responding to the order processing flow in real time. The core parameters of the module need to be calibrated by engineering: the attenuation constant is set to 0.1 (unit: 1 / hour), a value determined by fitting the actual attenuation curves of all sudden orders in the past 6 months, ensuring that the coefficient change is consistent with the actual attenuation pattern; at the same time, the coefficient range is set to 0.1-0.95 to avoid the coefficient being too low, leading to excessive weakening of effective demand, or too high, failing to offset sudden noise. The core workflow of the dynamic attenuation coefficient calculation module is to automatically output the attenuation weight coefficient based on the time difference between the order generation time and the current processing time. The order generation time is the millisecond-level timestamp when the order is submitted, which is directly extracted from the time-stamped order geographic dataset to ensure time accuracy. The current processing time is the system's millisecond-level timestamp when the module starts the calculation, which is synchronized in real time by the system clock. The time difference is the difference between the two. To facilitate calculation and fit the attenuation cycle, it is converted to hourly units (less than 1 hour is counted as 1 hour). For example, if the order is generated at 10:30 and the current processing time is 14:45, the time difference is 5 hours.The specific calculation process is as follows: The module first reads the timestamp of the target sudden order's generation time, synchronously obtains the current system processing timestamp, calculates the time difference between the two and converts it into hours; the time difference is input into the preset negative exponential function logic, and the function will output the corresponding attenuation coefficient according to the size of the time difference. When the time difference is 0 hours (the order has just been generated, and the sudden impact is the strongest), the coefficient is 0.95, retaining most of the sales weight but not exaggerating the peak; when the time difference is 6 hours, the coefficient is about 0.55, weakening nearly half of the abnormal impact; when the time difference is 12 hours, the coefficient is about 0.3, further reducing the weight; when the time difference reaches 24 hours or more, the coefficient is fixed at 0.1, at which point the sudden impact has basically disappeared, retaining only a small amount of weight to avoid missing potential effective demand; if the time difference exceeds 48 hours, the coefficient remains at 0.1 and no longer attenuates. The final output attenuation weighting coefficient is a value between 0.1 and 0.95. This coefficient will be used to multiply the original sales value of the sudden order to achieve quantitative attenuation of the abnormal peak. This ensures that the dynamically corrected sales data retains the reasonable demand component in the sudden order while eliminating noise interference that deviates excessively from the normal trend, laying the foundation for subsequent integration with regular sales data.
[0024] After obtaining the dynamic attenuation coefficient and the hysteresis fluctuation correction amount, the sudden noise processing unit needs to perform a double correction on the sales value of sudden orders. This both weakens the numerical interference of abnormal peaks and corrects the fluctuation deviation in the time dimension, ultimately generating dynamically corrected sales data that conforms to normal demand patterns. The specific implementation method is as follows: First, a sales volume decay adjustment is performed. The dynamic decay coefficient is multiplied by the original sales volume value of the sudden order to obtain the decayed sales volume value. The original sales volume value of the sudden order is the actual purchase quantity (unit: pieces) recorded when the order is submitted. It is directly extracted from the time-stamped order geographic dataset to ensure a unique correspondence with the order ID without any correction processing. The dynamic decay coefficient is a weighting coefficient between 0.1 and 0.95 calculated based on the negative exponential function. The value is positively correlated with the time difference between the order generation time and the current processing time. The larger the time difference, the smaller the coefficient and the stronger the decay. The specific multiplication process is as follows: Extract the original sales value (e.g., 100 units) and the corresponding dynamic decay coefficient (e.g., 0.55) of a single sudden order. Perform a multiplication operation on the two, and keep the result as an integer (rounding is used because the sales volume is a discrete number of units). For example, 100 × 0.55 = 55 units, which gives the sales value after decay adjustment. If the result is less than 1, it is retained as 1 unit to avoid the effective demand being completely eliminated due to a sales value of 0. Through this step, the abnormal peak of sudden orders is weakened while retaining the reasonable demand components. At the same time, the lag fluctuation correction is converted into a time dimension offset. The lag fluctuation correction is a minute-level parameter (e.g., 15 minutes) previously queried from the two-dimensional relationship matrix, which reflects the time lag or advance of the sudden order sales fluctuation relative to the normal demand period. The time dimension offset is the conversion of the minute-level correction into a millisecond-level time adjustment value, which is used to directly correct the order's timestamp. Because the timestamp is a millisecond-level timestamp, a unified unit is required to perform the translation operation. The specific conversion process is as follows: Multiply the lag fluctuation correction amount (in minutes) by 60,000 (1 minute = 60,000 milliseconds) to obtain a millisecond-level offset, for example, 15 minutes × 60,000 = 900,000 milliseconds; the offset direction is determined according to the physical meaning of the correction amount. If the correction amount represents lag (i.e., the sudden sales peak is later than the normal demand period), the offset direction is backward (timestamp added to offset); if it represents advance, it is forward (timestamp subtracted from offset). Here, based on the general pattern of sudden orders in community group buying, the default offset direction is backward. If adjustments are needed in special scenarios, they can be switched through the system's preset rules to ensure that the time correction meets the actual needs. Subsequently, the timestamps of the sales data after attenuation adjustment are shifted forward and backward. The timestamps are the original millisecond-level timestamps of the orders (format: year-month-day hour:minute:second.millisecond), recording the precise time of order submission; the forward and backward shifting operation is based on the original timestamps, performing addition and subtraction operations according to the time dimension offset to adjust the time allocation of the sales data, so that the corrected sales can correspond to the actual demand period it affects. The specific operation process is as follows: Extract the original timestamp corresponding to the sales data after attenuation adjustment (e.g., 2024-05-20 10:30:00.000), convert it to a millisecond-level value (e.g., 1716186600000), and then add the previously calculated millisecond-level offset (e.g., 900000 milliseconds) to obtain the adjusted millisecond-level timestamp (1716187500000). Then convert it back to year-month-day hour:minute:second millisecond format (e.g., 2024-05-20 10:45:00.000). If the shifted timestamp exceeds 24:00 of the current day, it will be automatically shifted to the corresponding time period of the next day (e.g., 23:50 shifted 15 minutes becomes 00:05 of the next day), ensuring logical consistency of time and no invalid time markers. Finally, all the corrected data is integrated to output a dynamically corrected sales dataset. This dataset is a structured database table that retains the core fields of the order geographic dataset and adds four new fields: corrected sales, adjusted timestamp decay coefficient, and time offset, to ensure data traceability. The specific integration process is as follows: For each sudden order, the order ID, latitude and longitude coordinates, discretization radius level identifier, sales value after attenuation adjustment, translated timestamp, corresponding dynamic attenuation coefficient, and time dimension offset are associated one by one to form a complete corrected data record. After performing the above correction process on all sudden orders, all records are sorted according to the order of the adjusted timestamps. At the same time, data validation and deduplication are performed to remove duplicate order records (by order ID) and filter invalid records with adjusted sales of 0, ensuring that the data in the dataset is unique and valid. The final output dynamic corrected sales dataset not only solves the problem of abnormal sales values in sudden orders, but also corrects the fluctuation deviation in the time dimension, achieving dual optimization of numerical attenuation and time correction. It can be seamlessly integrated with regular sales data, laying a high-quality data foundation for subsequent multi-source heterogeneous data integration and demand forecasting model input.
[0025] After the burst noise processing unit outputs a dynamically corrected sales dataset that has undergone sales attenuation and time axis correction, the prediction execution unit needs to further integrate multi-source heterogeneous data, standardize it, and input it into the prediction model to finally generate accurate regional demand prediction results, providing a basis for supply chain response decisions. The specific implementation method is as follows: First, the dynamically corrected sales dataset is aligned according to a time stamp sequence. Time stamp sequence alignment means sorting all records in the dataset in chronological order based on millisecond-level timestamps, while filling in missing values for periods with no orders (sales during periods with no orders are recorded as 0) to ensure continuity in the time dimension. Specifically, the adjusted timestamp field is extracted from the dynamically corrected sales dataset and converted to a uniform millisecond-level value. All records are then sorted in ascending order. The minimum interval for the time series is set to 10 minutes (to match the data collection frequency). The sorted dataset is iterated, and if the time interval between two adjacent records exceeds 10 minutes, a supplementary record with a sales value of 0 is inserted during the interval. This ensures an uninterrupted sequence from the earliest to the latest timestamps, avoiding incomplete model input data due to time gaps. The aligned dataset was then integrated with e-commerce real-time transaction data and store inventory data into a unified spatiotemporal dimension multi-source data stream. The e-commerce real-time transaction data is sales data of target area products obtained from the open interface of e-commerce platforms, including fields such as order ID, timestamp, geographic unit ID, sales value, and payment status, which complements the community group buying data. The store inventory data is dynamic inventory data uploaded by each physical store every hour, including fields such as store ID, product ID, inventory quantity, inventory warning threshold, and geographic unit ID, reflecting the current supply and demand balance. The unified spatiotemporal dimension means that all data are associated with a three-dimensional primary key of timestamp-geographic unit ID-product ID, ensuring that data from different sources of the same product at the same time and in the same region can be matched. The specific integration process is as follows: First, the fields of the three types of data are standardized, and the coding rules of geographic unit IDs (all adopt 1-kilometer grid IDs), timestamp format (millisecond level), and classification standards of product IDs are unified. Then, using timestamp-geographic unit ID-product ID as a joint index, the sales field in the dynamically corrected sales data and e-commerce real-time transaction data is summed (the sales under the same index are accumulated to reflect the total demand of multiple channels), and associated with the inventory quantity and warning threshold fields in the store inventory data. This forms a structured multi-source data stream in which each record contains timestamp, geographic unit ID, product ID, total sales, current inventory, and inventory warning threshold, realizing the spatiotemporal linkage between demand and inventory data. After integration, the data stream is normalized using a time window slicing mechanism. This mechanism divides a continuous time series into multiple non-overlapping time windows of fixed duration. Each window serves as an input sample for the model, capturing demand characteristics within the time period. The window duration was experimentally calibrated to 1 hour, ensuring that each window contains sufficient data while accurately reflecting short-term demand fluctuations. The normalization process maps numerical features (total sales, current inventory) within the window to the 0-1 range, eliminating the impact of differences in feature units (e.g., sales in units of pieces, inventory in units of boxes) on model training and preventing any single feature from dominating gradient updates.The specific process is as follows: The multi-source data stream is sliced according to an hourly duration. Each window contains the total sales and current inventory data of all geographical units and all products within that hour. If there is no data for a product in a certain geographical unit within a window, the total sales are recorded as 0, and the current inventory uses the data from the previous window. For the total sales field in each window, Min-Max normalization is used, and the current window sales are mapped to 0-1 based on the maximum and minimum sales of the product in the target area over the past 90 days. For the current inventory field, the maximum inventory capacity of the product (store inventory limit) is used as the benchmark, and it is also mapped to 0-1. After normalization, the data of each time window forms a feature matrix of geographical units × product number × 2, which serves as the input sample for the model. Next, the processed feature matrix is input into a demand forecasting model based on a recurrent neural network architecture. Here, a Long Short-Term Memory (LSTM) network is chosen for the recurrent neural network architecture. Its core advantage is that it can capture long-term dependencies in time series data and adapt to the trend and cyclical characteristics of supply chain demand. The model consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is consistent with the dimension of the feature matrix (the number of features for each sample is 2, i.e., the normalized total sales and current inventory). There are 3 hidden layers, each containing 64 neurons, and the ReLU activation function is used to enhance the model's nonlinear fitting ability. The number of neurons in the output layer matches the prediction dimension (the predicted sales value for each region for the next seven days). The specific implementation process is as follows: First, the model is pre-trained. The training data uses multi-source heterogeneous data from the target region over the past 12 months (after being sliced and normalized). During training, the mean squared error between the predicted value and the actual sales is used as the loss function. The model parameters are iteratively updated using the Adam optimizer, with 100 iterations and an initial learning rate of 0.001. The learning rate decreases by 50% every 20 iterations until the loss function converges (the loss change is less than 0.0001 after 10 consecutive iterations). After pre-training, the real-time processed time window feature matrix is input into the model in sequence. The model extracts time-dependent features (such as demand patterns at the same time of day and trend changes over several consecutive days) through the LSTM layer, and integrates spatial features (demand differences between different geographical units) through the fully connected layer, finally outputting the prediction result. The model outputs regional supply chain demand forecasts for the next seven days. The regional division is based on a previously set 1-kilometer grid geographical unit, with each geographical unit corresponding to a set of daily demand forecasts. The demand forecasts are normalized inverse mapping results, that is, by using the Min-Max baseline values during training, the 0-1 range values output by the model are restored to actual sales (unit: pieces). At the same time, the prediction confidence level (such as 95% confidence interval) is given to reflect the reliability of the prediction results.The specific output process is as follows: The model outputs normalized forecast values for each geographic unit and each product for the next seven days. The system calls the pre-stored baseline values (maximum sales volume, minimum sales volume) and reverse-calculates the normalized values to obtain the actual sales forecast values. The forecasts are summarized by geographic unit and product category, generating a structured forecast report with date-geographic unit-product-forecasted sales volume-confidence level. The daily regional forecast values need to be combined with inventory data to calculate the demand gap (demand gap = forecasted sales volume - current inventory). If the demand gap exceeds the inventory warning threshold, it is marked as needing replenishment. Finally, automatic replenishment decisions are triggered based on the forecast results. Automatic replenishment decisions are supply chain execution instructions generated by the execution unit based on the forecasted demand gap and inventory status, requiring no manual intervention and directly connecting to warehousing and logistics systems. The specific process is as follows: The response execution unit traverses the replenishment mark records in the forecast report, aggregates the daily total demand gap by geographical unit, and determines the replenishment trigger time by combining the replenishment lead time of the stores (24 hours for the same city, 48 hours for cross-regional areas) (if the demand gap of a certain geographical unit is predicted to be large in the next 3 days, replenishment will be triggered 2 days in advance); based on the warehouse distribution of the goods (the principle of replenishment by proximity), the optimal replenishment quantity is calculated (replenishment quantity = demand gap + safety stock, where the safety stock is 1.5 times the average daily sales of the product in the past 7 days); an automatic replenishment instruction is generated, which includes the replenishment time, target store (the physical store corresponding to the geographical unit), product ID, replenishment quantity, and delivery route. This instruction is sent to the warehouse management system and logistics scheduling system through the interface to drive warehouse preparation and logistics distribution, so as to realize the rapid response of the supply chain to the forecast demand and ensure that the inventory of goods can meet future demand without causing inventory backlog due to over-replenishment.
[0026] The second objective of this invention is to provide a system for implementing a supply chain demand forecasting method based on multi-source heterogeneous data fusion, comprising: Data acquisition unit 1 obtains order location coordinates and sales value in real time through the community group buying platform interface, performs reverse geocoding on the text address, and generates a time-stamped order geographic dataset; The sudden noise processing unit 2 identifies local sudden order events based on the rolling daily average sales benchmark and dynamic threshold multiple. It uses density clustering algorithm to calculate the order coverage radius and divides the discretized radius level identifier. It obtains the hysteresis fluctuation correction amount by mapping the radius level and noise intensity coefficient through a two-dimensional relationship matrix. It generates a dynamic attenuation coefficient by combining a negative exponential function. It applies the attenuation coefficient and time dimension offset to the sudden sales value simultaneously and outputs the dynamically corrected sales dataset. Prediction execution unit 3 aligns the dynamically corrected dataset with e-commerce transaction data and store inventory data according to time stamps. After time window slicing and normalization, it is input into the recurrent neural network prediction model to generate regional demand forecasts for the next seven days. The response execution unit 4 triggers automatic replenishment decisions based on demand forecasts and drives the supply chain to execute responses.
[0027] This invention.
[0028] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A supply chain demand forecasting method based on multi-source heterogeneous data fusion, characterized in that: Includes the following steps: S1. Real-time collection of location coordinates and sales volume of community group-buying orders; S2. By using a preset sales threshold, sudden noise is identified in community group buying orders. When a local sudden order event is identified, the order coverage radius is calculated and the order coverage radius level is divided based on the location coordinates of the community group buying order using geofencing technology. A preset noise intensity coefficient is matched according to the order coverage radius level. At the same time, the order coverage radius level and the preset noise intensity coefficient are combined to query a preset three-dimensional mapping table to obtain the hysteresis fluctuation correction amount. A dynamic attenuation coefficient is generated based on the instantaneous attenuation characteristics of the local sudden order. The sales value of the sudden order is synchronously applied with the dynamic attenuation coefficient and the hysteresis fluctuation correction amount to generate dynamically corrected sales data. S3. Integrate dynamically corrected sales data with regular sales data, input the forecasting model to output demand forecasting results, and execute the demand forecasting results to drive supply chain response.
2. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The real-time collection of location coordinates for community group-buying orders specifically includes: The platform uses the open data interface of the community group buying platform to obtain the original location information when the order is submitted in real time. When the original location information is latitude and longitude coordinates, it is recorded directly. When the original location information is a text address, it is converted into latitude and longitude coordinates through reverse geocoding service. The latitude and longitude coordinates are then bound to the sales value of the corresponding order to form a time-stamped order geographic dataset.
3. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 2, characterized in that: The step of identifying sudden noise in community group-buying orders by setting a preset sales threshold specifically includes: The daily average sales benchmark is calculated based on historical community group buying data of the target area and is updated on a rolling basis. The new order data stream is monitored in real time. When the cumulative sales of a single geographical unit in a continuous period exceeds the dynamic threshold multiple of the daily average sales benchmark, a local sudden order event is determined to have occurred, and the subsequent coverage radius calculation process is activated.
4. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 2, characterized in that: The location coordinates based on community group-buying orders are used to calculate the order coverage radius through geofencing technology, specifically including: For all latitude and longitude coordinates associated with sudden order events, a density-based spatial clustering algorithm is used to identify the core coordinate point of the coordinate point clustering area. The minimum enclosing circle radius covering the coordinate points within the clustering area is calculated by expanding outward from the core coordinate point, and the minimum enclosing circle radius is used as the order coverage radius.
5. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The classification of order coverage radius levels specifically includes: The order coverage radius is pre-defined as a graded interval boundary value. The order coverage radius is compared with the graded interval boundary value, and the corresponding discretized radius level identifier is output according to the specific interval range to which the order coverage radius value belongs.
6. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The process of simultaneously combining the order coverage radius level and the preset noise intensity coefficient to query a preset three-dimensional mapping table to obtain the hysteresis fluctuation correction amount specifically includes: A two-dimensional relational matrix is established with the discretized radius level identifier as the row dimension and the preset noise intensity coefficient as the column dimension. Each cell in the two-dimensional relational matrix stores the corresponding hysteresis correction amount. The discretized radius level identifier is input into the two-dimensional relational matrix to determine the noise intensity coefficient. Then, based on the combination of the discretized radius level identifier and the noise intensity coefficient, the hysteresis correction amount value stored in the two-dimensional relational matrix is queried.
7. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The generation of a dynamic attenuation coefficient based on the instantaneous attenuation characteristics of local burst orders specifically includes: The instantaneous decay characteristic refers to the nonlinear change law of sudden order sales rising and then falling back to the baseline level within a preset time. A dynamic decay coefficient calculation module based on a negative exponential function is constructed. The dynamic decay coefficient calculation module automatically outputs the decay weight coefficient according to the time difference between the order generation time and the current processing time.
8. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 7, characterized in that: The process of simultaneously applying a dynamic attenuation coefficient and a lag fluctuation correction to the sales value of sudden orders to generate dynamically corrected sales data specifically includes: The dynamic decay coefficient is multiplied by the original sales value of the sudden order to obtain the decay-adjusted sales value. At the same time, the lag fluctuation correction amount is converted into a time dimension offset. The time stamp of the decay-adjusted sales data is shifted forward and backward. Finally, the dynamically corrected sales dataset after sales value decay and time axis correction is output.
9. The supply chain demand forecasting method based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The dynamically corrected sales dataset is aligned with the time stamp sequence and integrated with real-time e-commerce transaction data and store inventory data into a multi-source data stream with a unified spatiotemporal dimension. The data stream is normalized through a time window slicing mechanism, input into a demand forecasting model based on a recurrent neural network architecture, and outputs regional supply chain demand forecasts for the next seven days, triggering automatic replenishment decisions.
10. A system for implementing a supply chain demand forecasting method based on multi-source heterogeneous data fusion as described in any one of claims 1-9, characterized in that, include: The data acquisition unit (1) obtains the order location coordinates and sales value in real time through the community group buying platform interface, performs reverse geocoding on the text address, and generates a time-stamped order geographic dataset; The sudden noise processing unit (2) identifies local sudden order events based on the rolling daily average sales benchmark value and dynamic threshold multiple. It uses density clustering algorithm to calculate the order coverage radius and divides the discretized radius level identifier. It obtains the hysteresis fluctuation correction amount by mapping the radius level and noise intensity coefficient through a two-dimensional relation matrix. It generates a dynamic attenuation coefficient by combining a negative exponential function. It applies the attenuation coefficient and time dimension offset to the sudden sales value simultaneously and outputs the dynamic corrected sales dataset. The prediction execution unit (3) aligns the dynamically corrected dataset with e-commerce transaction data and store inventory data according to time stamps, and after time window slicing normalization, inputs it into the recurrent neural network prediction model to generate regional demand prediction values for the next seven days. The response execution unit (4) triggers automatic replenishment decisions based on demand forecasts and drives the supply chain to execute responses.