Fast fashion women's garment cost measuring and calculating system and method based on multi-source heterogeneous data

By using multi-source heterogeneous data processing and random forest models, the problems of low efficiency, large error, and lack of quantifiable reliability in traditional apparel cost estimation have been solved. This has enabled accurate cost calculation and risk assessment for fast fashion women's ready-to-wear garments, improving the scientific nature and efficiency of enterprise cost management.

CN120996849APending Publication Date: 2025-11-21WANSHANG (NINGBO) FASHION CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511100050.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional garment cost estimation methods are inefficient, prone to errors, lack sufficient data utilization, have unquantifiable reliability, and have unclear influence from characteristics. Existing methods are insufficient for accurate calculation and risk assessment.

Method used

A cost calculation system for fast fashion women's clothing based on multi-source heterogeneous data is adopted, including modules for raw data acquisition, data processing, calculation model and confidence assessment. The mean, standard deviation and confidence interval of the calculated values ​​are calculated by random forest model and bootstrap method. TF-IDF and one-hot encoding are combined to process text and numerical features to achieve accurate calculation and confidence assessment.

Benefits of technology

It improves the efficiency and accuracy of calculations, quantifies key cost-influencing factors, provides reliable calculation results and risk assessments, and supports enterprises in optimizing resource allocation and pricing strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996849A_ABST
    Figure CN120996849A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of multi-source heterogeneous data, and discloses a system and method for measuring and calculating the cost of fast fashion women's ready-made clothes based on multi-source heterogeneous data, and the system comprises an original data acquisition module, a data processing module, a measurement and calculation model module, a confidence evaluation module, and a storage and display module. The method comprises the following steps: acquiring price checking data through a multi-source data interface and formatting the price checking data; 11 core features are extracted after data cleaning, and text type features, classification type features and numeric type features are processed respectively; calculating the cost by adopting a random forest model, and outputting a mean value, a standard deviation and a 99% confidence interval by combining Bootstrap re-sampling; weighting and calculating a confidence level based on the garment type, the fabric and the labor price similarity; and finally, storing a result and visually displaying the result. According to the invention, the problems of low efficiency, large error and insufficient data utilization of traditional manual measurement and calculation are solved, and high-precision and high-reliability automatic cost measurement and calculation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multi-source heterogeneous data technology, and in particular relates to a cost calculation system and method for fast fashion women's ready-to-wear garments based on multi-source heterogeneous data. Background Technology

[0002] In the process of garment production and operation, cost estimation is a core element in formulating pricing strategies, controlling production expenses, and optimizing resource allocation. Traditional garment cost estimation mainly relies on manual experience, estimating the total cost by adding up the costs of individual items such as fabric, accessories, and labor hours. This method has the following significant drawbacks:

[0003] Inefficient and prone to errors: Manual calculations rely on individual experience, which can easily lead to errors in cost estimation for complex styles or combinations of multiple materials, and it is difficult to meet the needs of rapid calculation for large batches of styles.

[0004] Insufficient data utilization: The pricing table contains more than 70 original features such as fabric composition, unit consumption, labor cost, and transportation cost, but it lacks systematic organization. Key features (such as fabric market fluctuations and process complexity) have not been effectively mined, resulting in a waste of data value.

[0005] The reliability of the calculation is not quantified: Existing methods only output a single cost estimate, which cannot reflect the credibility of the calculation results and is difficult to support enterprises in risk assessment in pricing and order decisions;

[0006] The impact of features is unclear: it is impossible to quantify the weight of the impact of factors such as fabric unit price, transportation cost, and labor cost on the total cost, which is not conducive to targeted optimization of cost structure (such as reducing the expenditure of high-impact factors).

[0007] In existing technologies, some companies have attempted to use simple statistical models (such as linear regression) for cost calculation, but these methods still fail to address issues such as insufficient data standardization and a lack of confidence assessment, resulting in limited accuracy and practicality. Therefore, there is an urgent need for a system that integrates multi-source features to achieve accurate calculations and quantify reliability, thereby improving the scientific rigor and efficiency of apparel cost management. Summary of the Invention

[0008] The purpose of this invention is to provide a cost calculation system and method for fast fashion women's clothing based on multi-source heterogeneous data, so as to solve the above-mentioned technical problems.

[0009] To address the aforementioned technical problems, the specific technical solution of this invention—a cost calculation system and method for fast fashion women's ready-to-wear garments based on multi-source heterogeneous data—is as follows:

[0010] A cost calculation system for fast fashion women's clothing based on multi-source heterogeneous data includes a raw data acquisition module, a data processing module, a calculation model module, a confidence assessment module, and a storage and display module. The raw data acquisition module integrates methods for acquiring data from multiple platforms to obtain clothing pricing data and outputs it in a formatted format to the data processing module. The data processing module cleans the received pricing data, extracting 11 core features; it processes textual, categorical, and numerical features separately, converting them into modelable numerical forms. The calculation model module uses the processed feature data as input and total cost as the target variable to construct a random forest calculation model. It calculates the mean, standard deviation, 99% confidence interval, and extended interval of the calculated values ​​through Bootstrap resampling and outputs the calculation results. The confidence assessment module calculates the similarity of each indicator based on the feature data output by the data processing module, taking the average value as the confidence level of the calculation results. The storage and display module stores the calculation results in a database and displays them using visualization technology.

[0011] This invention also discloses a method for calculating the cost of fast fashion women's ready-to-wear garments based on a multi-source heterogeneous data system, comprising the following steps:

[0012] S1, the raw data acquisition module integrates multiple text files, Excel files, and CSV files obtained from multiple business systems, extracts clothing pricing data from them, verifies the data, and outputs the formatted data to the data processing module;

[0013] In the S2 data processing module, the received pricing data is first cleaned and 11 core features are extracted. Then, the textual, categorical, and numerical features are processed separately and converted into a modelable numerical form. The processed feature data is then output to the calculation model module.

[0014] In the S3 calculation model module, the processed feature data is used as input and the total cost is used as the target variable to construct a random forest calculation model, calculate the calculated value of the garment price, and output the calculation results to the confidence assessment module and the storage and display module.

[0015] In the S4 confidence assessment module, based on the feature data output by the data processing module, the similarity of garment type, fabric, and labor cost is calculated respectively. The three similarities are fused, and the result after feature importance normalization is used as the weight to calculate the confidence of the measurement result. The Percentile Bootstrap method is used to calculate the 99% confidence interval and output it to the storage and display module.

[0016] In the S5 storage and display module, the calculation results are displayed on the front end using visualization technology and then saved to the database.

[0017] Furthermore, step S1 includes the following steps:

[0018] S11. Integrate multi-source data acquisition methods to obtain various pricing data from multiple business systems, and encapsulate these functions to ensure compatibility with pricing data in different formats;

[0019] S12. Validate the input pricing data. If any abnormal fields are found, a clear error message is displayed. If all fields are present, the data is formatted into a uniform data frame, indexed by the item number, and output to the data processing module. Further, S2 includes the following steps:

[0020] S21. Perform data cleaning. For the total cost column, delete invalid values ​​and reset the index, and convert the total cost column to floating point type. For numeric features, first convert them to string type, remove non-digit and decimal characters using regular expressions, then replace invalid values ​​with NaN, convert them to numeric type, and fill missing values ​​with the median. Categorical features and text features retain their original format for subsequent processing.

[0021] S22. Perform feature processing: use one-hot encoding for categorical features to generate a data frame; for text features, perform word segmentation and convert them into TF-IDF vectors to form a text feature matrix; merge the processed categorical features, text features and numerical features by row to construct a complete feature matrix and output it to the measurement model module.

[0022] Furthermore, step S3 includes the following steps:

[0023] S31. Model Training and Dataset Partitioning: A random forest model is used, with features such as classification, fabric, fabric unit price, fabric unit consumption, auxiliary material cost, labor cost, tax rate, transportation, and total cost. The test set and training set are allocated using the "80 / 20 rule".

[0024] S32. Calculation of measurement results: For the test set samples, use the random forest model to output the measurement results of the garment price, so that each style number can get a corresponding measurement value.

[0025] S33. Feature Importance Evaluation: In random forests, feature importance is calculated based on the reduction in Gini impurity. For each feature, its Gini impurity is... The reduction in Gini impurity is ΔG = G 父节点 -(ω·G 子节点1 +ω·G 子节点2 The importance of this feature is the average of the sum of the reductions in Gini impurity caused by this feature across all decision trees.

[0026] Furthermore, step S4 includes the following steps:

[0027] S41. For garment type similarity calculation, the similarity within the same category is recorded as 1, and no similarity is calculated for different categories;

[0028] S42. Fabric similarity calculation: First, the Chinese fabric description is processed by word segmentation.

[0029] The formula for calculating term frequency (TF) is:

[0030]

[0031] The formula for calculating Inverse Document Frequency (IDF) is as follows:

[0032]

[0033] S43. Then convert it to a numerical vector and calculate the TF-IDF value of each word. The TF-IDF vector of the fabric description d is:

[0034] [TF(t1,d)×IDF(t1),TF(t2,d)×IDF(t2),…,TF(t n ,d)×IDF(t n )]

[0035] S44. Finally, cosine similarity is calculated only for fabric texts within the same category. The formula is:

[0036]

[0037] Where A·B is the dot product of vectors A and B, |A| and |B| are the magnitudes of vectors A and B respectively, and n is the dimension of the vector, which ranges from [-1, 1]. The closer the value is to 1, the more similar the two vectors are; the closer the value is to 0, the less similar the two vectors are.

[0038] S45. Price similarity calculation: Similarity is calculated only for prices within the same category. First, prices are directly treated as one-dimensional numerical vectors, i.e., price 1 is [price1], price 2 is [price2], and then the similarity between the two price vectors is calculated.

[0039]

[0040] S46. Calculate the confidence score: For clothing of the same category, normalize the importance of its features (garment type, fabric, and cost). Let the importance of features for garment type, fabric, and cost be w1, w2, and w3, respectively. The normalized weights are the type weights w... t Fabric weight w f , labor cost weight w cThe calculation follows the normalization formula:

[0041]

[0042] in Ensure that w1+w2+w3=1 to complete weight normalization;

[0043] As the weights of the similarity among the three, the weighted confidence score is calculated: Let the type similarity be s. t Fabric similarity is s f The similarity of labor costs is s. c Type weight w t Fabric weight w f , labor cost weight w c The confidence level C is then calculated as follows:

[0044] C = s t ×w t +s f ×w f +s c ×w c

[0045] S47. Calculate the confidence interval. For each sample, calculate the fabric similarity, type similarity, and cost similarity with other samples in the same category. Using the Percentile Bootstrap method, randomly draw N samples with replacement from the original dataset D to create a new Bootstrap sample set D. * This process is repeated B times, generating B Bootstrap sample sets. Here, B is set to 500. The 500 similarity values ​​are sorted, and the 0.5% quantile and 99.5% are used as the 99% confidence interval for the similarity. The average of the three similarity values ​​is taken as the confidence level, as shown in the following formula:

[0046]

[0047] Furthermore, step S5 includes the following steps:

[0048] S51. Compilation of calculation information: Based on the full amount of original data, all calculation information is obtained, including model number, actual total cost, calculated value, and confidence level;

[0049] S52. Data storage: Store the original pricing data, processed feature data, and all calculation results in the database;

[0050] S53. Visualization: Line charts show the comparison between calculated and actual values, bar charts show the ranking of feature importance, and heatmaps show the fabric similarity matrix of samples within the same category, intuitively presenting the calculation results and key influencing factors.

[0051] The present invention provides a cost calculation system and method for fast fashion women's ready-to-wear garments based on multi-source heterogeneous data, which has the following advantages:

[0052] 1. Improve measurement efficiency and accuracy

[0053] By integrating multi-source heterogeneous data and automating the processing, the system significantly reduces the time and errors of manual calculations, making it particularly suitable for the need for rapid cost calculation of large batches of women's clothing styles.

[0054] Using the random forest model can capture nonlinear relationships and improve the accuracy of the measurement results, which is superior to the traditional linear regression method.

[0055] 2. Maximize data utilization

[0056] Eleven core features were extracted from more than 70 original features, and textual and categorical data were transformed into modelable numerical forms through one-hot encoding, TF-IDF and other technologies to fully explore the value of the data.

[0057] Feature importance assessment (such as the weighting of transportation costs and fabric unit prices) helps companies identify key cost influencing factors and optimize resource allocation.

[0058] 3. Quantitatively assess reliability

[0059] A confidence assessment module is introduced, which calculates a 99% confidence interval through similarity weighting and the Bootstrap method, providing enterprises with the credibility level of the measurement results and supporting risk decision-making.

[0060] Visual representations (such as heatmaps and line graphs) intuitively present the deviation and confidence interval between the calculated and actual values, enhancing the interpretability of the results.

[0061] 4. Flexibility and scalability

[0062] The system is compatible with multiple data formats (Excel, CSV, text, etc.), and its modular design facilitates subsequent functional expansion (such as adding new features or models).

[0063] The evaluation is targeted by calculating similarity within the same category for women's clothing styles that are applicable to different categories.

[0064] 5. Optimize cost management

[0065] By prioritizing features (e.g., transportation costs having the highest proportion), companies can target cost drivers and develop more scientific pricing and production strategies.

[0066] The historical calculation results storage function supports long-term cost trend analysis and strategy adjustment.

[0067] In summary, this invention solves the problems of low efficiency, insufficient data utilization, and lack of quantifiable reliability in traditional methods, providing the fast fashion women's apparel industry with an efficient, accurate, and transparent cost calculation tool. Attached Figure Description

[0068] Figure 1 This is a flowchart of a fast fashion women's ready-to-wear cost calculation system and method based on multi-source heterogeneous data according to the present invention;

[0069] Figure 2 This is a comparison chart of the calculated and actual values ​​of the present invention;

[0070] Figure 3 A heatmap ranking the importance of features of this invention;

[0071] Figure 4 This is a scatter plot showing the confidence levels of the present invention. Detailed Implementation

[0072] To better understand the purpose, structure, and function of this invention, the following detailed description, in conjunction with the accompanying drawings, provides a fast fashion women's clothing cost calculation system and method based on multi-source heterogeneous data.

[0073] like Figure 1 As shown, this invention discloses a cost calculation system for fast fashion women's clothing based on multi-source heterogeneous data, comprising a raw data acquisition module, a data processing module, a calculation model module, a confidence assessment module, and a storage and display module. The raw data acquisition module integrates methods for acquiring data from multiple platforms, obtaining clothing pricing data containing fields such as style number, category, and fabric, and then formats and outputs it to the data processing module. The data processing module cleans the received pricing data, extracting 11 core features; it processes textual, categorical, and numerical features respectively, converting them into modelable numerical forms. The calculation model module uses the processed feature data as input and total cost as the target variable to construct a random forest calculation model. It calculates the mean, standard deviation, 99% confidence interval, and extended interval of the calculated values ​​through Bootstrap resampling, and outputs the calculation results. The confidence assessment module calculates the similarity of each indicator based on the feature data output by the data processing module, and takes the average value as the confidence level of the calculation result. The storage and display module stores the calculated values, confidence levels, similarity of each dimension, and feature importance in the database and displays them through visualization technology.

[0074] The present invention provides a method for cost calculation of fast fashion women's ready-to-wear garments based on multi-source heterogeneous data, comprising the following steps:

[0075] S1, the raw data acquisition module integrates multiple text files, Excel files, and CSV files obtained from multiple business systems, including the financial system, production system, and sales system. It extracts clothing pricing data containing fields such as style number, category, and fabric, verifies the data, and outputs the formatted data to the data processing module.

[0076] S11. Integrate multi-source data acquisition methods to obtain various pricing data from multiple business systems, including financial systems, production systems, and sales systems, using platforms such as text files, Excel, CSV, and databases as data interfaces. Encapsulate these functions to ensure compatibility with pricing data in different formats (such as tabular data containing fields such as style number, category, fabric, and workmanship).

[0077] S12. Validate the input pricing data. If there are any abnormal fields, throw a clear error message. If all fields are present, format the data into a uniform data frame, indexed by the style number, and output it to the data processing module.

[0078] In the S2 data processing module, the received pricing data is first cleaned, and 11 core features are extracted. Then, textual, categorical, and numerical features are processed separately, converted into modelable numerical forms, and the processed feature data is output to the calculation model module.

[0079] S21. Perform data cleaning. For the total cost column, delete invalid values ​​such as "not found" and "none" and reset the index. Convert the total cost column to floating point type. For numerical features (such as fabric unit price and transportation cost), first convert them to string type, remove non-digit and decimal characters using regular expressions, then replace invalid values ​​with NaN, convert them to numeric type, and fill missing values ​​with the median. Categorical features and text features retain their original format for subsequent processing.

[0080] S22. Perform feature processing: use one-hot encoding for categorical features to generate a data frame; for text features, perform word segmentation and convert them into TF-IDF vectors to form a text feature matrix; merge the processed categorical features, text features and numerical features by row to construct a complete feature matrix and output it to the measurement model module.

[0081] In the S3 calculation model module, the processed feature data is used as input and the total cost is used as the target variable to construct a random forest calculation model, calculate the calculated value of the garment price, and output the calculation results to the confidence assessment module and the storage and display module.

[0082] S31. Model Training and Dataset Partitioning: A random forest model is used, trained with features including classification, fabric, fabric unit price, fabric unit consumption, accessory cost, labor cost, tax rate, transportation, and total cost. The "80 / 20 rule" is used to allocate the test set and training set, with the first 80% of the data serving as the training set D. train The last 20% of the data was used as the test set D. test Set the number of decision trees in the random forest to 100, the maximum depth to None, and train the model using the default parameters.

[0083] S32. Calculation of Measurement Results: For the test set samples, the random forest model is used to output the measurement results of the garment prices. The output of the garment price measurement results ensures that each style number can obtain a corresponding measurement value.

[0084] S33. Feature Importance Evaluation: In random forests, feature importance is calculated based on the reduction in Gini impurity. For each feature, its Gini impurity is... The reduction in Gini impurity is ΔG = G 父节点 -(ω·G 子节点1 +ω·G 子节点2 The importance of this feature is the average of the sum of the reductions in Gini impurity caused by this feature across all decision trees.

[0085] Table 1. Ranking of Feature Importance

[0086]

[0087] In the S4 confidence assessment module, based on the feature data output by the data processing module, the similarity of garment type, fabric, and labor cost is calculated separately. These three similarities are then fused, and the result after feature importance normalization is used as the weight to calculate the confidence level of the measurement result. The Percentile Bootstrap method is used to calculate the 99% confidence interval, and the result is then output to the storage and display module.

[0088] S41. For garment type similarity calculation, the similarity within the same category is recorded as 1 (e.g., the type similarity between tops is directly recorded as 1.0), and no similarity is calculated between different categories (only subsequent confidence assessment is performed within the same category), thereby improving the efficiency and accuracy of category similarity assessment.

[0089] S42. Fabric similarity calculation: First, the Chinese fabric description is processed by word segmentation (e.g., "98 polyester 2 spandex blend" is split into "98 polyester", "2 spandex", and "blended").

[0090] The formula for calculating term frequency (TF) is:

[0091]

[0092] The formula for calculating Inverse Document Frequency (IDF) is as follows:

[0093]

[0094] S43. Then convert it to a numerical vector and calculate the TF-IDF value of each word. The TF-IDF vector of the fabric description d is:

[0095] [TF(t1,d)×IDF(t1),TF(t2,d)×IDF(t2),…,TF(t n ,d)×IDF(t n )]

[0096] S44. Finally, cosine similarity is calculated only for fabric texts within the same category. The formula is:

[0097]

[0098] Where A·B is the dot product of vectors A and B, |A| and |B| are the magnitudes of vectors A and B respectively, and n is the dimension of the vector. Its value ranges between [-1, 1]. The closer the value is to 1, the more similar the two vectors are; the closer the value is to 0, the less similar the two vectors are.

[0099] S45. Price similarity calculation: Similarity is calculated only for prices within the same category. First, prices are directly treated as one-dimensional numerical vectors, i.e., price 1 is [price1], price 2 is [price2], and then the similarity between the two price vectors is calculated.

[0100]

[0101] S46. Calculate the confidence score: For clothing of the same category, normalize the feature importance of garment type, fabric, and cost. Let the feature importance of garment type, fabric, and cost be w1, w2, and w3, respectively. The normalized weights (i.e., type weights w1, w2, and w3) are calculated as follows: t Fabric weight w f , labor cost weight w c The calculation follows the normalization formula:

[0102]

[0103] in Ensure w1+w2+w3=1 to complete weight normalization.

[0104] As the weights of the similarity among the three, the weighted confidence score is calculated: Let the type similarity be s. t Fabric similarity is s f The similarity of labor costs is s. c Type weight w tFabric weight w f , labor cost weight w c The confidence level C is then calculated as follows:

[0105] C = s t ×w t +s f ×w f +s c ×w c

[0106] S47. Calculate the confidence interval. For each sample and other samples within the same category, calculate the fabric similarity, type similarity (fixed at 1.0), and cost similarity. Using the Percentile Bootstrap method, randomly draw N samples with replacement from the original dataset D to create a new Bootstrap sample set D. * This process is repeated B times, generating B Bootstrap sample sets. Here, B is set to 500. The 500 similarity values ​​are sorted, and the 0.5% quantile and 99.5% are used as the 99% confidence interval for the similarity. The average of the three similarity values ​​is then used as the confidence level. The formula is as follows:

[0107]

[0108] In the S5 storage and display module, visualization technology is used to display the calculated values, confidence levels, similarity of each dimension, and feature importance on the front end, and then save them to the database.

[0109] S51. Compilation of calculation information: Based on the full amount of original data, all calculation information is obtained, including model number, actual total cost, calculated value, and confidence level;

[0110] S52. Data storage: Store the original pricing data, processed feature data, and all calculation results (including item number, actual value, calculated value, confidence level, etc.) in the database;

[0111] S53, Visual Display: such as Figure 2-4 As shown, a line chart is used to compare the calculated values ​​with the actual values ​​(the horizontal axis represents the style number and the vertical axis represents the cost), a bar chart is used to show the ranking of feature importance, and a heat map is used to show the fabric similarity matrix of samples within the same category, intuitively presenting the calculation results and key influencing factors.

[0112] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A cost calculation system for fast fashion women's ready-to-wear garments based on multi-source heterogeneous data, characterized in that, It includes a raw data acquisition module, a data processing module, a calculation model module, a confidence assessment module, and a storage and display module. The raw data acquisition module integrates methods for acquiring data from multiple platforms to obtain clothing pricing data and outputs it in a formatted manner to the data processing module. The data processing module cleans the received pricing data and extracts 11 core features. Textual, categorical, and numerical features are processed separately and converted into modelable numerical forms. The calculation model module uses the processed feature data as input and the total cost as the target variable to construct a random forest calculation model. It calculates the mean, standard deviation, 99% confidence interval, and extended interval of the calculated values ​​through Bootstrap resampling and outputs the calculation results. The confidence assessment module calculates the similarity of each indicator based on the feature data output by the data processing module and takes the average value as the confidence level of the calculation results. The storage and display module stores the calculation results in the database and displays them through visualization technology.

2. A method for calculating the cost of fast fashion women's ready-to-wear garments using a cost calculation system based on multi-source heterogeneous data as described in claim 1, characterized in that, Includes the following steps: S1, the raw data acquisition module integrates multiple text files, Excel files, and CSV files obtained from multiple business systems, extracts clothing pricing data from them, verifies the data, and outputs the formatted data to the data processing module; In the S2 data processing module, the received pricing data is first cleaned and 11 core features are extracted. Then, the textual, categorical, and numerical features are processed separately and converted into a modelable numerical form. The processed feature data is then output to the calculation model module. In the S3 calculation model module, the processed feature data is used as input and the total cost is used as the target variable to construct a random forest calculation model, calculate the calculated value of the garment price, and output the calculation results to the confidence assessment module and the storage and display module. In the S4 confidence assessment module, based on the feature data output by the data processing module, the similarity of garment type, fabric, and labor cost is calculated respectively. The three similarities are fused, and the result after feature importance normalization is used as the weight to calculate the confidence of the measurement result. The Percentile Bootstrap method is used to calculate the 99% confidence interval and output it to the storage and display module. In the S5 storage and display module, the calculation results are displayed on the front end using visualization technology and then saved to the database.

3. The method for calculating the cost of fast fashion women's ready-to-wear garments according to claim 2, characterized in that, S1 includes the following steps: S11. Integrate multi-source data acquisition methods to obtain various pricing data from multiple business systems, and encapsulate these functions to ensure compatibility with pricing data in different formats; S12. Validate the input pricing data. If any abnormal fields are found, throw a clear error message. If all fields are present, the data will be formatted into a uniform data frame, indexed by the style number, and output to the data processing module.

4. The method for calculating the cost of fast fashion women's ready-to-wear garments according to claim 2, characterized in that, S2 includes the following steps: S21. Perform data cleaning. For the total cost column, delete invalid values ​​and reset the index, and convert the total cost column to floating point type. For numeric features, first convert them to string type, remove non-digit and decimal characters using regular expressions, then replace invalid values ​​with NaN, convert them to numeric type, and fill missing values ​​with the median. Classification features and text features retain their original format for further processing. S22. Perform feature processing: use one-hot encoding for categorical features to generate a data frame; for text features, perform word segmentation and convert them into TF-IDF vectors to form a text feature matrix; merge the processed categorical features, text features and numerical features by row to construct a complete feature matrix and output it to the measurement model module.

5. The method for calculating the cost of fast fashion women's ready-to-wear garments according to claim 2, characterized in that, S3 includes the following steps: S31. Model Training and Dataset Partitioning: A random forest model is used, with features such as classification, fabric, fabric unit price, fabric unit consumption, auxiliary material cost, labor cost, tax rate, transportation, and total cost. The test set and training set are allocated using the "80 / 20 rule". S32. Calculation of measurement results: For the test set samples, use the random forest model to output the measurement results of the garment price, so that each style number can get a corresponding measurement value. S33. Feature Importance Evaluation: In random forests, feature importance is calculated based on the reduction in Gini impurity. For each feature, its Gini impurity is... The reduction in Gini impurity is ΔG = G 父节点 -(ω·G 子节点1 +ω·G 子节点2 The importance of this feature is the average of the sum of the reductions in Gini impurity caused by this feature across all decision trees.

6. The method for calculating the cost of fast fashion women's ready-to-wear garments according to claim 2, characterized in that, S4 includes the following steps: S41. For garment type similarity calculation, the similarity within the same category is recorded as 1, and no similarity is calculated for different categories; S42. Fabric similarity calculation: First, the Chinese fabric description is processed by word segmentation. The formula for calculating term frequency (TF) is: The formula for calculating Inverse Document Frequency (IDF) is as follows: S43. Then convert it to a numerical vector and calculate the TF-IDF value of each word. The TF-IDF vector of the fabric description d is [TF(t1,d)×IDF(t1),TF(t2,d)×IDF(t2),…,TF(t…]. n ,d)×IDF(t n )] S44. Finally, cosine similarity is calculated only for fabric texts within the same category. The formula is: Where A·B is the dot product of vectors A and B, |A| and |B| are the magnitudes of vectors A and B respectively, and n is the dimension of the vector, which ranges from [-1, 1]. The closer the value is to 1, the more similar the two vectors are; the closer the value is to 0, the less similar the two vectors are. S45. Price similarity calculation: Similarity is calculated only for prices within the same category. First, prices are directly treated as one-dimensional numerical vectors, i.e., price 1 is [price1], price 2 is [price2], and then the similarity between the two price vectors is calculated. S46. Calculate the confidence score: For clothing of the same category, normalize the importance of its features (garment type, fabric, and cost). Let the importance of features for garment type, fabric, and cost be w1, w2, and w3, respectively. The normalized weights are the type weights w... t Fabric weight w f , labor cost weight w c The calculation follows the normalization formula: in Ensure that w1+w2+w3=1 to complete weight normalization; As the weights of the similarity among the three, the weighted confidence score is calculated: Let the type similarity be s. t Fabric similarity is s f The similarity of labor costs is s. c Type weight w t Fabric weight w f , labor cost weight w c The confidence level C is then calculated as follows: C=s t ×w t +s f ×w f +s c ×w c S47. Calculate the confidence interval. For each sample, calculate the fabric similarity, type similarity, and cost similarity with other samples in the same category. Using the Percentile Bootstrap method, randomly draw N samples with replacement from the original dataset D to create a new Bootstrap sample set D. * This process is repeated B times, generating B Bootstrap sample sets. Here, B is set to 500. The 500 similarity values ​​are sorted, and the 0.5% quantile and 99.5% are used as the 99% confidence interval for the similarity. The average of the three similarity values ​​is taken as the confidence level, as shown in the following formula:

7. The method for calculating the cost of fast fashion women's ready-to-wear garments according to claim 2, characterized in that, S5 includes the following steps: S51. Compilation of calculation information: Based on the full amount of original data, all calculation information is obtained, including model number, actual total cost, calculated value, and confidence level; S52. Data storage: Store the original pricing data, processed feature data, and all calculation results in the database; S53. Visualization: Line charts show the comparison between calculated and actual values, bar charts show the ranking of feature importance, and heatmaps show the fabric similarity matrix of samples within the same category, intuitively presenting the calculation results and key influencing factors.