Risk early warning method for bulk commodities
Through multi-source data association consistency detection and cross-feature extraction, combined with feature interaction modeling and principal component analysis, a dynamic risk scoring curve for commodities is generated, which solves the problem that existing technologies cannot fully consider multiple market factors and achieves fast and accurate risk warning.
Patent Information
- Application Number
- CN202510816692.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies are unable to fully consider the complex interactions of multiple market factors in commodity markets, resulting in inconsistent data quality, oversimplified models and slow responses to emergencies when market sentiment fluctuates drastically, information asymmetry or emergencies occur.
By acquiring multi-source data for association consistency detection, extracting cross-feature candidate sets, performing feature interaction modeling and principal component analysis, generating comprehensive feature vectors, cross-validating using risk assessment sub-models, generating dynamic risk scoring curves and generating risk warning reports.
It improves the response speed and accuracy of commodity risk warnings, can comprehensively identify multiple risk factors, deeply explore the relationship between data sources, enhance the accuracy and robustness of risk assessment models, quickly identify potential risks and generate real-time warnings.
Smart Images

Figure CN120707185A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of bulk commodities, and in particular to a risk early warning method for bulk commodities. Background Art
[0002] Commodities, a vital component of the global economy, encompass a wide range of resource commodities, including oil, natural gas, metals, and agricultural products. Due to their high price volatility, complex market supply and demand relationships, and numerous risk factors, effective risk warning and timely understanding of market trends are crucial for investors, manufacturers, and government agencies. In recent years, with the deepening of global economic integration, commodity markets have been impacted by multiple factors, including international policies, emergencies, and climate change, leading to more intense price fluctuations and placing higher demands on market stability and forecast accuracy.
[0003] Among relevant technical approaches, existing technologies for commodity risk early warning primarily rely on traditional statistical analysis methods, economic models, and machine learning techniques. Common approaches include leveraging historical price and trading volume data to predict future price fluctuations through traditional methods such as regression analysis, time series models (such as ARIMA), and volatility models (such as GARCH). Additionally, market risk is estimated using economic indicators (such as exchange rates, interest rates, and inventory levels).
[0004] Regarding the above-mentioned technical solutions, although existing technical solutions can predict commodity price fluctuations through regression analysis of historical data or machine learning modeling, helping market participants understand risk levels and reduce potential losses, in actual applications, traditional methods usually rely on a single data source (such as historical prices or a single economic indicator) and cannot fully consider the complex interactions of multiple market factors. As a result, when market sentiment changes drastically, there is information asymmetry or the influence of sudden events, there are problems such as inconsistent data quality, over-simplified models and slow response to sudden events. Summary of the Invention
[0005] In order to improve the problems in actual applications where traditional methods are unable to fully consider the complex interactions of multiple market factors, resulting in inconsistent data quality, oversimplified models, and slow response to emergencies when market sentiment changes drastically, information asymmetry, or emergencies occur, this application provides a risk warning method for commodities.
[0006] The present invention provides a risk warning method for bulk commodities, comprising: acquiring multi-source data of target bulk commodities, performing correlation consistency detection on the multi-source data to obtain a high-credibility data set; extracting a cross-feature candidate set from the high-credibility data set, performing feature interaction modeling on the cross-feature candidate set and the historical fluctuation characteristics of the bulk commodities to obtain a feature interaction matrix, performing principal component analysis and information gain evaluation on the feature interaction matrix to obtain a comprehensive feature vector; generating different structured input sequences based on the comprehensive feature vector, inputting all the structured input sequences into a preset corresponding risk assessment sub-model to obtain a local risk response matrix and a correlation factor sensitivity matrix, cross-validating the local risk response matrix and the correlation factor sensitivity matrix to generate a comprehensive risk assessment result; generating a dynamic risk scoring curve for the target bulk commodity based on the comprehensive risk assessment result, judging the risk warning level according to the dynamic risk scoring curve, and generating a risk warning report based on the risk warning level.
[0007] As a preferred solution, the steps of obtaining multi-source data of a target commodity, performing association consistency testing on the multi-source data, and obtaining a high-credibility data set include: obtaining multi-source data of the target commodity within a preset time window, wherein the multi-source data includes historical transaction data, macroeconomic indicator data, news information data, and social media text data; performing time alignment processing on the historical transaction data and the macroeconomic indicator data to obtain structured time series data; performing text cleaning and entity recognition processing on the news information data and the social media text data to obtain structured text feature data; performing association consistency testing on the structured time series data and the structured text feature data to obtain time consistency features and semantic consistency features; using the time consistency features and the semantic consistency features to correct and remove outliers in the historical transaction data, the macroeconomic indicator data, the news information data, and the social media text data to construct a corrected multi-source data set; performing statistical stability testing and cointegration analysis on the corrected multi-source data set to obtain a correlation matrix and confidence weights, screening out data source samples based on the correlation matrix and the confidence weights, and constructing a high-credibility data set based on the data source samples.
[0008] As a preferred solution, the steps of extracting a cross-feature candidate set from the high-credibility data set, performing feature interaction modeling on the cross-feature candidate set and the historical fluctuation characteristics of commodities to obtain a feature interaction matrix, performing principal component analysis and information gain evaluation on the feature interaction matrix to obtain a comprehensive feature vector include: extracting cross-domain numerical feature pairs between structured data and sentiment trend change rates of unstructured text data based on the high-credibility data set, generating a cross-feature candidate set based on the cross-domain numerical feature pairs and the sentiment trend change rates; performing feature interaction modeling on the highly significant cross items in the cross-feature candidate set and the historical fluctuation characteristics of commodities to obtain a feature interaction matrix, and performing principal component analysis and information gain evaluation on the feature interaction matrix to obtain a comprehensive feature vector. The feature interaction matrix is subjected to principal component analysis and information gain evaluation to construct a cross-feature vector set, wherein the highly significant cross-term refers to a significantly correlated feature combination in a cross-domain data source; the cross-feature vector set is subjected to dynamic window clustering and hierarchical encoding in chronological order to obtain a feature evolution vector and a feature structure label, and a multi-level feature fusion framework is constructed using the feature evolution vector and the feature structure label; the multi-level feature fusion framework is combined with the industry fluctuation cycle to perform feature grouping to obtain a hierarchical fusion feature set, and a comprehensive feature vector is generated based on the hierarchical fusion feature set, wherein the industry fluctuation cycle refers to a time parameter determined according to the seasonal and cyclical laws within the commodity industry.
[0009] As a preferred solution, the steps of extracting cross-domain numerical feature pairs and sentiment trend change rates of unstructured text data between structured data based on the high-credibility data set, and generating a cross-feature candidate set based on the cross-domain numerical feature pairs and the sentiment trend change rates include: performing multivariate Pearson correlation coefficient analysis and Granger causality test on the structured data in the high-credibility data set to extract cross-domain numerical feature pairs, wherein the structured data is historical transaction data and macroeconomic indicator data; performing sentiment analysis on the unstructured text data in the high-credibility data set, calculating the sentiment score of each text fragment, and generating a sentiment trend change rate based on the sentiment score of each fragment, wherein the unstructured text data is news information data and social media text data; performing time series alignment and weighted integration on the sentiment trend change rate and the cross-domain numerical feature pairs to obtain a fused sentiment and numerical composite feature set, inputting the sentiment and numerical composite feature set into a preset rolling window regression model to obtain key cross-terms; performing information gain analysis on the key cross-terms to screen out cross-feature terms, and generating a cross-feature candidate set based on the cross-feature terms.
[0010] As a preferred solution, the steps of generating different structured input sequences based on the comprehensive feature vector, inputting all the structured input sequences into a preset corresponding risk assessment sub-model to obtain a local risk response matrix and a correlation factor sensitivity matrix, and cross-validating the local risk response matrix and the correlation factor sensitivity matrix to generate a comprehensive risk assessment result include: time slicing and context-enhancing the comprehensive feature vector according to the preset model input requirements to obtain different structured input sequences, inputting all the structured input sequences into the corresponding preset risk assessment sub-model to obtain a number of risk response data and a number of correlation factor data; using all the risk response data to construct a local risk response matrix, performing a horizontal comparison analysis on the local risk response matrix to obtain a factor inconsistency index; using all the correlation factor data to construct a correlation factor sensitivity matrix, performing weighted averaging and normalization on the correlation factor sensitivity matrix to obtain a factor influence distribution map; cross-validating the factor inconsistency index with the factor influence distribution map to obtain a credible risk signal set and a model confidence deviation set, and generating a comprehensive risk assessment result based on the credible risk signal set and the model confidence deviation set.
[0011] As a preferred solution, the steps of time slicing and context enhancing the comprehensive feature vector according to the preset model input requirements to obtain different structured input sequences, inputting all the structured input sequences into the corresponding preset risk assessment sub-model to obtain a number of risk response data and a number of correlation factor data include: dividing the comprehensive feature vector into multiple time windows through a sliding window algorithm to obtain a number of time slice sequences, in each of the time slice sequences, context enhancement is performed in combination with the context information of external macroeconomic events and industry cycles to obtain an enhanced time slice sequence; standardizing all the enhanced time slice sequences, inputting all the standardized time slices into the corresponding preset risk assessment sub-model, and outputting the corresponding number of risk response data and a number of correlation factor data.
[0012] As a preferred solution, the steps of generating a dynamic risk scoring curve for the target commodity based on the comprehensive risk assessment results, determining the risk warning level according to the dynamic risk scoring curve, and generating a risk warning report based on the risk warning level include: performing risk index quantification processing on the comprehensive risk assessment results to obtain a multi-dimensional risk score subset, normalizing the multi-dimensional risk score subset according to preset indicator categories to generate a dynamic risk scoring curve for the target commodity; calculating the score volatility and the score jump factor based on the dynamic risk scoring curve, inputting the score volatility and the score jump factor into a preset risk level threshold set for mapping to obtain the current risk warning level, and generating a risk warning report based on the risk warning level.
[0013] As a preferred solution, the step of calculating the score volatility and the score jump factor based on the dynamic risk scoring curve, inputting the score volatility and the score jump factor into a preset risk level threshold set for mapping, and obtaining the current risk warning level includes: performing a time series analysis on the dynamic risk scoring curve, calculating the score volatility, measuring the score stability of the dynamic risk scoring curve by a weighted coefficient of variation method to obtain a scoring curve, calculating the jump amplitude and frequency of the score in the scoring curve according to a differential algorithm to obtain a score jump factor; inputting the score volatility and the score jump factor into a preset risk level threshold set to map the current score of the dynamic risk scoring curve to obtain the current risk warning level.
[0014] Compared with the prior art, the present application has the following beneficial effects: fast response speed and accurate early warning. By combining the correlation consistency detection and cross-feature extraction of multi-source data, it is possible to comprehensively identify a variety of risk factors that affect commodity price fluctuations, and through feature interaction modeling, the relationship between historical fluctuation characteristics and other data sources can be deeply explored. Through principal component analysis and information gain evaluation, the most predictive features are further screened out, and an efficient risk assessment system is constructed. The structured input sequence generated based on the comprehensive feature vector can further enhance the accuracy of the risk assessment model, and the robustness of the model is ensured by cross-validation. After generating the dynamic risk scoring curve, the potential risks in the commodity market can be quickly identified by judging the risk warning level, and a warning report is generated in real time to improve the situation. In actual applications, traditional methods cannot fully consider the complex interactions of multiple market factors, resulting in inconsistent data quality, oversimplified models, and slow response to emergencies under the influence of drastic changes in market sentiment, information asymmetry, or emergencies. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not intended to limit the conditions under which the present invention can be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportional relationships, or adjustments in size should still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and objectives that can be achieved by the present invention.
[0017] Figure 1 It is a flowchart of a commodity risk warning method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0020] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0021] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0022] The technical solution of the present invention will be further described below with reference to the accompanying drawings and through specific implementation methods.
[0023] Example 1: like Figure 1 As shown, the present application provides a commodity risk warning method, including steps S100 to S400.
[0024] Step S100: Acquire multi-source data of target commodities, perform correlation consistency detection on the multi-source data, and obtain a high-reliability data set.
[0025] In this step, we acquire commodity-related data from various sources to ensure data diversity and comprehensiveness. Specifically, this data includes historical price data, trading volume data, macroeconomic data (such as interest rates, exchange rates, and inventory levels), news data (such as financial news and policy changes), and social media data (such as market sentiment data from Twitter and Weibo). This data comes from different systems and data sources, so it is necessary to perform correlation consistency testing to eliminate errors or conflicts between different data sources and obtain a highly reliable data set.
[0026] For example, historical price data can be delayed due to different collection times, social media data can contain noise or sentiment bias, and macroeconomic data can be intermittently missing or delayed. Therefore, we first use time alignment techniques and correlation analysis methods to ensure that data from different time scales are synchronized within the same time window. Next, we use data consistency detection algorithms (such as the Pearson correlation coefficient or cointegration tests) to verify the correlation and consistency between different data sources, thereby removing outliers and invalid data, ultimately obtaining a highly reliable multi-source data set.
[0027] Step S200: extract a cross-feature candidate set from a high-confidence data set, perform feature interaction modeling on the cross-feature candidate set and the historical volatility characteristics of commodities to obtain a feature interaction matrix, perform principal component analysis and information gain evaluation on the feature interaction matrix to obtain a comprehensive feature vector.
[0028] In this step, we conduct in-depth analysis of high-confidence data sets to extract cross-features from various data sources. Through feature interaction modeling, we derive the relationships between the historical volatility characteristics of commodities and other data sources. Specifically, we first use multivariate statistical methods (such as the Pearson correlation coefficient or mutual information method) to extract cross-domain numerical feature pairs between structured data (such as the relationship between historical prices and macroeconomic indicators) and the rate of change of sentiment trends in unstructured text data (such as market sentiment fluctuations in social media data). These cross-features are then combined with the historical volatility characteristics of commodities through feature interaction modeling to obtain a feature interaction matrix.
[0029] For example, for historical price and volume data, feature interactions include linear or nonlinear relationships between price and volume, and the relationship between sentiment fluctuations in news data and price fluctuations. Principal component analysis (PCA) and information gain evaluation are performed on the resulting feature interaction matrix to identify features that best reflect risk changes. Based on these features, a comprehensive feature vector is generated for subsequent model training.
[0030] Step S300: Generate different structured input sequences based on the comprehensive feature vector, input all structured input sequences into the preset corresponding risk assessment sub-model, obtain the local risk response matrix and the correlation factor sensitivity matrix, cross-validate the local risk response matrix and the correlation factor sensitivity matrix to generate a comprehensive risk assessment result.
[0031] In this step, different structured input sequences are generated based on the obtained comprehensive feature vector and input into the pre-set risk assessment sub-model to obtain a local risk response matrix and an associated factor sensitivity matrix. Specifically, the comprehensive feature vector is first time-sliced and context-enhanced to generate structured input sequences for multiple time periods. These input sequences are then input into the pre-set risk assessment sub-model (such as a support vector machine (SVM), decision tree, or random forest). Based on the input sequence, the sub-model outputs a local risk response matrix (representing the impact of each feature on the risk) and an associated factor sensitivity matrix (representing the sensitivity of different risk factors to the outcome).
[0032] For example, using a sliding time window approach, a feature set can be generated for each time slice to assess commodity risk. The resulting local risk response matrix and associated factor sensitivity matrix are then verified and optimized using cross-validation techniques to produce the final comprehensive risk assessment results.
[0033] Step S400: Generate a dynamic risk scoring curve for the target commodity based on the comprehensive risk assessment results, determine the risk warning level based on the dynamic risk scoring curve, and generate a risk warning report based on the risk warning level.
[0034] In this step, a dynamic risk score curve is generated for the target commodity based on the comprehensive risk assessment results. Specifically, the local risk response matrix and the correlation factor sensitivity matrix are quantified and converted into a risk score. The corresponding score volatility and jump factor are then calculated. Based on the score volatility and jump factor, a mapping algorithm (such as the weighted coefficient of variation method) is used to generate a dynamic risk score curve, which is then used to determine the risk warning level. Finally, a risk warning report is generated, which includes the risk level, potential risk sources, and response strategies.
[0035] For example, suppose the risk score curve of a certain commodity shows large fluctuations at a certain moment, and the score suddenly changes. At this time, the potential risks faced by the commodity can be judged by mapping the risk level threshold (such as the risk warning level is high risk), and corresponding risk response recommendations can be provided based on actual data.
[0036] In this embodiment, a high-credibility data set is obtained by acquiring multi-source data of the target commodity and performing correlation consistency detection on the multi-source data; then, a cross-feature candidate set is extracted from the high-credibility data set, and feature interaction modeling is performed on the cross-feature candidate set and the historical fluctuation characteristics of the commodity to obtain a feature interaction matrix; then, principal component analysis and information gain evaluation are performed on the feature interaction matrix to obtain a comprehensive feature vector; different structured input sequences are generated based on the comprehensive feature vector, and all the structured input sequences are input into a preset corresponding risk assessment sub-model to obtain a local risk response matrix and a correlation factor sensitivity matrix; next, the local risk response matrix and the correlation factor sensitivity matrix are cross-validated to generate a comprehensive risk assessment result; finally, a dynamic risk scoring curve of the target commodity is generated based on the comprehensive risk assessment result, the risk warning level is determined according to the dynamic risk scoring curve, and a risk warning report is generated based on the risk warning level.
[0037] By comprehensively analyzing multi-source data on target commodities, the accuracy and real-time nature of commodity risk warnings have been effectively improved. Through correlation consistency detection and cross-feature extraction of multi-source data, the high credibility of the data used can be ensured, and the complex relationships between different data sources can be deeply explored, overcoming the limitations of traditional methods that rely on a single data source. The combination of feature interaction modeling and principal component analysis makes the relationship between the historical volatility characteristics of commodities and multi-source data clearer, improving the model's ability to identify risk patterns. The multi-level structured input sequence generated based on the comprehensive feature vector can enhance the accuracy of the risk assessment sub-model. The cross-validation results further improve the robustness of risk assessment and improve practical applications. Traditional methods are unable to fully consider the complex interactions of multiple market factors, resulting in inconsistent data quality, oversimplified models, and slow responses to emergencies when market sentiment fluctuates drastically, information asymmetry occurs, or emergencies occur.
[0038] Example 2: In step S100, the steps of obtaining multi-source data of the target commodity, performing correlation consistency detection on the multi-source data, and obtaining a high-reliability data set specifically include: Obtain multi-source data on the target commodity within a preset time window, where the multi-source data includes historical transaction data, macroeconomic indicator data, news information data, and social media text data.
[0039] When acquiring multi-source data on target commodities, the first step is to determine the data's time window. This time window can be preset based on specific market requirements, for example, by daily, weekly, or monthly granularity. Within the preset time window, historical transaction data and macroeconomic indicator data are retrieved through designated API interfaces or database access modules. Historical transaction data includes features such as price and trading volume, while macroeconomic indicator data includes interest rate changes, exchange rate fluctuations, inflation indicators, and inventory levels. Furthermore, news and information data is captured using crawler technology. This data is sourced from publicly available and reliable financial news websites or government announcements. Furthermore, relevant market sentiment data is obtained through social media data interfaces (such as the Twitter API or similar frameworks), allowing for real-time capture and storage of social discussions surrounding the target commodity.
[0040] For example, captured historical transaction data is stored in CSV format, recording daily transaction prices and volumes. Macroeconomic indicator data is stored in a database (such as MySQL), with each record including a timestamp and the corresponding macro variable value. News and social media data is stored as JSON files, containing news headlines, release times, social media posts, and sentiment analysis preprocessing results. This unified collection and storage of multi-source data ensures comprehensive and accurate basic data for subsequent processing.
[0041] Perform time alignment processing on historical transaction data and macroeconomic indicator data to obtain structured time series data.
[0042] Historical transaction data and macroeconomic indicator data are aligned along the time dimension to ensure consistent time labels across different data sources. First, the transaction data is standardized, and missing trading day data (e.g., missing holidays) is interpolated. Then, time window alignment is performed on the macroeconomic indicator data, using linear interpolation or exponential smoothing algorithms to break down long-term data (e.g., monthly interest rate changes) into shorter time granularity (e.g., daily data). After time alignment, a sliding window is used to construct time series data, representing data trends within the sliced time window. The resulting structured time series data is ready for subsequent analysis.
[0043] For example, if transaction data is daily and macroeconomic indicator data is monthly, linear interpolation can be applied to expand the monthly data to daily data, while ensuring that each day's transaction data and macroeconomic indicator data have the same time labels. Furthermore, the Python library Pandas can be used to implement time window slicing, such as the pd.resample() function, to process monthly data, convert it to daily data, and generate a time series file in a unified format.
[0044] Perform text cleaning and entity recognition on news information data and social media text data to obtain structured text feature data.
[0045] Noise is removed from news and social media text data, and valuable features are extracted from the text using standardized processing methods. First, the data is segmented to convert it into a series of words. This step utilizes existing natural language processing libraries (such as NLTK or SpaCy) for word segmentation and stop word removal. Secondly, part-of-speech tagging and entity recognition techniques (such as named entity recognition (NER)) are used to extract key entity information, such as company names, policy terms, and economic indicators. Finally, keyword extraction algorithms (such as TF-IDF) are used to identify high-importance terms in the text and generate a structured text feature file.
[0046] For example, the keyword "Federal Reserve interest rate hike" was detected in news data, and the evaluative term "oil supply and demand balance" was detected in social media data. These keywords were assigned corresponding importance weights using the TF-IDF algorithm. Furthermore, NER recognition analysis classified "Federal Reserve" as an institutional entity and "interest rate hike" as a policy term. This method generates a structured text feature file, in which each record contains the time tag of the text feature, a keyword list, and feature weight information.
[0047] The structured time series data and structured text feature data are correlated and checked for consistency to determine whether there is information content deviation within the same time window, and the temporal consistency features and semantic consistency features are obtained.
[0048] For structured time series data and structured text feature data, the data distribution consistency within the same time window is calculated. For time series data, the stability is analyzed using standard deviation and mean change rate. For text feature data, the relevance of text content within the same time window is assessed based on positional similarity (Jaccard similarity or cosine similarity). These results are then cross-validated to identify data groups that exhibit consistent or mutually supportive behavior, and to construct temporal consistency features and semantic consistency features.
[0049] For example, by calculating the standard deviation of daily trading price fluctuations in time series data, we found that price fluctuations were small within a certain time window, which is highly semantically consistent with the evaluation of a "stable market" in news text from the same period. Using cosine similarity to calculate the semantic relevance between news documents, we screened out data windows with consistent market evaluations, effectively marking them as temporally consistent and semantically consistent features.
[0050] Temporal consistency features and semantic consistency features are used to correct and remove anomalies in historical transaction data, macroeconomic indicator data, news information data, and social media text data to construct a corrected multi-source data set.
[0051] Outlier detection is performed based on temporal and semantic consistency features, identifying sudden extreme values in historical trading data, unusual fluctuations in macroeconomic indicators, erroneous or false information in news data, and extreme emotional content in social media data. For trading data, box plots are used to detect extreme values, while for economic indicators, a sliding mean method is used to identify data points outside the normal range. For text data, sentiment analysis models are used to remove content with extreme emotions or semantic inconsistencies. Finally, a revised multi-source data set is constructed for subsequent analysis.
[0052] For example, the box plot method was used to detect abnormal peaks in transaction prices, and combined with the "market panic" sentiment analysis in the news data, the information was confirmed as an outlier; abnormal fluctuations in inventory data were detected through the sliding mean and the outlier was eliminated; by filtering out extreme emotional posts in social media, a revised multi-source data set was constructed.
[0053] The revised multi-source data set is subjected to statistical stability testing and cointegration analysis to obtain the correlation matrix and confidence weights between the data sources. Based on the correlation matrix and confidence weights, the data source samples are screened out, and a high-credibility data set is constructed based on the data source samples.
[0054] The revised multi-source data set is then tested for fluctuation patterns and stability through statistical stability testing and cointegration analysis. A correlation matrix is constructed for each data source in the data set, and the correlations between data sources are calculated using the Pearson correlation coefficient. Confidence interval analysis (e.g., confidence weighting) is also used to assess the importance of each data source to the overall analysis results. The correlation matrix and confidence weights are then screened to select sample data sources, ultimately constructing a highly reliable data set.
[0055] For example, Pearson correlation coefficients were calculated for revised price and trading volume data, revealing a high correlation. Cointegration analysis of news and social media data confirmed their long-term stable synergy. Confidence interval analysis was used to retain samples with higher weights to construct a high-confidence data set.
[0056] In step S200, a cross-feature candidate set is extracted from a high-credibility data set, a feature interaction model is performed between the cross-feature candidate set and the historical volatility characteristics of commodities to obtain a feature interaction matrix, and principal component analysis and information gain evaluation are performed on the feature interaction matrix to obtain a comprehensive feature vector. Specifically, the steps include: Based on the high-confidence data set, cross-domain numerical feature pairs between structured data and the sentiment trend change rate of unstructured text data are extracted, and a cross-feature candidate set is generated according to the cross-domain numerical feature pairs and the sentiment trend change rate.
[0057] Feature extraction and fusion are performed using structured and unstructured text data from high-confidence datasets. For structured data (such as historical transaction data and macroeconomic indicators), multivariate analysis methods (such as the Pearson correlation coefficient and Granger causality test) are used to assess the correlation and statistical significance between features from different data sources, including transaction prices, transaction volumes, interest rates, and inventory levels. Time series analysis is used to further extract feature pairs with temporal leading or significant correlation, where one variable can predict the trend of another, thereby forming cross-domain numerical feature pairs. For unstructured text data (such as news and social media data), natural language processing techniques, such as sentiment analysis models, are used to represent text content through word vector computation (such as Word2Vec or BERT). Sentiment scores are then extracted from the text using sentiment dictionaries or classification models to generate a sequence of sentiment change trends. Finally, the cross-domain numerical feature pairs are mapped and integrated with the sentiment trend change rates of the text data to form a cross-feature candidate set.
[0058] For example, a study of commodity trading prices and macroeconomic indicator data over a specific time period found a Pearson correlation coefficient of 0.85 between trading prices and macroeconomic interest rates. Granger causality tests confirmed that interest rate changes significantly influence trading price trends. Furthermore, sentiment scores extracted from news and social media data showed an upward trend in sentiment over the same time window, positively correlated with price fluctuations. Text features generated using Word2Vec can be mapped to numerical relationships, integrating trading prices, interest rates, and sentiment change rates into a candidate set of cross-features.
[0059] The highly significant cross-items in the cross-feature candidate set are modeled with the historical volatility characteristics of commodities to obtain a feature interaction matrix. Principal component analysis and information gain evaluation are performed on the feature interaction matrix to screen out the main cross-features and joint features that are strongly correlated with the volatility characteristics to construct a cross-feature vector set. Among them, highly significant cross-items refer to feature combinations that are significantly correlated in cross-domain data sources.
[0060] When modeling the interaction between highly significant cross-items in a cross-feature candidate set and historical commodity volatility characteristics, key feature dimensions are defined based on the source of the cross-feature candidate set (e.g., sentiment-value composite features, historical volatility data). A feature relationship matrix is constructed using multivariate statistical models (e.g., covariance matrix or stepwise regression) to represent the contribution of each cross-item term to the overall feature space. Furthermore, frequent pattern mining algorithms (e.g., FP-Growth) are combined to identify feature combinations with high support and confidence across different data sources. Principal component analysis (PCA) is used to reduce the feature interaction matrix and identify key cross-items. Information gain is then used to assess the importance of each feature to the target task, thereby retaining joint features that are strongly correlated with volatility characteristics. Finally, a set of cross-feature vectors is constructed.
[0061] For example, the covariance matrix between trading volume and price shows a high linear correlation. Text sentiment analysis detected significant support for sentiment fluctuations and high-frequency price changes (FP-Growth confidence value reached 0.92). After PCA dimensionality reduction, key cross-features such as sentiment volatility and trading volume change rate were ultimately selected. The joint features that contributed most to the prediction task were retained to form a cross-feature vector set.
[0062] The cross-feature vector set is dynamically window clustered and hierarchically encoded in chronological order to obtain feature evolution vectors and feature structure labels across time scales. The feature evolution vectors and feature structure labels are used to construct a multi-level feature fusion framework.
[0063] The cross-feature vector set is input into a dynamic window clustering algorithm (such as DBSCAN or K-Means) to dynamically cluster the data features according to the time series dimension. For each clustering result, a cross-timescale feature evolution vector is constructed based on the temporal trend of the features, representing the behavioral characteristics of the features at different time periods. Furthermore, feature structure labels are generated using hierarchical encoding techniques (such as classification trees or hierarchical clustering). These labels reflect the classification hierarchy of the features. Ultimately, the feature evolution vectors and feature structure labels are integrated into a multi-level feature fusion framework, effectively organizing the different data features within the time dimension and facilitating subsequent processing.
[0064] For example, we used the K-Means clustering algorithm to perform cluster analysis on cross-feature vectors by weekly time windows, extracting patterns in feature changes within each week and generating a weekly timescale evolution vector. We also used a classification tree to generate a hierarchical encoding of context. For example, strong correlations between price changes and trading volume were assigned to the same structural label, which was then associated with sentiment fluctuations. Ultimately, we constructed a feature fusion framework organized by timescale and classification relevance.
[0065] The multi-level feature fusion framework is combined with the industry fluctuation cycle to perform feature grouping to obtain a hierarchical fusion feature set. A comprehensive feature vector is generated based on the hierarchical fusion feature set. The industry fluctuation cycle refers to the time parameter determined according to the seasonal and cyclical laws within the commodity industry.
[0066] Leveraging the established multi-level feature fusion framework, features are grouped according to industry fluctuation cycles. First, industry fluctuation cycles are analyzed, and seasonal analysis methods (such as time series decomposition) are used to calculate the cyclical patterns of each commodity. Based on the demand cycle and market supply patterns of commodities, the time parameter range is determined. Second, cross-timescale features and industry cycle features are mapped and integrated, and the feature set is fused hierarchically to form a hierarchical fused feature set. Finally, the hierarchical fused feature set is used as input to generate a comprehensive feature vector for subsequent risk assessment.
[0067] For example, using a metal commodity as an example, we use time series decomposition to derive its quarterly demand cycle and annual supply cycle. By combining the hierarchical feature labels within each quarter (such as sentiment volatility and trading volume change rate), we group the intra-quarter features by cycle and form a fused feature set. Ultimately, by integrating relevant features such as intra-quarter demand changes, we generate a comprehensive feature vector that fully captures market volatility.
[0068] The steps of extracting cross-domain numerical feature pairs between structured data and sentiment trend change rates of unstructured text data based on a high-confidence data set, and generating a cross-feature candidate set based on the cross-domain numerical feature pairs and sentiment trend change rates, specifically include: Multivariate Pearson correlation coefficient analysis and Granger causality test are performed on the structured data in the high-credibility data set to extract cross-domain numerical feature pairs with time series leading and high statistical significance. Among them, the structured data includes historical transaction data and macroeconomic indicator data.
[0069] Multivariate statistical analysis is performed on structured data (including historical transaction data and macroeconomic indicator data) from a high-confidence dataset. First, the prices and trading volumes in the historical transaction data are normalized by standard deviation and mean to eliminate unit scale differences. Then, Pearson correlation coefficient analysis is used to calculate the linear correlation between prices and macroeconomic indicators (such as interest rates and inventory levels) to determine whether there is a significant correlation between prices and indicators. Next, Granger causality tests are used to assess causal relationships between variables to identify the temporal influence of macroeconomic indicators on transaction data trends. Finally, cross-domain numerical feature pairs with temporal leading and statistical significance are selected for use in constructing a candidate cross-feature set.
[0070] For example, an analysis of historical transaction data (including daily transaction prices and volumes) and macroeconomic indicator data (including interest rates and inventories) yielded a Pearson correlation coefficient of 0.78 between interest rates and transaction prices, indicating a strong linear correlation between the two. A Granger causality test revealed that interest rate changes significantly led price changes in the time series (p-value < 0.05), indicating that interest rates are a time-leading variable in transaction data changes. Ultimately, based on these results, the interest rate-price feature pair was selected as part of the cross-domain numerical feature pair.
[0071] Sentiment analysis is performed on unstructured text data in a high-credibility data set, the sentiment score of each text fragment is calculated, and the sentiment trend change rate is generated based on the sentiment score of each fragment. The unstructured text data includes news information data and social media text data.
[0072] Sentiment analysis is performed on unstructured textual information in news and social media data. First, the text data is preprocessed, including word segmentation, stop word removal, and grammatical normalization. Then, using a sentiment analysis model (such as a rule-based model based on a sentiment lexicon or a deep learning-based LSTM sentiment classification model), a sentiment score is extracted for each text item. The sentiment score reflects the intensity of positive and negative emotions within the text. Next, based on the constructed sentiment time series, the sentiment trend change rate is calculated by performing a differential operation on the continuous sentiment scores to reflect the changing trends in news or social media sentiment over time.
[0073] For example, using Google's BERT model for sentiment classification on social media data (e.g., tweets about commodities), we detected that the tweet "Oil reserves are running out" had a high negative sentiment score (-0.75). By combining the time tags and performing a differential calculation on the sentiment scores of several consecutive tweets, we found a negative sentiment trend change rate (indicating an increasing negative trend). The same sentiment analysis method was applied to the news headline "Demand plummets" to generate a news sentiment score and calculate the sentiment trend change rate. The results were then used to construct a candidate set of cross-features.
[0074] The emotion trend change rate and cross-domain numerical feature pairs are time-series aligned and weightedly integrated to obtain a fused emotion and numerical composite feature set. The emotion and numerical composite feature set is input into a preset rolling window regression model to identify significant correlations between emotion fluctuations and numerical fluctuations, and obtain key cross-terms in the composite feature set.
[0075] Time series alignment is performed on sentiment trend change rates and cross-domain numerical feature pairs to ensure that both are analyzed within the same time window. Once aligned, feature weighting methods (such as those based on entropy weighting or the Delphi method) are used to assign weights to each feature to reflect the influence of sentiment and numerical features within different time windows. The weighted sentiment and numerical feature data are then input into a rolling window regression model (with an adjustable window length, such as 7 or 30 days) to assess correlations and changing trends between features. The model outputs a set of significant factors, and statistical screening is used to retain key cross-terms (e.g., the significant contribution of key feature combinations to risk).
[0076] For example, we analyzed the relationship between transaction price and sentiment change using a 7-day sliding window model. The inputs were transaction price and sentiment score change rates. The results showed a significant positive correlation within certain time windows. The weights assigned using the entropy weighting method indicated a weight of 0.65 for sentiment features and 0.35 for numerical features. Ultimately, we selected sentiment volatility-price volatility and inventory change rate-negative sentiment score change rates as key cross-features.
[0077] Information gain analysis is performed on key cross-items to screen out the most explanatory cross-feature items after combining sentiment and numerical features, and a cross-feature candidate set is generated based on the cross-feature items.
[0078] Information gain analysis is performed on the key cross-terms output by the rolling window regression model to measure the contribution of feature combinations to the ultimate goal (such as risk assessment accuracy). First, using a benchmark evaluation model (such as a decision tree or random forest), feature combinations are tested step by step to calculate the split information gain value for each feature combination. The split information gain value is calculated as follows: in, Information gain is the amount of uncertainty or information confusion a feature reduces when dividing data. The higher the information gain, the greater the contribution of the feature to classification or prediction.
[0079] is the entropy (information confusion) of the parent node, that is, the overall uncertainty of the data set before using a certain feature. The higher the entropy value, the greater the uncertainty in the data.
[0080] is the weighted sum of the entropy of the child nodes. For each partition of the parent node, calculate the entropy value of the child node and the probability of the child node appearing Then, the entropy value of each child node is multiplied by its probability of occurrence and summed to obtain the weighted entropy of the child node.
[0081] For child nodes The probability of occurrence indicates the probability of data being divided into this child node. It can be calculated by the ratio of the number of samples in the child node to the total number of samples in the parent node.
[0082] For child nodes The entropy value represents the uncertainty or information confusion of the data in the child node. It is obtained by calculating the category distribution in the child node. If the samples in a node all belong to the same category, the entropy value is 0; if the samples are evenly distributed, the entropy value is the maximum.
[0083] Then, the feature items with the highest information gain value are selected as part of the cross feature candidate set.
[0084] For example, when calculating the information gain at the splitting node of the random forest model for the feature combination "price volatility - sentiment volatility" and "inventory change rate - negative sentiment score change rate," the former had an information gain of 0.92, while the latter had an information gain of 0.74. Therefore, the feature with the higher information gain value was retained to generate the important factors in the final cross-feature candidate set.
[0085] In step S300, different structured input sequences are generated based on the comprehensive feature vector, all structured input sequences are input into the corresponding preset risk assessment sub-model to obtain a local risk response matrix and a correlation factor sensitivity matrix, and the local risk response matrix and the correlation factor sensitivity matrix are cross-validated to generate a comprehensive risk assessment result. The steps specifically include: The comprehensive feature vector is time-sliced and context-enhanced according to the preset model input requirements to obtain different structured input sequences. All structured input sequences are input into the preset corresponding risk assessment sub-model to obtain a number of risk response data and a number of correlation factor data.
[0086] The comprehensive feature vector is processed according to the input requirements of the risk assessment sub-model (e.g., support vector machines, LSTM networks, random forests, and other machine learning models). First, a sliding window algorithm is used to time-slice the comprehensive feature vector, segmenting the continuous feature data into fixed-length time windows (e.g., 7 or 30 days) to capture short-term or long-term feature trends. Furthermore, additional contextual information, such as industry cycles, macroeconomic event labels, or seasonal fluctuations, is incorporated into each time slice, using contextual enhancement techniques to generate a more explanatory input sequence. Furthermore, to ensure data standardization and consistency, each slice is normalized (e.g., using Z-score normalization) to remove scale differences. Finally, the processed time slices are input into the corresponding risk assessment sub-model for training or testing, outputting risk response data (representing the risk assessment results) and correlation factor data (representing the key factors the model focuses on).
[0087] For example, the trading volume change rate, sentiment trend change rate, and price volatility in the comprehensive feature vector are divided into 7-day time windows at daily intervals to generate a set of sliding time slices. Subsequently, macroeconomic event labels (such as "interest rate hike decision" or "inventory drop") are introduced into each time slice to form a context-enhanced dataset. Finally, the context-enhanced time slices are input into a pre-set random forest model for risk assessment. The model outputs a daily risk score (risk response data) and key factor contribution values (correlated factor data), such as inventory levels contributing 30% to the risk score.
[0088] A local risk response matrix is constructed using all risk response data, and a horizontal comparison analysis is performed on the local risk response matrix to identify the key factors that cause evaluation differences between models and obtain factor inconsistency indicators.
[0089] Based on the output risk response data of all risk assessment sub-models, a local risk response matrix is constructed. Each row of the matrix corresponds to the risk score result of a sub-model, and each column corresponds to the risk assessment value of a different time segment. Subsequently, the evaluation results of different sub-models in the matrix are compared horizontally, and the differences between the models in the risk assessment are identified through statistical indicators (such as absolute error, mean square error or Pearson correlation coefficient). Furthermore, combined with the results of the difference analysis, the sensitivity of the risk factors in each model is quantified for the features with significant differences in contribution of each model, and the factor inconsistency index is obtained (for example, the weight of a certain feature in sub-model A is higher than that in sub-model B, indicating that there is a difference in its impact on the risk score).
[0090] For example, suppose that within a constructed local risk response matrix, the risk scores from the random forest model and the LSTM network differ by a mean square error of 12% for a certain time window (e.g., January 1, 2023, to January 7, 2023). A variance analysis reveals that price volatility contributes 50% to the random forest model during this period, while it contributes only 28% to the LSTM network. This indicates that price volatility is an indicator of high factor inconsistency.
[0091] All the correlation factor data are used to construct the correlation factor sensitivity matrix, and the correlation factor sensitivity matrix is weighted averaged and normalized to obtain the factor influence distribution map.
[0092] Based on the output correlation factor data of all submodels, a correlation factor sensitivity matrix is constructed. Each row in the matrix corresponds to a submodel, and each column represents the sensitivity weight of different correlation factors to the model output. The sensitivity weights of each factor in the matrix are weighted averaged to ensure that high-confidence submodels have higher weights, thereby weakening the influence of low-confidence submodels. Next, the weighted results are normalized (for example, using Min-Max normalization) to ensure that the sensitivity weights range between 0 and 1. Finally, a factor impact distribution map is generated, which intuitively shows the impact strength and distribution characteristics of each factor on the overall risk assessment.
[0093] For example, calculating the sensitivity matrix for factors such as trading volume change rate, inventory level, and sentiment volatility yielded the following results: the sensitivity of trading volume change rate was 0.6 in the LSTM model and 0.4 in the support vector machine; the sensitivity of inventory level was 0.5 in the LSTM model and 0.7 in the random forest model. A factor influence distribution map generated through weighted averaging and normalization showed that the impact of inventory level on the comprehensive risk assessment was 0.64, while the impact of sentiment volatility was only 0.48.
[0094] The factor inconsistency index is cross-validated with the factor influence distribution map to obtain a set of credible risk signals and a set of model confidence deviations, and a comprehensive risk assessment result is generated based on the set of credible risk signals and the set of model confidence deviations.
[0095] By combining factor inconsistency indicators and factor impact distribution maps, we cross-validate the contribution and assessment consistency of each factor to screen for a set of credible risk signals and a set of model confidence deviations. During the cross-validation process, we use a logistic regression model or Bayesian model to quantify the confidence of each factor. When a factor performs consistently across multiple sub-models and exhibits high sensitivity, it is added to the credible risk signal set; otherwise, it is marked as a factor in the confidence deviation set. Ultimately, the credible risk signal set is converted into the model's final risk score, and the confidence deviation results are used to correct or adjust certain anomaly scores, ultimately outputting the final comprehensive risk assessment result.
[0096] For example, sentiment volatility and inventory levels both appear to be highly consistent sensitive factors in horizontal comparisons (confidence levels greater than 90%) and are therefore included in the credible risk signal set. However, trading volume change exhibits sensitivity bias in some models (confidence level is only 65%) and is therefore labeled a model confidence bias factor. By combining credible risk signals and confidence bias information, the final average risk score is 8.2, indicating a medium-risk market volatility forecast for the target commodity.
[0097] The steps of time slicing and context enhancing the comprehensive feature vector according to the preset model input requirements to obtain different structured input sequences, inputting all structured input sequences into the corresponding preset risk assessment sub-models, and obtaining a plurality of risk response data and a plurality of correlation factor data specifically include: The comprehensive feature vector is divided into multiple time windows through the sliding window algorithm to obtain several time slice sequences. In each time slice sequence, the context information of external macroeconomic events and industry cycles is combined for context enhancement to obtain the enhanced time slice sequence.
[0098] First, a sliding window algorithm is applied to partition the comprehensive feature vector into multiple time windows. This algorithm selects a fixed-length time window (such as 7 or 30 days) and slides it across the dataset, gradually generating a sequence of different time slices. For each time slice, in combination with the feature data within the current window, contextual enhancement is performed to include external macroeconomic events (such as policy changes, monetary policy adjustments, and sudden natural disasters) and industry cycles (such as seasonal changes in commodity demand and supply chain fluctuations). Using data fusion techniques (such as time series feature embedding), relevant macroeconomic events and industry cycle information are encoded as contextual information and added to each time slice, thereby enhancing the expressiveness of the features.
[0099] For example, suppose the comprehensive feature vector includes data such as commodity prices, trading volumes, and inventory levels, and the window size is set to 7 days. Each time the time slice is slid, it will include data from the current 7 days. If, within a certain time window, a macroeconomic event is labeled "interest rate hike," and industry cycle information is labeled "peak season demand," this contextual information will be encoded into the feature sequence for that time window along with the current price fluctuation data, thereby enhancing the timeliness of the feature sequence and its ability to reflect market changes.
[0100] All enhanced time slice sequences are standardized, all standardized time slices are input into the corresponding preset risk assessment sub-model, and corresponding risk response data and correlation factor data are output.
[0101] After time slicing and context augmentation, all time slice sequences need to be normalized to ensure that different time windows and features have the same scale when input into the model. Normalization can use Z-score normalization or Min-Max normalization to adjust each feature value in each time slice to a uniform standard range (for example, Z-score normalization transforms the data into a standard normal distribution with mean 0 and variance 1). The normalized time slices are then input into the pre-set risk assessment sub-model for training or prediction. The sub-model will output the corresponding risk response data and correlation factor data.
[0102] Among them, the risk assessment sub-model uses the following models: Random Forest Model: A random forest is an ensemble learning model based on decision trees. It constructs multiple decision trees and combines their results to perform classification or regression. In commodity risk assessment, the random forest model is often used to process high-dimensional, nonlinear data. It trains multiple decision trees, performs multiple sampling and partitioning steps, and ultimately determines the risk level based on the voting results of all the trees. This model has the advantage of being able to handle complex feature relationships and has high accuracy and stability.
[0103] LSTM (Long Short-Term Memory): LSTM is a deep learning model based on recurrent neural networks (RNNs) and is particularly well-suited for processing time series data. In commodity risk assessment, LSTM can predict future trends by capturing long-term and short-term dependencies in historical data. LSTM can handle nonlinear changes and complex temporal dependencies, demonstrating strong predictive power for commodity price fluctuations and sentiment trends.
[0104] Support Vector Machine (SVM): A supervised learning model commonly used for classification problems, SVMs construct a hyperplane to maximize the margin between classes. For commodity risk assessment, SVMs effectively handle nonlinear relationships between features, particularly in multi-class classification problems, enabling accurate classification of risk levels (e.g., "low risk," "medium risk," and "high risk").
[0105] XGBoost (Extreme Gradient Boosted Trees): XGBoost is an ensemble learning method based on the Gradient Boosted Trees (GBDT) algorithm. By optimizing loss functions and regularization strategies, it can more accurately model complex data relationships. For commodity risk assessment, XGBoost can capture complex interactions between multiple features, thereby improving the accuracy and robustness of predictions.
[0106] For example, suppose a time slice contains contextual information about transaction prices, trading volumes, inventory levels, and macroeconomic events (such as "interest rate hike"). During the normalization phase, these features are converted to numerical values with a mean of 0 and a standard deviation of 1. These features are then input into a pre-defined risk assessment sub-model (such as a random forest model or an LSTM model). Based on the input features, this sub-model outputs a risk score (such as "medium risk") and a correlation factor (such as the impact of inventory levels on the risk score of 0.6 and the impact of trading volume on the risk score of 0.3) for the current time window.
[0107] In this example, the LSTM model is used as an example. Assume that there is a feature slice for a 7-day time window, containing the following data: Transaction prices: [100, 101, 98, 95, 102, 100, 97]; Volume: [500,520,510,490,530,540,510] Stock levels: [200,210,215,205,220,215,210]; Macroeconomic events (context information): ['interest rate hike', 'no change', 'interest rate hike', 'policy adjustment', 'no change', 'interest rate hike', 'no change']; 1. Standardization processing: Use the Z-score method to standardize these features. Assume that the standardized data is: Trading prices: [0.35, 0.42, -0.18, -0.60, 0.62, 0.35, -0.53]; Volume: [0.56, 0.75, 0.67, -0.34, 1.03, 1.15, 0.68] Inventory levels: [0.23, 0.43, 0.62, 0.33, 0.86, 0.62, 0.43]; 2. Input to the preset LSTM model: The standardized data is input into the LSTM model. The model captures the time dependency between price and trading volume based on historical data, outputs a corresponding risk score (such as "medium risk"), and uses a regression model to conclude that the impact of inventory levels on risk is 0.4, and the impact of trading volume on risk is 0.6.
[0108] Finally, the generated output is: Risk score: "Medium risk", associated factor data: "Inventory level contribution: 0.4, transaction volume contribution: 0.6".
[0109] In step S400, a dynamic risk scoring curve for the target commodity is generated based on the comprehensive risk assessment results, a risk warning level is determined based on the dynamic risk scoring curve, and a risk warning report is generated based on the risk warning level. Specifically, the steps include: The comprehensive risk assessment results are quantified into risk indexes and combined with the corresponding factor influence distribution maps to obtain a multi-dimensional risk score subset. The multi-dimensional risk score subset is normalized according to the preset indicator categories, and the risk score sequence is reconstructed according to the time label to generate a dynamic risk score curve for the target commodity.
[0110] First, the comprehensive risk assessment results are quantified using a risk index. The assessment results include various risk dimensions (such as price volatility, trading volume changes, and sentiment fluctuations). The risk contribution of each dimension influences the final risk score, so a multi-dimensional risk score subset is extracted based on the factor impact distribution map. Subsequently, the multi-dimensional risk score subset is normalized according to pre-set indicator categories (such as market volatility, sentiment fluctuations, and macroeconomic impact) to ensure that the risk scores of all dimensions are within the same range (e.g., between 0 and 1). Next, the risk score sequence is reconstructed based on time tags to obtain the dynamic risk score of the target commodity at different time points and construct a dynamic risk score curve.
[0111] For example, assume that the comprehensive risk assessment results include the following dimensions: Price volatility risk score: 0.75; Trading volume volatility risk score: 0.65; Mood swing risk score: 0.80; The factor impact distribution map shows that sentiment fluctuations significantly contribute to the overall risk assessment. Therefore, sentiment fluctuations are weighted relatively higher when generating the multi-dimensional risk score subset. After normalizing all dimension scores, we obtain a comprehensive risk score at each time point and construct a dynamic risk score curve based on time tags.
[0112] Based on the dynamic risk scoring curve, the score volatility and score jump factor are calculated, and the score volatility and score jump factor are input into the preset risk level threshold set for mapping to obtain the current risk warning level, and a risk warning report is generated based on the risk warning level.
[0113] After the dynamic risk score curve is generated, the score volatility and score jump factor need to be calculated. The score volatility measures the degree of change in the risk score over time and can be measured using the standard deviation or weighted coefficient of variation; the score jump factor represents the dramatic change in the risk score within a short period of time and can be determined by calculating the maximum amplitude change of the score curve. These two indicators are mapped to a set of preset risk level thresholds, which usually include levels such as "low risk", "medium risk", and "high risk". The current risk warning level is determined based on these thresholds. Finally, a detailed risk warning report is generated based on the determined risk warning level, which includes an explanation of the risk level and response recommendations.
[0114] For example, suppose the standard deviation of the dynamic risk score curve is 0.12, the volatility is 0.09, and the calculated score jump factor is 0.15 (indicating a significant score change over a short period of time). Based on the preset risk level thresholds (e.g., volatility greater than 0.1 indicates "high risk"), this situation will be mapped to a "high risk" warning level. The risk warning report will include an analysis of the current market conditions, the impact of relevant factors (such as sentiment fluctuations and trading volume changes) on risk, and proposed corresponding risk response strategies.
[0115] The steps of calculating the score volatility and the score jump factor based on the dynamic risk score curve, inputting the score volatility and the score jump factor into a preset risk level threshold set for mapping, and obtaining the current risk warning level specifically include: A time series analysis is performed on the dynamic risk scoring curve to calculate the scoring volatility. The score stability of the dynamic risk scoring curve is measured using the weighted coefficient of variation method to obtain the scoring curve. The score jump amplitude and frequency in the scoring curve are calculated using the differential algorithm to obtain the score jump factor.
[0116] Performing time series analysis on the dynamic risk score curve can well capture its fluctuation characteristics over time. First, the volatility of the dynamic score curve is calculated using the historical score values (such as the daily score series) in the dynamic score curve. , the formula is as follows: in, Dynamic risk scoring curve in time The rating value of the point. is the average value of the scoring curve. is the number of ratings within the time window.
[0117] In order to measure stability more accurately, the weighted coefficient of variation method is used to evaluate the fluctuation stability of the score. The formula of the weighted coefficient of variation is: Weighted coefficient of variation in, It is the weighted standard deviation, which is evaluated by calculating different feature weights. is the mean of the scoring curve.
[0118] At the same time, the difference algorithm is used to differentiate the score curve value point by point to calculate the score jump amplitude (indicating the maximum change in the score between adjacent time points) and the jump frequency (indicating the number of times the score changes significantly within the time range). The difference algorithm formula is as follows: Combining the volatility, jump amplitude and jump frequency, the changing characteristics of the score curve are comprehensively evaluated, and the score volatility and score jump factor are obtained.
[0119] The score volatility and score jump factor are input into the preset risk level threshold set to map the score of the current dynamic risk score curve to obtain the current risk warning level.
[0120] Compare the calculated score volatility and score jump factor with a set of pre-set risk level thresholds. The pre-set risk level threshold set typically includes multiple levels (e.g., "low risk," "medium risk," and "high risk"), and map the scores based on the volatility and jump factor. The mapping rule can use logistic regression or interval mapping, for example: Low risk: volatility < 0.05 and jump amplitude < 0.10; Medium risk: 0.05 ≤ Volatility < 0.15 or 0.10 ≤ Jump Amplitude < 0.20; High risk: volatility ≥ 0.15 or jump amplitude ≥ 0.20; Based on the mapping results, the risk level corresponding to the current dynamic score is determined. Finally, a detailed risk warning report is generated based on the mapping results. The report will include the current risk level, analysis of the changing trends of risk characteristic factors, and targeted response recommendations.
[0121] In this embodiment, an innovative method for commodity risk assessment is proposed through comprehensive analysis and processing of high-credibility data sets. First, in response to the complexity of multi-source data, an association consistency detection method is adopted. Through time alignment, text cleaning, entity recognition, and cointegration analysis, the deviations between different data sources are eliminated, and a high-credibility data set is constructed to ensure the accuracy and reliability of the analysis basis. Secondly, by extracting cross-domain numerical feature pairs of structured data and sentiment trend change rates of unstructured text data from the high-credibility data set, combined with principal component analysis and information gain evaluation technology, a cross-feature vector set including sentiment and volatility features is effectively generated. In addition, based on the time window dynamic clustering and hierarchical coding method, a multi-level feature fusion framework is constructed in combination with industry cycles, which further improves the adaptability and comprehensiveness of feature modeling. During the risk assessment phase, a sliding window algorithm and context enhancement techniques are used to process the comprehensive feature vector into a structured input sequence, which is then fed into various risk assessment sub-models (such as random forest models and LSTM models). This constructs a local risk response matrix and a relationship factor sensitivity matrix. Through horizontal comparison and normalization, a factor influence distribution map is generated, providing strong support for the comprehensive risk results. Finally, based on the dynamic risk scoring curve, the score volatility is calculated using the weighted coefficient of variation method, and the score jump factor is calculated using a differential algorithm to assess the stability and jump characteristics of the score. This factor is then mapped to a set of preset thresholds to determine the risk warning level and generate a risk warning report. In summary, this embodiment integrates multi-source data analysis, multi-dimensional feature modeling, and dynamic risk assessment methods to achieve accurate quantitative assessment and real-time warning of commodity market risks, comprehensively enhancing the scientific and practical nature of risk management.
[0122] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not intended to limit the conditions under which the present invention can be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportional relationships, or adjustments in size should still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and objectives that can be achieved by the present invention.
[0123] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A commodity risk early warning method, characterized in that: include: Acquire multi-source data of target commodities, perform correlation consistency detection on the multi-source data, and obtain a high-reliability data set; Extracting a cross-feature candidate set from the high-credibility data set, performing feature interaction modeling on the cross-feature candidate set and the historical volatility characteristics of commodities to obtain a feature interaction matrix, performing principal component analysis and information gain evaluation on the feature interaction matrix to obtain a comprehensive feature vector; Generating different structured input sequences based on the comprehensive feature vector, inputting all the structured input sequences into a preset corresponding risk assessment sub-model to obtain a local risk response matrix and a correlation factor sensitivity matrix, and cross-validating the local risk response matrix and the correlation factor sensitivity matrix to generate a comprehensive risk assessment result; A dynamic risk scoring curve for the target commodity is generated based on the comprehensive risk assessment result, a risk warning level is determined according to the dynamic risk scoring curve, and a risk warning report is generated based on the risk warning level.
2. The commodity risk early warning method according to claim 1, characterized in that: The step of obtaining multi-source data of the target commodity, performing correlation consistency detection on the multi-source data, and obtaining a high-credibility data set includes: Obtain multi-source data on the target commodity within a preset time window, wherein the multi-source data includes historical transaction data, macroeconomic indicator data, news information data, and social media text data; Performing time alignment processing on the historical transaction data and the macroeconomic indicator data to obtain structured time series data; Performing text cleaning and entity recognition processing on the news information data and the social media text data to obtain structured text feature data; Performing association consistency detection on the structured time series data and the structured text feature data to obtain a time consistency feature and a semantic consistency feature; Using the temporal consistency feature and the semantic consistency feature, the anomalies in the historical transaction data, the macroeconomic indicator data, the news information data, and the social media text data are corrected and removed to construct a corrected multi-source data set; Perform statistical stability testing and cointegration analysis on the modified multi-source data set to obtain a correlation matrix and confidence weights, screen out data source samples based on the correlation matrix and the confidence weights, and construct a high-credibility data set based on the data source samples.
3. The commodity risk early warning method according to claim 1, characterized in that: extracting a cross-feature candidate set from the high-credibility data set, and performing feature interaction modeling on the cross-feature candidate set and the historical fluctuation characteristics of commodities; The steps of obtaining a feature interaction matrix, performing principal component analysis and information gain evaluation on the feature interaction matrix, and obtaining a comprehensive feature vector include: Extracting cross-domain numerical feature pairs between structured data and sentiment trend change rates of unstructured text data based on the high-confidence data set, and generating a cross-feature candidate set based on the cross-domain numerical feature pairs and the sentiment trend change rates; Performing feature interaction modeling on the highly significant cross-items in the cross-feature candidate set and the historical volatility characteristics of commodities to obtain a feature interaction matrix, and performing principal component analysis and information gain evaluation on the feature interaction matrix to construct a cross-feature vector set, wherein the highly significant cross-items refer to significantly correlated feature combinations in cross-domain data sources; Performing dynamic window clustering and hierarchical encoding on the cross feature vector set in chronological order to obtain feature evolution vectors and feature structure labels, and constructing a multi-level feature fusion framework using the feature evolution vectors and the feature structure labels; The multi-level feature fusion framework is combined with the industry fluctuation cycle to perform feature grouping to obtain a hierarchical fusion feature set, and a comprehensive feature vector is generated based on the hierarchical fusion feature set, wherein the industry fluctuation cycle refers to a time parameter determined according to the seasonal and cyclical laws within the commodity industry.
4. The commodity risk early warning method according to claim 3, characterized in that: The step of extracting cross-domain numerical feature pairs between structured data and sentiment trend change rates of unstructured text data based on the high-confidence data set, and generating a cross-feature candidate set based on the cross-domain numerical feature pairs and the sentiment trend change rates, comprises: Performing multivariate Pearson correlation coefficient analysis and Granger causality test on the structured data in the high-credibility data set to extract cross-domain numerical feature pairs, wherein the structured data includes historical transaction data and macroeconomic indicator data; Performing sentiment analysis on the unstructured text data in the high-credibility data set, calculating a sentiment score for each text segment, and generating a sentiment trend change rate based on the sentiment score of each segment, wherein the unstructured text data is news information data and social media text data; Performing time series alignment and weighted integration on the emotion trend change rate and the cross-domain numerical feature pair to obtain a fused emotion and numerical composite feature set, inputting the emotion and numerical composite feature set into a preset rolling window regression model to obtain key cross-terms; An information gain analysis is performed on the key cross-items to screen out cross-feature items, and a cross-feature candidate set is generated based on the cross-feature items.
5. The commodity risk early warning method according to claim 1, characterized in that: The steps of generating different structured input sequences based on the comprehensive feature vector, inputting all the structured input sequences into a preset corresponding risk assessment sub-model to obtain a local risk response matrix and a correlation factor sensitivity matrix, and cross-validating the local risk response matrix and the correlation factor sensitivity matrix to generate a comprehensive risk assessment result include: The comprehensive feature vector is time-sliced and context-enhanced according to preset model input requirements to obtain different structured input sequences, and all the structured input sequences are input into corresponding preset risk assessment sub-models to obtain a plurality of risk response data and a plurality of correlation factor data; Constructing a local risk response matrix using all of the risk response data, performing a horizontal comparison analysis on the local risk response matrix, and obtaining a factor inconsistency index; Utilizing all correlation factor data to construct a correlation factor sensitivity matrix, performing weighted averaging and normalization on the correlation factor sensitivity matrix to obtain a factor impact distribution map; The factor inconsistency index is cross-validated with the factor impact distribution map to obtain a credible risk signal set and a model confidence deviation set, and a comprehensive risk assessment result is generated based on the credible risk signal set and the model confidence deviation set.
6. The commodity risk early warning method according to claim 5, characterized in that: The steps of time slicing and context enhancing the comprehensive feature vector according to preset model input requirements to obtain different structured input sequences, and inputting all the structured input sequences into corresponding preset risk assessment sub-models to obtain a plurality of risk response data and a plurality of correlation factor data include: The comprehensive feature vector is divided into multiple time windows by a sliding window algorithm to obtain a number of time slice sequences. In each of the time slice sequences, context enhancement is performed by combining the context information of external macroeconomic events and industry cycles to obtain an enhanced time slice sequence. All the enhanced time slice sequences are standardized, all the standardized time slices are input into the corresponding preset risk assessment sub-model, and corresponding risk response data and correlation factor data are output.
7. The commodity risk early warning method according to claim 1, characterized in that: The steps of generating a dynamic risk scoring curve for the target commodity based on the comprehensive risk assessment result, determining a risk warning level according to the dynamic risk scoring curve, and generating a risk warning report based on the risk warning level include: Performing risk index quantification processing on the comprehensive risk assessment results to obtain a multi-dimensional risk score subset, and normalizing the multi-dimensional risk score subset according to preset indicator categories to generate a dynamic risk score curve for the target commodity; Based on the dynamic risk scoring curve, the scoring volatility and the score jump factor are calculated, and the scoring volatility and the score jump factor are input into a preset risk level threshold set for mapping to obtain the current risk warning level, and a risk warning report is generated according to the risk warning level.
8. The commodity risk early warning method according to claim 7, characterized in that: The step of calculating the score volatility and the score jump factor based on the dynamic risk score curve, inputting the score volatility and the score jump factor into a preset risk level threshold set for mapping, and obtaining the current risk warning level includes: Performing a time series analysis on the dynamic risk scoring curve to calculate the score volatility, measuring the score stability of the dynamic risk scoring curve using a weighted coefficient of variation method to obtain a scoring curve, and calculating the score jump amplitude and frequency in the scoring curve using a differential algorithm to obtain a score jump factor; The score volatility and the score jump factor are input into a preset risk level threshold set to map the score of the current dynamic risk score curve to obtain the current risk warning level.
Citation Information
Patent Citations
Multi-source data processing method and system applied to supplier evaluation
CN118967202A
Method and device for formulating power transaction strategy
CN119671622A
Financial transaction anomaly detection and risk assessment method and device based on artificial intelligence
CN119693111A
Platform knowledge graph construction method and system based on bulk commodity transaction information
CN119808930A
Data risk management system and method based on large model
CN119918065A