A Lithium Carbonate Price Prediction Method Based on Multimodal Data Collaboration

By using a multimodal data collaborative prediction method, integrating multiple data types and constructing a dual-branch collaborative prediction model, the problems of insufficient data utilization and poor model adaptability in lithium carbonate price prediction are solved, achieving high-precision, dynamically adaptable, and interpretable prediction results.

CN122089357APending Publication Date: 2026-05-26SHANXI DONGTUO NEW ENERGY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610153927.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-03
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for predicting lithium carbonate prices suffer from insufficient data utilization, single modalities and poor synergy, lack of specificity and generalization ability in the prediction models, large data noise interference, lack of dynamic adaptive adjustment mechanisms, poor interpretability of prediction results, and limited practicality.

Method used

A multimodal data collaborative prediction method is adopted, which integrates structured, semi-structured and unstructured data. Through personalized preprocessing, feature extraction and weighted collaborative fusion, a two-branch collaborative prediction model is constructed. Combined with collaborative attention mechanism and dynamic correction mechanism, personalized noise filtering and outlier correction strategy is designed, and interpretability analysis and visualization methods are adopted.

Benefits of technology

It achieves high-precision and stable lithium carbonate price forecasting, taking into account both long-term trends and short-term fluctuations, improving forecast accuracy and anti-interference capabilities, and possessing dynamic adaptability and strong interpretability to meet the needs of enterprise business decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089357A_ABST
    Figure CN122089357A_ABST
Patent Text Reader

Abstract

This invention provides a lithium carbonate price prediction method based on multimodal data collaborative driving, comprising the following steps: multimodal data source identification and data acquisition; multimodal data classification and standardization; personalized noise filtering and outlier correction of multimodal data; extraction of trend and correlation features from structured data; extraction of semantic and visual / speech features from unstructured data; extraction of key information and feature quantification from semi-structured data; weighted collaborative fusion and feature optimization of multimodal features; construction and training of a two-branch collaborative prediction model; model prediction and preliminary verification of prediction results; optimization and dynamic correction of prediction results; interpretability analysis and visualization of prediction results; and method effectiveness verification and continuous improvement. This method, based on multimodal data collaborative driving, addresses the pain points of existing lithium carbonate price prediction technologies, forming a comprehensive, high-precision, and highly practical prediction solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention specifically relates to a lithium carbonate price prediction method based on multimodal data collaborative driving. Background Technology

[0002] As a core raw material for lithium-ion batteries, pharmaceuticals, ceramics, and other fields, the price fluctuations of lithium carbonate directly impact the development and economic benefits of upstream and downstream industries such as new energy vehicles, energy storage, and consumer electronics. In recent years, with the advancement of global "dual-carbon" goals, the penetration rate of new energy vehicles has continued to rise, and the energy storage industry has expanded rapidly, leading to explosive growth in the demand for lithium carbonate. Simultaneously, influenced by a combination of factors including mining difficulty, purification technology, supply chain efficiency, policy guidance, and speculation in the capital market, lithium carbonate prices have experienced significant volatility and randomness. For example, from 2021 to 2022, the price of lithium carbonate surged from less than 100,000 yuan / ton to over 500,000 yuan / ton, then quickly fell back to around 150,000 yuan / ton in 2023, before showing a fluctuating upward trend in 2024. This dramatic fluctuation has posed significant challenges to the production planning, inventory optimization management, and precise cost control of upstream and downstream enterprises.

[0003] Currently, lithium carbonate price forecasting has become a research hotspot in the industry. Existing forecasting methods are mainly divided into three categories: First, traditional statistical methods, such as time series analysis and regression analysis, which rely on historical price data for trend fitting and have limited adaptability; second, machine learning methods, such as support vector machines, random forests (RF), and neural networks (CNN, LSTM), which gradually introduce supply and demand data and macroeconomic data to assist in forecasting, but still suffer from the problem of single data dimension; and third, simple multi-source data fusion methods, which simply splice together a small amount of different types of data and input them into the forecasting model, lacking in-depth mining and synergistic utilization of the characteristics of multimodal data, and failing to fully release the value of the data.

[0004] Currently, the industry's requirements for the accuracy of lithium carbonate price forecasts are continuously increasing. Enterprises urgently need forecasting methods that can adapt to complex market environments, balance short-term fluctuations and long-term trends, and possess strong generalization capabilities to assist in business decision-making and reduce operational risks. Meanwhile, with the iterative development of big data and artificial intelligence technologies, the cost of acquiring multimodal data (such as structured supply and demand data, semi-structured industry reports, and unstructured news text and image data) has significantly decreased, providing a solid data foundation for building more accurate and practical forecasting methods.

[0005] Prior to this method, the field of lithium carbonate price forecasting faced numerous technical bottlenecks, resulting in existing forecasting methods failing to meet the actual needs of the industry in terms of accuracy, generalization ability, and practicality. Specifically: First, data utilization is insufficient, modalities are singular, and coordination is poor. Most existing methods rely solely on single-modal data (such as historical price data or simple supply and demand data), ignoring the core characteristic that lithium carbonate prices are influenced by multiple dimensions and types of factors. Unstructured / semi-structured data, such as policy texts (e.g., new energy subsidy policies, mineral mining control policies), news and public opinion (e.g., mining safety accidents, supply chain disruptions), and image data (e.g., lithium mine sites, battery production line load), can all directly reflect changes in the supply and demand of the lithium carbonate market. However, existing methods either fail to incorporate this type of data or simply quantify it before directly inputting it into the model, without deeply exploring the inherent relationships between different modalities, thus failing to fully realize the value of the data.

[0006] Second, the predictive models lack specificity and have weak generalization ability. Lithium carbonate price fluctuations are influenced by both short-term unforeseen factors (such as public opinion events and logistical disruptions) and long-term trend factors (such as industrial policies, technological advancements, and changes in supply and demand). Most existing models employ single-structure predictive networks, which cannot simultaneously adapt to the randomness of short-term fluctuations and the regularity of long-term trends. This results in significant differences in predictive accuracy and insufficient generalization ability across different market environments. For example, while LSTM-based models can capture long-term dependencies in time series, their response speed to short-term unforeseen factors is slow; while CNN-based models can extract local features from data, they struggle to effectively capture long-term price trends.

[0007] Third, the data suffers from significant noise interference, resulting in suboptimal preprocessing. Multimodal data related to lithium carbonate originates from complex sources, including price data from different platforms, news texts from various media outlets, and statistical reports from different institutions, containing a large amount of noisy data (such as false price information, duplicate news, and statistical errors). Existing preprocessing methods mostly employ simple filtering and deletion techniques, failing to design personalized preprocessing strategies tailored to the noise characteristics of different modalities. Consequently, noisy data entering the prediction model severely impacts prediction accuracy.

[0008] Fourth, the lack of a dynamic adaptive adjustment mechanism makes it unable to adapt to market changes. The lithium carbonate market environment is constantly evolving; adjustments in supply and demand, policy updates, and technological breakthroughs all alter the market's operating logic, and the weights of different factors on prices change dynamically over time. However, most existing forecasting methods use fixed model parameters and feature weights, failing to dynamically adjust the model structure and parameters according to changes in the market environment. This results in a significant decrease in forecast accuracy when major changes occur in the market landscape.

[0009] Fifth, the predictive results have poor interpretability and limited practicality. Most existing deep learning-based prediction methods are "black box models," which cannot clearly explain the impact of various factors (such as lithium mine production, news sentiment, and macroeconomic indicators) on the lithium carbonate price prediction results. This makes it difficult for corporate decision-makers to understand and trust the prediction results, greatly limiting the practical application value of the methods.

[0010] In summary, this application proposes a lithium carbonate price prediction method based on multimodal data collaborative driving to solve the above problems. Summary of the Invention

[0011] The purpose of this invention is to address the shortcomings of existing technologies by providing a lithium carbonate price prediction method based on multimodal data collaborative driving, which can effectively solve the aforementioned problems.

[0012] To achieve the above requirements, the technical solution adopted by the present invention is as follows: A lithium carbonate price prediction method based on multimodal data collaborative driving is provided, which includes the following steps: S1: Steps for determining multimodal data sources and acquiring data; S2: Steps for multimodal data classification and standardization; S3: Steps for personalized noise filtering and outlier correction of multimodal data; S4: Steps for extracting trend and correlation features from structured data; S5: Steps for extracting semantic and visual / speech features from unstructured data; S6: Steps for extracting key information and quantifying features from semi-structured data; S7: Steps for performing multimodal feature weighted collaborative fusion and feature optimization; S8: Steps for constructing and training a dual-branch collaborative prediction model; S9: Steps for performing model prediction and preliminary verification of prediction results; S10: Steps for optimizing and dynamically correcting prediction results; S11: Steps for interpreting and visualizing the prediction results; S12: Steps for validating the effectiveness of the method and for continuous improvement.

[0013] This method, driven by multimodal data collaboration, addresses the pain points of existing lithium carbonate price prediction technologies, forming a comprehensive, high-precision, and highly practical prediction solution. Its core benefits are as follows: I. More comprehensive data utilization, overcoming the challenge of single-modality data. Integrating three major categories of multimodal data—structured (supply and demand, macroeconomic, price data), semi-structured (industry reports, corporate announcements), and unstructured (public opinion, image, and voice data)—covering both long-term trend factors and short-term emergencies, this approach fully leverages the intrinsic value of various data types through personalized preprocessing, feature extraction, and weighted collaborative fusion. This addresses the problems of existing methods, such as single data sources, poor coordination, and insufficient release of data value, providing comprehensive data support for accurate forecasting.

[0014] II. Higher Prediction Accuracy, Balancing Long-Term Trends and Short-Term Fluctuations. A dual-branch collaborative prediction model is constructed (the long-term trend branch is based on an improved Transformer, and the short-term fluctuation branch is based on an improved CNN-LSTM). Combined with a collaborative attention mechanism to achieve dynamic weight allocation, it can accurately capture long-term price trends (such as quarterly and annual supply and demand changes) and quickly respond to short-term fluctuations (such as public opinion and the impact of sudden accidents). With the addition of a composite loss function, hyperparameter optimization, and dynamic correction mechanism, the prediction accuracy is significantly improved. The short-term daily prediction MAPE is ≤5%, and the medium-to-long-term monthly prediction MAPE is ≤8%, which is superior to existing single-structure models.

[0015] Third, it has stronger anti-interference capabilities and more guaranteed data quality. Personalized noise filtering and outlier correction strategies are designed to address the noise characteristics of different modalities, distinguishing between reasonable fluctuations and spurious noise, and between real and spurious anomalies. This effectively eliminates false, redundant, and erroneous data, reducing noise interference. Simultaneously, data augmentation and cross-validation enhance the model's generalization ability, addressing the problems of poor preprocessing, weak model generalization, and susceptibility to noise in existing methods, ensuring the stability of prediction results.

[0016] Fourth, it exhibits strong dynamic adaptability, readily adapting to market dynamics. A dynamic correction mechanism for forecast results is established, setting different correction intervals based on the forecast scale (short-term / medium-to-long-term). Combined with online fine-tuning, transfer learning, and expert experience correction, it can adapt in real-time to changes in the market landscape (such as policy adjustments and sudden supply-demand shifts), dynamically adjusting model parameters, feature weights, and forecast results. This addresses the pain point of existing methods having fixed parameters and being unable to adapt to dynamic market changes, ensuring stable long-term forecast performance.

[0017] V. Enhanced Interpretability and Practical Value. Employing multiple interpretability methods such as SHAP value analysis and modality / branch contribution analysis, this method clarifies the degree and direction of influence of each feature, modality, and branch on the prediction results, overcoming the "black box" problem of deep learning models. Simultaneously, it designs diverse visualization reports (professional technical reports and decision-making application reports) to simplify the technical threshold, making it easier for enterprise decision-makers to understand the prediction logic and core influencing factors. This provides clear support for business decisions such as inventory adjustment, production planning, and cost control, enhancing the practical application value of the prediction method.

[0018] VI. Enhanced practicality and adaptability to actual industry needs. It enables multi-scale forecasting (1-7 day daily, 1-3 month), outputting forecast deviation range, confidence level, and application suggestions to meet the decision-making needs of enterprises in different scenarios; it optimizes forecasting efficiency, with a single batch forecasting time of ≤10 seconds, and can quickly output standardized forecasting reports; simultaneously, through effectiveness verification and continuous improvement mechanisms, it forms a closed loop of "verification-improvement-re-verification," continuously adapting to industry needs and solving the problems of limited practicality and difficulty in implementation of existing methods.

[0019] VII. The technology is technologically advanced and possesses significant value for industry promotion. Integrating cutting-edge technologies such as deep learning, natural language processing, and computer vision, and optimizing feature extraction, model structure, and fusion strategies, it exhibits significant advantages over existing ARIMA, single LSTM, and simple multi-source fusion models in terms of prediction accuracy, generalization ability, interpretability, and dynamic adaptability. It can be widely applied to price decisions for upstream and downstream enterprises in the lithium carbonate industry (lithium mining, lithium carbonate production, and battery manufacturing), and can also provide technical reference for price forecasting of other commodities, thus possessing strong value for industry promotion. Attached Figure Description

[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, use the same reference numerals to denote the same or similar parts. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic flowchart of a lithium carbonate price prediction method based on multimodal data collaborative driving according to an embodiment of this application is shown. Detailed Implementation

[0021] To make the objectives, technical solutions and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and specific embodiments.

[0022] In the following description, references to "an embodiment," "an embodiment," "an example," "example," etc., indicate that the described embodiment or example may include a particular feature, structure, characteristic, property, element, or limitation, but not every embodiment or example necessarily includes that particular feature, structure, characteristic, property, element, or limitation. Furthermore, the repeated use of the phrase "an embodiment according to this application," while possibly referring to the same embodiment, does not necessarily refer to the same embodiment.

[0023] For simplicity, certain technical features known to those skilled in the art are omitted in the following description.

[0024] Example 1: According to one embodiment of this application, a lithium carbonate price prediction method based on multimodal data collaborative driving is provided, such as... Figure 1 As shown, it includes the following steps: Step 1: Determining the Multimodal Data Source and Acquiring Data This step serves as the foundation of the entire forecasting method. Its core objective is to clarify the types and sources of multimodal data required for lithium carbonate price forecasting, establish a standardized data collection process, and ensure that the collected data is comprehensive, accurate, and effective, providing high-quality data support for all subsequent steps. The operation and working principle of this step are as follows: First, based on the core influencing factors of lithium carbonate price fluctuations, three major categories of multimodal data sources were identified, and the collection scope and standards for each type of data were clearly defined to avoid data redundancy or missing key data. The first category is structured data. This type of data has a clear numerical format and statistical standards, mainly reflecting the supply and demand fundamentals of the lithium carbonate market, the macroeconomic environment, and the operation of the upstream and downstream of the industrial chain. Specifically, it includes: historical lithium carbonate price data (daily, weekly, and monthly prices of industrial-grade and battery-grade lithium carbonate, collected over the past 10 years, with data accuracy to yuan / ton), lithium mine supply and demand data (global and major producing areas' lithium mine mining volume, inventory, import volume, and export volume, collected monthly), battery industry data (new energy vehicle production, power battery installation volume, and battery recycling volume, collected monthly), macroeconomic data (GDP, PPI, RMB exchange rate, and commodity composite index, collected quarterly), and policy quantitative data (new energy subsidy intensity, number of mining approvals, and environmental control levels, collected annually / quarterly, and converted into numerical values ​​using a quantitative scoring method). The second category is semi-structured data. This type of data has a certain structural framework but lacks a standard numerical format. It mainly includes: industry research reports (lithium carbonate market analysis reports released by securities firms and industry associations, focusing on extracting key information such as supply and demand forecasts, price trend judgments, and core influencing factor analysis); and announcements from companies in the industrial chain (capacity announcements, performance announcements, and capacity expansion / contraction information from lithium mining companies, lithium carbonate producers, and battery companies, collected quarterly / annually). The third category is unstructured data. This type of data has no fixed format and mainly reflects the impact of short-term unforeseen factors on prices. Specifically, it includes: news and public opinion data (news texts released by mainstream financial media and industry media regarding lithium mining, lithium carbonate production, supply chain logistics, policy regulation, safety accidents, etc., collected daily); image data (images of lithium mining sites, lithium carbonate production lines, and new energy vehicle production workshops, collected weekly to determine capacity utilization); and audio data (recordings of speeches on the lithium carbonate market at industry seminars and company press conferences, collected monthly to extract key viewpoints).

[0025] Secondly, a multi-source data collection channel was established to ensure the real-time nature and integrity of the data. For structured data, automatic data collection was conducted via API interfaces by connecting to industry databases (such as Baichuan Information and Longzhong Information), government statistical department websites, and corporate websites, with fixed collection times set daily, weekly, and monthly to ensure timely data updates. For data that could not be collected via API interfaces, a combination of manual entry and verification was used, establishing standardized data entry specifications to avoid errors. For semi-structured data, web crawling technology was employed to selectively crawl research reports and announcements from designated securities firms, industry associations, and corporate websites. Key information was extracted using PDF and HTML parsing technologies and converted into a processable text format. Simultaneously, a report and announcement update monitoring mechanism was established, initiating the collection process immediately upon the release of new content. For unstructured data, news and public opinion data are collected automatically by connecting to public opinion monitoring platforms (such as Baidu Public Opinion and Sina Public Opinion) and using keyword search methods (keywords include "lithium carbonate price", "lithium mining", "power battery installation", "lithium mine safety accident", "lithium carbonate supply chain", etc.), filtering out redundant content such as advertisements and irrelevant information; image data is obtained through industry cooperation channels, filtering out clear and effective images and removing blurry and irrelevant images; audio data is obtained by recording industry seminars and corporate press conferences, or by connecting to relevant audio platforms and using professional audio acquisition tools to ensure clear and noise-free audio.

[0026] Finally, a data collection and verification mechanism was established to eliminate false, erroneous, and duplicate data. For each type of data collected, specific verification rules were formulated: for structured data, the focus was on verifying numerical range (e.g., lithium carbonate prices must not be lower than reasonable cost prices and must not be negative), data integrity (avoiding missing key fields, such as missing dates in price data or missing production areas in production data), and data consistency (the deviation of the same indicator from different sources should not exceed 5%; if the deviation is too large, manual verification is required to confirm the correct data); for semi-structured data, the focus was on verifying source reliability (prioritizing content released by authoritative institutions, well-known securities firms, and core enterprises) and information integrity (extracted key information must be complete, avoiding misinterpretation); for unstructured data, the focus was on verifying authenticity (eliminating fake news, forged images, and synthesized speech), relevance (eliminating content unrelated to the lithium carbonate market), and completeness (avoiding incomplete news text, blurry images, and broken speech). Data that failed verification was marked as abnormal data, and the reasons for the abnormality were recorded. It was either corrected through supplementary data collection or directly removed to ensure that the collected multimodal data was comprehensive, authentic, and effective, providing a reliable foundation for subsequent data preprocessing steps. The core working principle of this step is to: clarify the type and scope of multimodal data sources, cover long-term trend factors and short-term sudden factors affecting lithium carbonate prices, adopt diversified collection methods to ensure real-time and complete data, and eliminate noisy data through verification mechanisms to provide high-quality data input for subsequent steps, thus solving the problems of single data sources and poor data quality in existing methods.

[0027] Step 2: Multimodal data classification and standardization This step builds upon the multimodal data collected in Step 1. Its core objective is to classify and organize multimodal data of different types and formats, employing targeted standardization methods to transform all data types into data of a unified format and scale. This eliminates the influence of data format differences and units, laying the foundation for subsequent data preprocessing and feature extraction steps. The operation and working principle of this step are as follows: First, the multimodal data collected in Step 1 is categorized and organized to establish a classification database. Based on the three main data types identified in Step 1, the collected data is stored in structured, semi-structured, and unstructured databases, respectively. Within each database category, data is further subdivided by sub-type: Structured databases are subdivided into "price data," "supply and demand data," "macroeconomic data," and "policy quantitative data," with each sub-type sorted by time dimension (daily, weekly, monthly, quarterly, annual); semi-structured databases are subdivided into "industry research reports" and "company announcements," with each sub-type sorted by publication time and issuing institution; unstructured databases are subdivided into "news and public opinion data," "image data," and "audio data," with news and public opinion data sorted by publication time and keywords, image data sorted by collection time and location, and audio data sorted by recording time and speaker. Simultaneously, a unique identifier (including data type, collection time, and source) is added to each data entry to facilitate subsequent data traceability and management and prevent data confusion.

[0028] Secondly, targeted standardization methods are employed for different types of multimodal data to eliminate format differences and dimensional variations. For structured data, since it already has a numerical format, the focus is on dimensional standardization to eliminate dimensional differences between different indicators (e.g., lithium carbonate price is in yuan / ton, lithium ore production is in 10,000 tons, and macroeconomic data is in 100 million yuan; the dimensional differences are significant, and directly inputting them into the model will lead to a bias in model weights towards indicators with large dimensional variations). This step uses a combination of two standardization methods: for structured data that follows a normal distribution (e.g., monthly lithium carbonate price, monthly power battery installations), the Z-score standardization method is used to transform the data into standardized data with a mean of 0 and a standard deviation of 1; for structured data that does not follow a normal distribution (e.g., policy quantitative scores, lithium ore inventory), the Min-Max standardization method is used to map the data to the [0,1] interval. In practice, a normality test (using the Shapiro-Wilk test) is first performed on the data of each structured indicator, and the corresponding standardization method is selected based on the test results to ensure the rationality of the standardization process. Meanwhile, for missing values ​​in structured data (such as missing monthly lithium ore import volumes), linear interpolation is used to supplement them, avoiding the impact of missing values ​​on subsequent processing. Linear interpolation constructs a linear function by utilizing the effective data adjacent to the missing value, and calculates the estimated value of the missing value, ensuring that the supplemented data conforms to the data change trend, which is better than the existing method of simply deleting missing values.

[0029] For semi-structured data, the core is to convert it into a structured or standardized text format for subsequent collaborative processing. For industry research reports, through PDF parsing and text extraction techniques, key information (such as supply and demand forecast values, price trend judgments, core influencing factors) is extracted. Numerical information is converted into structured data (for example, "It is expected that the lithium carbonate production in 2025 will be 1 million tons" is converted into the structured format of "In 2025, lithium carbonate production forecast, 1 million tons"), and text information is converted into standardized text (removing punctuation marks, redundant words, and duplicate content, and unifying the font and encoding format). For enterprise announcements, key information such as production capacity, performance, and production capacity adjustment is extracted. Similarly, numerical information is converted into structured data, and text information is converted into standardized text to ensure that semi-structured data can be processed collaboratively with other types of data.

[0030] For unstructured data, diverse standardized processing methods are adopted to convert it into a format that can be processed subsequently. For news sentiment text data, text standardization processing is adopted: removing special symbols, punctuation marks, and irrelevant words (such as advertising terms and modal particles) in the text, performing Chinese word segmentation (using the jieba word segmentation tool), removing stop words (removing words with no practical meaning such as "de", "le", "shi", etc.), and converting the text into UTF-8 encoding and lowercase format to ensure the uniformity of news text formats from different sources. For image data, image standardization processing is adopted: unifying the size of all images (such as 224×224 pixels), unifying the color space (RGB color space), normalizing the image (mapping the pixel values of the image to the [0,1] interval), and removing noise in the image (using the Gaussian filtering method) to ensure the consistency and clarity of the image data. For voice data, voice standardization processing is adopted: unifying the sampling rate (such as 16kHz), unifying the bit depth (16 bits), removing background noise in the voice (using the Wiener filtering method), performing frame processing on the voice (frame length 25ms, frame shift 10ms), and converting the voice into a standardized audio frame format to lay a foundation for subsequent voice feature extraction.

[0031] Finally, a standardized data verification mechanism is established to ensure that the standardized data meets the requirements. For each type of standardized data, the key checks are: format uniformity (e.g., whether the units of structured data are consistent, whether the encoding of text data is consistent, and whether the size of image data is consistent), data integrity (e.g., whether new missing values ​​appear after standardization), and data rationality (e.g., whether the standardized values ​​are within a reasonable range, whether key information is retained in text data, and whether image data is clear and identifiable). Data that fails verification is re-standardized until it meets the requirements. The core working principle of this step is: different modal data have significantly different formats and units, making direct collaborative processing impossible. By classifying and organizing data to establish a standardized database, and designing personalized standardization strategies for different types of data, the effects of format differences and units are eliminated, missing values ​​are supplemented, and redundant information is removed. This ensures that all types of data can be collaboratively input into subsequent steps, solving the problem of inconsistent multimodal data formats and inability to be collaboratively utilized in existing methods.

[0032] Step 3: Personalized noise filtering and outlier correction for multimodal data This step, based on the standardized multimodal data from step 2, aims to design personalized noise filtering strategies tailored to the noise characteristics of different modalities. This involves removing noisy data, correcting outliers, further improving data quality, and preventing noise and outliers from interfering with subsequent feature extraction and model prediction, thus ensuring prediction accuracy. The operation and working principle of this step are as follows: First, we analyze the noise types and sources of different modalities of data to clarify the key points and difficulties of noise filtering. Noise in structured data mainly originates from statistical errors, data entry errors, and data source deviations (such as significant differences in data for the same indicator published on different platforms). Outliers are mainly manifested as sudden numerical changes (such as a single-day surge or plunge of more than 30% in lithium carbonate prices, exceeding the normal market fluctuation range). Noise in semi-structured data mainly originates from redundant information, vague statements, and false information in reports or announcements (such as exaggerated or false supply and demand forecasts in research reports published by some non-authoritative institutions). Outliers are mainly manifested as contradictory information (such as inconsistent statements about production capacity in two announcements issued by the same company). The sources of noise in unstructured data are the most complex. Noise in news and public opinion data mainly consists of fake news, irrelevant information, and duplicate news. Outliers are mainly extremely negative / positive false public opinion (such as fabricated news about lithium mine safety accidents). Noise in image data mainly consists of blurred images, noise points, and irrelevant backgrounds (such as images of lithium mine sites containing a large number of irrelevant personnel or equipment). Outliers are mainly image distortion and fabricated images. Noise in voice data mainly consists of environmental noise and recording interference. Outliers are mainly voice breaks, fabricated voice, and irrelevant remarks.

[0033] Secondly, for structured data, a dual processing strategy of "noise filtering + outlier correction" is designed. The first step, noise filtering, employs an adaptive filtering method based on a sliding window. Different sliding window sizes are set for the time characteristics of different structured indicators (e.g., a 7-day sliding window for daily price data, and a 3-month sliding window for monthly production data). The mean and standard deviation of the data within the sliding window are calculated. Data deviating from the mean by more than twice the standard deviation is marked as suspected noise data. This is verified in conjunction with the reliability of the data source. If confirmed as noise data (e.g., data entry errors, statistical errors), the mean within the window is used to replace it; if the data exhibits reasonable fluctuations, it is retained. Simultaneously, for the same structured indicator data from different sources, a weighted average method is used to merge the data (the weight of authoritative data sources is set to 0.7, and the weight of ordinary data sources is set to 0.3) to eliminate noise caused by data source bias. The second step is outlier correction: Outliers are identified using the Grubbs test (significance level α=0.05). For identified outliers, a judgment is made based on the actual market situation. If the outlier is a reasonable anomaly caused by short-term, sudden factors (such as a price surge due to a major policy adjustment), the data is retained and marked as an outlier for targeted processing in subsequent models. If the outlier is unreasonable (such as data entry errors or false data), a trend-fitting correction method is used. A quadratic polynomial fitting function is constructed using valid data before and after the outlier to calculate the corrected value, replacing the original outlier and ensuring that the corrected data conforms to market trends. For example, if the price of lithium carbonate on a certain day is displayed as 1 million yuan / ton due to an entry error (far exceeding the normal price range), a reasonable price correction value is calculated by fitting the price data of the 10 days before and after the error using a quadratic polynomial, replacing the erroneous data.

[0034] For semi-structured data, a noise filtering strategy of "source verification + information filtering + contradiction correction" is designed. The first step, source verification, involves establishing an authoritative source database (including well-known securities firms, industry associations, and core industry chain companies) to verify the issuing institutions of industry research reports and company announcements, eliminating content from non-authoritative institutions to reduce false information noise. The second step, information filtering, uses a combination of keyword matching and semantic analysis to filter out key information related to lithium carbonate prices, eliminating redundant information and vague statements (such as industry background descriptions in reports unrelated to prices and supply and demand), retaining core content. The third step, contradiction correction, compares the statements of the same indicator (such as company capacity and supply and demand forecasts) in different reports or announcements. If contradictions exist, they are verified in conjunction with authoritative sources and actual market conditions to confirm the correct information and correct contradictory data. If verification is not possible, the data is marked as suspicious and temporarily removed to avoid interfering with subsequent processing.

[0035] For unstructured data, personalized noise filtering and outlier correction strategies are employed. For news and public opinion text data: First, duplicate news is removed using text similarity calculation (cosine similarity). News with a similarity exceeding 0.8 is considered duplicate news, retaining only the earliest published and most authoritative source. Second, a fake news identification method based on the BERT model is used. The news text is input into a pre-trained BERT model to identify fake news (the model is trained on labeled real / fake news datasets and can capture semantic contradictions and false statements in the text), removing fake news. Third, sentiment analysis is used to mark news with extreme sentiment (sentiment value > 0.9 or < 0.1) and no actual basis as abnormal public opinion. This is verified in conjunction with actual market conditions; if it is false extreme public opinion, it is removed; if it is a real extreme event (such as a major safety accident), it is retained and marked. For image data: First, Gaussian filtering is used to remove image noise, and an edge detection algorithm (Canny algorithm) is used to extract core regions (such as lithium mining equipment and production lines) from the image, removing irrelevant backgrounds and retaining the core regions. Second, a CNN-based image distortion recognition model is used to identify blurry, distorted, and forged images, treating them as abnormal images and removing them. Third, the retained images are scored in terms of quality (scores are given in three dimensions: clarity, completeness, and relevance, with a maximum score of 1 point), retaining images with a score > 0.7 to ensure the validity of the image data. For speech data: First, Wiener filtering is used to remove environmental noise, and a speech enhancement algorithm is used to improve speech clarity. Second, a speech anomaly recognition method based on MFCC features is used to identify speech breaks, forged speech, and irrelevant remarks, treating them as abnormal speech and removing them. Third, the retained speech is segmented and preprocessed to ensure that the speech data can be used for subsequent feature extraction.

[0036] Finally, the multimodal data after noise filtering and outlier correction are comprehensively validated. The noise removal rate and outlier correction rate for each type of data are calculated to ensure that the noise removal rate is >90% and the outlier correction rate is >95%. Simultaneously, the integrity and rationality of the data are checked to avoid data loss due to over-filtering. The core working principle of this step is that the noise characteristics of different modalities vary significantly. Existing methods using a uniform noise filtering approach are ineffective. This step designs personalized processing strategies based on the noise type and source of each type of data, distinguishing between reasonable fluctuations and noise, and between real and false anomalies. While removing noise and false anomalies, it retains real sudden anomaly data, further improving data quality and providing high-quality data input for subsequent feature extraction. This addresses the problems of poor noise filtering and inability to distinguish between reasonable and false anomalies in existing methods.

[0037] Step 4: Extraction of trend and correlation features from structured data This step, based on the structured data obtained from noise filtering and outlier correction in Step 3, aims to extract trend and correlation features reflecting lithium carbonate price fluctuations from the structured data. It seeks to uncover the inherent patterns within the structured data, providing structured feature support for subsequent multimodal feature fusion. Simultaneously, it analyzes the correlation between various structured indicators and lithium carbonate prices, providing a basis for subsequent feature weight allocation. The operation and working principle of this step are as follows: First, clarify the scope and objectives of feature extraction from structured data. Structured data includes five main categories: historical lithium carbonate price data, lithium ore supply and demand data, battery industry data, macroeconomic data, and policy quantitative data. The extraction objectives of this step are: to extract long-term price trends, short-term fluctuations, and cyclical patterns from historical price data; to extract price-related correlation features from the other four categories of data; and to analyze the correlation between various indicators and prices, selecting highly correlated features and avoiding feature redundancy.

[0038] Secondly, we extract trend features from the structured data, focusing on historical lithium carbonate price data. The first step is long-term trend feature extraction: using a combination of moving average (MA) and exponential smoothing (EMA), we calculate short-term moving averages (7-day and 30-day MA), long-term moving averages (90-day and 180-day MA), and an exponential smoothing value (α=0.3, where α is the smoothing coefficient used to balance the weights of recent and long-term data). By analyzing the crossover relationship between short-term and long-term MAs (e.g., a short-term MA crossing above a long-term MA is considered an upward signal, and a crossing below is considered a downward signal), we extract long-term price trend features (upward trend, downward trend, and sideways trend). Simultaneously, we use linear regression to fit the price time series and calculate regression coefficients to quantify the strength of the price trend (a positive regression coefficient indicates an upward trend, with a larger coefficient indicating a stronger trend; a negative regression coefficient indicates a downward trend, with a larger absolute value indicating a stronger trend). The second step is to extract short-term volatility characteristics: calculate the daily volatility and weekly volatility of prices (volatility = (current price - previous price) / previous price × 100%), and use standard deviation and variance to quantify the volatility amplitude; at the same time, extract the extreme value characteristics of prices (daily high price, low price, closing price, weekly / monthly high price, low price, average price), as well as the symmetry of volatility (measured by the skewness coefficient; skewness > 0 indicates right skewness, that is, upward volatility is greater than downward volatility; skewness < 0 indicates left skewness, that is, downward volatility is greater than upward volatility), to comprehensively capture the short-term volatility characteristics of prices. The third step is to extract periodic features: Fourier transform (FT) is used to decompose the price time series into trend, periodic, and stochastic components. By analyzing the frequency and amplitude of the periodic component, the periodic features of the price (such as quarterly and annual cycles) are extracted to clarify the periodic patterns of price fluctuations. For example, Fourier transform reveals that lithium carbonate prices have a clear annual cycle, which is closely related to the seasonality of lithium mining (such as the difficulty of mining in winter, resulting in decreased output and increased prices).

[0039] Then, the correlation features of the structured data are extracted to analyze the degree of correlation between each indicator and lithium carbonate prices. The first step involves extracting correlation features from lithium mine supply and demand data, battery industry data, macroeconomic data, and policy quantitative data: For lithium mine supply and demand data, features such as supply-demand difference (supply - demand), inventory turnover rate, and import dependence (imports / total supply) are extracted; for battery industry data, features such as the growth rate of installed power battery capacity, the growth rate of new energy vehicle production, and battery recycling rate are extracted; for macroeconomic data, features such as GDP growth rate, PPI growth rate, and RMB exchange rate fluctuation range are extracted; and for policy quantitative data, features such as policy support strength score and mining control intensity score are extracted. The second step uses correlation analysis to analyze the degree of correlation between each extracted feature and lithium carbonate prices, selecting features with high correlation (significance level α = 0.05, absolute value of correlation coefficient > 0.5 is considered highly correlated). Specifically, a combination of Pearson correlation coefficient and Spearman rank correlation coefficient is used: Pearson correlation coefficient is used to analyze linear correlation, and Spearman rank correlation coefficient is used to analyze nonlinear correlation, avoiding omissions caused by a single correlation coefficient. For example, correlation analysis revealed that the correlation coefficients between the supply and demand gap of lithium ore, the growth rate of power battery installations, and the price of lithium carbonate were -0.72 and 0.81, respectively, which are highly correlated features and were retained. However, the correlation coefficient between the GDP growth rate and the price of lithium carbonate was 0.35, which is a low-correlation feature and was removed to avoid feature redundancy.

[0040] Finally, the extracted trend and correlation features are normalized (using the Min-Max normalization method, mapping to the [0,1] interval) to eliminate dimensional differences between features and establish a structured feature set, providing support for subsequent multimodal feature fusion steps. Simultaneously, the correlation coefficients between each feature and lithium carbonate prices are recorded as an important basis for subsequent feature weight allocation. The core working principle of this step is: structured data is the core data reflecting the long-term trend of lithium carbonate prices, containing rich trend patterns and correlations. By extracting trend and correlation features through various mathematical methods, the intrinsic value of structured data can be fully explored. Correlation analysis is used to screen high-value features, avoiding feature redundancy and providing high-quality structured features for subsequent multimodal feature fusion. This solves the problem of existing methods only utilizing a single price feature and failing to explore the intrinsic correlations within structured data.

[0041] Step 5: Semantic and Visual / Speech Feature Extraction from Unstructured Data This step builds upon the unstructured data obtained from noise filtering and outlier correction in Step 3. Its core objective is to employ deep learning and natural language processing techniques to extract semantic, visual, and speech features from unstructured data such as news and public opinion texts, images, and audio. This transforms the unstructured data into quantifiable feature vectors, uncovering short-term, emerging factors within the unstructured data. This provides unstructured feature support for subsequent multimodal feature fusion, overcoming the shortcomings of existing methods that neglect the value of unstructured data. The operation and working principle of this step are as follows: First, semantic and sentiment features are extracted from news and public opinion text data. News and public opinion text data contains a large amount of information reflecting short-term emergencies (such as safety accidents, policy adjustments, and supply chain disruptions). Its core features are semantic information and sentiment bias. This step uses a text feature extraction method based on an improved BERT model. The specific operations are as follows: First, the standardized and denoised news text is segmented, part-of-speech tagging is performed, and entity recognition is conducted (using jieba segmentation + BiLSTM-CRF entity recognition model). Core entities in the text are identified (such as "lithium mine," "lithium carbonate," "new energy vehicles," "policy," "safety accident," etc.), and entity types are labeled (such as industry entities, event entities, policy entities), laying the foundation for semantic feature extraction. Second, a pre-trained BERT model for the lithium carbonate industry is constructed (based on the general BERT model, using…). The model is fine-tuned using news articles and research reports from the lithium carbonate industry to improve its semantic understanding of industry texts. The third step involves inputting the news text into the fine-tuned BERT model and extracting the output vector of the last layer as the text's semantic feature vector (768 dimensions). This vector comprehensively captures the text's semantic information (e.g., the semantic feature of "lithium mining accidents leading to production decline" clearly reflects a signal of reduced supply). The fourth step uses an LSTM-based sentiment analysis model to score the news text's sentiment (score range [0,1], where 0 represents extreme negativity and 1 represents extreme positivity), extracting sentiment features to reflect the short-term impact of public opinion on lithium carbonate prices (e.g., negative public opinion leads to a short-term price drop, while positive public opinion leads to a short-term price increase). Simultaneously, the extracted semantic feature vector is dimensionality-reduced (using PCA principal component analysis to reduce the 768-dimensional vector to 128 dimensions) to eliminate feature redundancy and improve subsequent processing efficiency.

[0042] Secondly, for the image data, visual features are extracted, focusing on reflecting capacity utilization and supply-demand changes. Image data (lithium mining sites, lithium carbonate production lines, new energy vehicle production workshops) can intuitively reflect the capacity situation of the industry chain (such as the operating rate of mining equipment and the load of production lines). This step adopts a visual feature extraction method based on an improved CNN model, specifically as follows: First, an improved ResNet-50 model for the lithium carbonate industry chain is constructed (based on the ResNet-50 model, the last fully connected layer is removed, an adaptive average pooling layer is added, and fine-tuning is performed using industry image data to improve the model's feature extraction capability for industry images); Second, data augmentation processing is performed on the denoised and standardized image data (using random cropping, ... The process involves several steps: First, the enhanced image is flipped, rotated, and its brightness adjusted to improve generalization ability and avoid overfitting. Second, the enhanced image is input into the improved ResNet-50 model, and the output vector of the model's adaptive average pooling layer is extracted as the image's visual feature vector (2048 dimensions). This vector captures core visual information in the image (such as the number of lithium mining equipment in operation and the operating status of the production line, thus reflecting capacity utilization). Third, the visual feature vector is dimensionality reduced (using the t-SNE dimensionality reduction method to reduce the 2048-dimensional vector to 128 dimensions) to eliminate redundancy and ensure that visual features can be synergistically integrated with other modal features. Simultaneously, an image classification model is used to classify the image into three levels of capacity utilization (high, medium, and low). The classification results are then converted into numerical features (high = 1, medium = 0.5, low = 0) to supplement the quantitative information of the visual features and improve their practicality.

[0043] Then, for the speech data, speech features are extracted to capture key industry viewpoints. The speech data (speech from industry seminars and company press conferences) contains the judgments and predictions of industry experts and company leaders regarding the lithium carbonate market. Its core features are the Mel-frequency cepstral coefficients (MFCC) and semantic information. This step employs a dual approach of "speech feature extraction + semantic transformation," specifically as follows: First, the standardized and denoised speech data is framed and windowed (frame length 25ms, frame shift 10ms, using a Hanning window). The MFCC features of each frame are calculated (13-dimensional MFCC coefficients are extracted, along with first-order and second-order differences, for a total of 39 dimensions). MFCC features can capture the spectral characteristics of speech, reflecting... The first step involves reflecting the speaker's tone and emotions. The second step uses a speech-to-text model based on Wav2Vec2.0 to convert the speech data into text data. Then, the improved BERT model from step 5 is used to extract the semantic feature vector of the text (reduced to 128 dimensions) to capture the core viewpoints in the speech (such as a company leader announcing capacity expansion, which can reflect the signal of increased supply). The third step involves concatenating the MFCC features and semantic feature vectors to obtain the comprehensive feature vector of the speech (39+128=167 dimensions). Then, PCA is used to reduce the dimensions to 128 dimensions to ensure that the dimensions are consistent with other unstructured features, which is convenient for subsequent fusion.

[0044] Finally, the extracted text semantic features, sentiment features, image visual features, and speech comprehensive features are normalized (mapped to the [0,1] interval) to establish an unstructured feature set, providing support for subsequent multimodal feature fusion steps. Simultaneously, the effectiveness of each type of unstructured feature is tested (using analysis of variance to verify whether the features can effectively distinguish different price fluctuation scenarios), eliminating invalid features to ensure the practicality of the unstructured features. The core working principle of this step is: unstructured data contains a large amount of information reflecting short-term sudden factors and cannot be directly used in prediction models. Through deep learning and natural language processing techniques, unstructured data is transformed into quantifiable feature vectors, extracting semantic, visual, and speech information to explore the impact of short-term sudden factors on prices, providing unstructured feature support for subsequent multimodal feature fusion, and compensating for the deficiency of existing methods in ignoring the value of unstructured data.

[0045] Step 6: Extraction of key information and feature quantization from semi-structured data This step builds upon the semi-structured data obtained from noise filtering and outlier correction in Step 3. Its core objective is to extract key information from semi-structured data such as industry research reports and company announcements, quantifying it into processable feature vectors. This process uncovers information such as supply and demand forecasts and capacity adjustments within the semi-structured data, establishing a semi-structured feature set to supplement subsequent multi-modal feature fusion and achieve comprehensive collaboration among structured, unstructured, and semi-structured data. The operation and working principle of this step are as follows: First, the scope of key information extraction from semi-structured data is clearly defined. Considering the factors influencing lithium carbonate price fluctuations, the key information is categorized into three main types: supply and demand forecasts (such as forecasts of lithium carbonate production, demand, and price trends for the next 1-3 years in the report), capacity adjustment information (such as capacity expansion and contraction plans in company announcements, and changes in capacity utilization), and analysis of core influencing factors (such as the key factors affecting prices mentioned in the report and their degree of impact). These three types of information reflect market expectations for lithium carbonate prices and changes in the medium- to long-term supply and demand pattern, providing significant reference value for price forecasting and representing core content that existing methods have not fully explored.

[0046] Secondly, for industry research reports, a method of "key information positioning + semantic analysis + quantitative scoring" is used to extract and quantify key information. The first step, key information positioning, employs a keyword-based and regular expression-based approach to locate key paragraphs in the report, such as supply and demand forecasts, capacity adjustments, and influencing factor analyses. For example, the keywords "production forecast," "demand forecast," and "price forecast" are used to locate the supply and demand forecast paragraph; "capacity expansion," "capacity contraction," and "capacity utilization rate" are used to locate the capacity adjustment paragraph; and "influencing factors" and "price drivers" are used to locate the core influencing factor paragraph. Simultaneously, a text paragraph classification model (based on a CNN-BiLSTM model) is used to classify the report paragraphs, filtering out paragraphs containing key information and eliminating irrelevant paragraphs (such as industry background introductions and company profiles). The second step is semantic parsing: For the selected key paragraphs, the improved BERT model (consistent with the fine-tuned BERT model in step 5, which enhances the semantic understanding of industry texts) is used for semantic parsing to extract key information. For example, from the supply and demand forecast paragraph, the forecast time, forecast output, forecast demand, and forecast price range are extracted; from the capacity adjustment paragraph, the capacity adjustment range, adjustment time, and adjustment reasons are extracted; and from the influencing factors paragraph, the name of the influencing factors, the direction of influence (positive / negative), and the description of the degree of influence (such as "significant influence" or "minor influence") are extracted. The third step is key information quantification: the extracted key information is transformed into quantifiable features. The specific quantification rules are as follows: For supply and demand forecast information, the forecast output and forecast demand are transformed into numerical features (unit is uniformly 10,000 tons), and the forecast price range is transformed into a price forecast index (e.g., if the forecast price increases by 10%-20%, the index is set to 1.15; if the forecast price decreases by 5%-10%, the index is set to 0.925; if the forecast price remains unchanged, the index is set to 1.0); For capacity adjustment information, the capacity expansion rate (e.g., expansion by 20%) and contraction rate (e.g., contraction by 10%) are transformed into numerical features, the capacity utilization rate is transformed into a quantitative value in the range [0,1], and the capacity adjustment time is transformed into the time elapsed since the current time (unit is months); For the analysis of core influencing factors, the description of the degree of influence is transformed into a quantitative score (significant influence = 1.0, moderate influence = 0.7, slight influence = 0.3), the direction of influence is transformed into a sign feature (positive influence = +1, negative influence = -1), and the name of the influencing factor is transformed into a unique thermal coding feature (e.g., "lithium mine output" corresponds to one coding dimension, "policy regulation" corresponds to another coding dimension).

[0047] Then, for company announcements, a method of "information extraction + authenticity verification + quantification" is used to extract and quantify key information. The first step, information extraction, employs PDF parsing and text extraction technologies to extract key information from the announcements, such as production capacity, performance, capacity adjustments, and cooperation agreements. For example, from production capacity announcements, current capacity, planned capacity, and capacity expansion / contraction timelines are extracted; from performance announcements, revenue, profit, and capacity utilization rates related to lithium carbonate businesses are extracted; from cooperation agreement announcements, the content of cooperation (such as lithium mining cooperation and capacity cooperation) and the scale of cooperation are extracted. This information reflects the company's operating status and capacity changes, thus affecting the supply and demand pattern of lithium carbonate. The second step, authenticity verification, combines past company announcements with actual industry conditions to verify the authenticity of the extracted information. For example, it verifies whether the company's capacity expansion plan has sufficient funding and technical support, and whether performance data is consistent with overall industry trends, avoiding interference from false announcements. False or questionable information is removed to ensure the authenticity of the extracted information. The third step is information quantification: the extracted key information is converted into quantitative features, and the specific rules are as follows: production capacity data is converted into numerical features (in ten thousand tons), capacity utilization rate is converted into a quantitative value in the range of [0,1], revenue and profit are converted into growth rate features (growth rate compared with the previous period), cooperation scale is converted into numerical features (e.g., if the scale of cooperative lithium mining is 50,000 tons / year, it is converted into 5.0), and capacity adjustment time is converted into the time from the current time (in months).

[0048] Finally, the quantified semi-structured features were integrated and optimized: First, correlation analysis (Pearson correlation coefficient) was used to analyze the degree of correlation between each semi-structured feature and the price of lithium carbonate, and features with an absolute correlation coefficient > 0.4 were selected, while redundant features (such as supply and demand forecast features repeated in different reports) were removed; Second, the selected features were normalized (Min-Max standardization, mapped to the [0,1] interval) to eliminate the influence of dimensions; Third, all quantified semi-structured features were concatenated to form a semi-structured feature set (with a dimension set to 128), ensuring that the dimensions were compatible with the structured and unstructured feature sets, providing supplementary support for subsequent multimodal feature fusion steps. Simultaneously, the source and credibility score of each semi-structured feature were recorded (credibility of authoritative sources = 1.0, credibility of ordinary sources = 0.7), serving as the basis for subsequent feature weight allocation. The core working principle of this step is as follows: Semi-structured data contains a large amount of market expectations and medium- to long-term supply and demand information, serving as a key bridge connecting structured data (long-term trends) and unstructured data (short-term emergencies). By extracting and quantifying key information, semi-structured data is transformed into processable feature vectors, mining market expectations and capacity change information within them, supplementing the deficiencies of structured and unstructured features, achieving comprehensive coverage of the three types of multimodal data, and solving the problem that existing methods do not fully utilize semi-structured data.

[0049] Step 7: Multimodal Feature Weighted Collaborative Fusion and Feature Optimization This step builds upon the structured feature set from step 4, the unstructured feature set from step 5, and the semi-structured feature set from step 6. Its core objective is to design a multimodal feature weighted collaborative fusion strategy. By combining the importance and reliability of each modality's features, reasonable feature weights are assigned, fusing the three types of features into a unified fused feature set. Simultaneously, the fused features are optimized, redundant features are removed, and feature quality is improved, providing high-quality input features for subsequent prediction models. This achieves collaborative driving of multimodal data and addresses the problem of simple multimodal feature fusion and poor collaborative performance in existing methods. The operation and working principle of this step are as follows: First, the characteristics and importance of the three types of multimodal features are analyzed to determine the basis for feature weight allocation. Structured features primarily reflect the long-term trend of lithium carbonate prices and supply and demand fundamentals, serving as the core for predicting long-term price trends and thus having the highest importance. Semi-structured features primarily reflect market expectations and medium- to long-term capacity changes, acting as a bridge between long-term trends and short-term emergencies, and thus have the next highest importance. Unstructured features primarily reflect short-term emergencies (such as public opinion and sudden accidents), significantly impacting short-term price fluctuations but having a weaker long-term impact, and thus have relatively lower importance. Furthermore, the importance of different features within each modality also varies (e.g., in structured features, the correlation between the growth rate of power battery installations and prices is higher than that of GDP growth; in unstructured features, news and public opinion from authoritative media are more important than those from ordinary media). Therefore, this step adopts a dual weighting strategy of "modal weight + feature weight" to ensure the rationality and relevance of the weight allocation.

[0050] Secondly, the weights of each modality are calculated using a modality weight allocation method based on information entropy. This method assigns weights according to the information content of each modality feature; the greater the information content, the higher the weight. The specific steps are as follows: First, calculate the information entropy of each modality feature set. Information entropy measures the uncertainty and information content of a feature; the smaller the information entropy, the greater the information content of the feature and the greater its contribution to prediction. The formula for calculating information entropy is:

[0051] in, Information entropy represents the feature set of a certain modality; This indicates the number of features in the feature set of this modality; Indicates the first The probability density of each feature (calculated using the kernel density estimation method, i.e., estimating the probability density of each feature based on the sample distribution of the feature). Indicates the first The higher the probability density of a feature, the higher its information content; the smaller the negative value, the lower the information entropy.

[0052] The second step is to calculate the weight of each mode based on the information entropy. The formula for calculating the mode weight is as follows:

[0053] in, Indicates the first Modal class ( For structured purposes, For unstructured, Weights for semi-structured data; Indicates the first Information entropy of modal feature sets; Representing three types of modes The sum of these values ​​is used for normalization to ensure that the sum of all modal weights is 1. ).

[0054] For example, suppose we calculate the information entropy of the structured feature set. Information entropy of unstructured feature sets Information entropy of semi-structured feature sets ,but: 1- =0.7, 1- =0.4, 1- =0.6 The total is 0.7 + 0.4 + 0.6 = 1.7 Structured Modal Weights Unstructured mode weights Semi-structured modal weights The sum of the weights is 1, which meets the requirements. This result indicates that the structured features have the largest information content and the highest weight, consistent with expectations.

[0055] Then, the weights of features within each mode are calculated using a feature weight allocation method based on correlation coefficients and confidence levels. The first step involves calculating the correlation coefficient between each feature within each mode and the price of lithium carbonate. (Pearson correlation coefficient, ranging from -1 to 1), and take the absolute value of the correlation coefficient to obtain The first step is to measure the correlation between features and prices; the second step is to combine the feature's credibility score. (Structured feature credibility = 1.0, semi-structured feature credibility is determined based on the source, unstructured features: authoritative source credibility = 1.0, ordinary source credibility = 0.7); Third step, calculate the weight of each feature. The calculation formula is:

[0056] in, Indicates the first In the class modality, the first The weights of each feature; Indicates the first The absolute value of the correlation coefficient between each characteristic and the price of lithium carbonate; Indicates the first Credibility score for each feature; Indicates the first The number of features in a class modality; Indicates the first All features in class modality The sum of these weights is used for normalization to ensure that the sum of the weights of all features within each modality is 1.

[0057] For example, in structured modes, the correlation coefficient of the installed capacity growth rate of power batteries Credibility Correlation coefficient of lithium ore supply and demand gap Credibility If the mode has a total of 10 features, all features The sum is 5.2, so the weight of the power battery installation growth rate is... Weight of lithium mine supply and demand imbalance Ensure that the sum of the feature weights within this modality is 1.

[0058] Next, we perform multimodal feature weighted collaborative fusion, specifically as follows: First, for the features within each modality, we use a weighted summation method to fuse all features of that modality into a single modal feature vector (128 dimensions) – that is, for the ... Modal class, fused modal feature vector = ,in Indicates the first A vector of features, The first step represents the weight of the feature; the second step is to use modal weights to weight and fuse the feature vectors of the three modalities to obtain the final fused feature set. The calculation formula is:

[0059] in, This represents the final fused feature set (128 dimensions). These represent the weights of the structured, unstructured, and semi-structured modes, respectively. These represent the feature vectors after the fusion of the three modalities. This formula achieves the synergistic fusion of three types of multimodal features, highlighting the core role of structured features while also considering the supplementary roles of semi-structured and unstructured features, thus maximizing the value of multimodal data.

[0060] Finally, the fused feature set is optimized to improve feature quality: First, analysis of variance (ANOVA) is used to test the effectiveness of the fused features and remove invalid features with variance less than 0.01 (such features cannot distinguish different price fluctuation scenarios); second, mutual information is used to analyze the redundancy between the dimensions within the fused features and remove redundant features with mutual information values ​​greater than 0.8 (to avoid mutual interference between features); third, principal component analysis (PCA) is used to reduce the dimensionality of the optimized fused features, adjusting the feature dimensions to 64 dimensions, which improves the training efficiency and generalization ability of the subsequent prediction model while retaining core information; fourth, the dimensionality-reduced fused features are normalized to ensure that all features are in the [0,1] interval, laying the foundation for the input of the subsequent prediction model. The core working principle of this step is that the fusion quality of multimodal features directly determines the prediction accuracy. Most existing methods use simple concatenation, which cannot reflect the differences in importance of each modality and feature. This step uses a dual weighting strategy, combining information entropy, correlation coefficient, and confidence, to reasonably allocate modality weights and feature weights, thereby achieving the collaborative fusion of multimodal features. At the same time, feature optimization is used to eliminate redundant and invalid features, improving feature quality and achieving collaborative driving of multimodal data. This solves the problem of simple multimodal feature fusion and poor synergy in existing methods.

[0061] Step 8: Construction and Training of the Two-Branch Collaborative Prediction Model This step, based on the optimized fusion feature set from step 7, aims to construct a dual-branch collaborative prediction model that captures both the long-term trend and short-term fluctuations in lithium carbonate prices. Through a branch-collaboration mechanism, it achieves accurate fitting between the long-term trend and short-term fluctuations, improving prediction accuracy and generalization ability, and addressing the problem that existing models cannot simultaneously adapt to both long-term trends and short-term fluctuations. The operation and working principle of this step are as follows: First, the overall structure of the dual-branch collaborative prediction model is defined. This model consists of a long-term trend prediction branch and a short-term fluctuation prediction branch. The two branches share the input of a fused feature set and achieve information exchange and weight allocation through a collaborative attention mechanism, ultimately outputting a fused price prediction result. Specifically, the long-term trend prediction branch is mainly used to capture the long-term trend of lithium carbonate prices (such as quarterly or annual trends), adapting to both structured and semi-structured features; the short-term fluctuation prediction branch is mainly used to capture short-term price fluctuations (such as daily or weekly fluctuations), adapting to unstructured features. The two branches work collaboratively, balancing the prediction accuracy of both long-term trends and short-term fluctuations.

[0062] Secondly, a long-term trend prediction branch is constructed using an improved Transformer model, focusing on capturing long-term price dependencies and medium- to long-term supply and demand patterns. The first step is model structure improvement: based on the standard Transformer model, a positional encoding optimization mechanism is introduced (using a concatenation of sine and cosine positional encoding with time feature encoding; the time feature encoding combines the daily, weekly, and monthly time attributes of price data to improve the model's sensitivity to time series); simultaneously, the encoder structure is simplified, reducing the number of attention heads (from 12 to 8), lowering model complexity, avoiding overfitting, while retaining the encoder's core ability to capture long-distance dependencies. The second step is input layer design: the optimized fusion feature set (64 dimensions) from step 7 is concatenated with trend features (such as long-term moving averages and regression coefficients) from the structured feature set as input to the long-term trend prediction branch (total dimensions are 64 + 16 = 80 dimensions), ensuring that the input features fully reflect long-term trend information. The third step is the output layer design: a linear regression layer is used as the output layer to output the predicted price of lithium carbonate for the next 1-3 months (monthly forecast). At the same time, a Dropout layer is added (dropout probability is set to 0.3) to improve the generalization ability of the model and avoid overfitting during the training process.

[0063] Then, a short-term fluctuation prediction branch is constructed, employing an improved CNN-LSTM hybrid model to focus on capturing short-term sudden fluctuations and random changes in prices. The first step is model structure improvement: a CNN model is used as the feature extraction front-end, employing three convolutional layers (kernel sizes of 3, 5, and 7) to extract local features from the fused feature set (such as unstructured features reflecting short-term public opinion and local fluctuation signals caused by sudden accidents). Max pooling layers (with a pooling kernel size of 2) are used to reduce feature dimensionality and retain key information. An LSTM model is used as the temporal prediction back-end, employing two LSTM layers (with 128 and 64 hidden units respectively) to capture the temporal variation patterns of local features, addressing the CNN model's inability to capture temporal dependencies. An attention mechanism (channel attention mechanism) is added between the CNN and LSTM layers to weight the local features extracted by the CNN, focusing on features that significantly impact short-term fluctuations (such as public opinion sentiment features and image capacity utilization features). The second step, input layer design: The optimized fusion feature set (64 dimensions) from step 7 is concatenated with the semantic, visual, and speech features from the unstructured feature set (short-term relevant features selected from the 128-dimensional features after dimensionality reduction, with a dimension of 32) as the input to the short-term fluctuation prediction branch (total dimension 64+32=96 dimensions), ensuring that the input features can fully reflect short-term fluctuation information. The third step, output layer design: A linear regression layer is used as the output layer to output the predicted lithium carbonate price for the next 1-7 days (daily prediction). A BatchNorm layer is added to accelerate model training convergence and further improve the model's generalization ability.

[0064] Next, a dual-branch collaboration mechanism is designed to achieve information interaction and prediction result fusion between the two branches. The first step is the collaborative attention mechanism design: a collaborative attention module is added before the output layers of the two branches. This module extracts the feature representations of the long-term trend prediction branch and the short-term volatility prediction branch, calculates the mutual information between the features of the two branches, and assigns attention weights based on the magnitude of the mutual information—when short-term volatility is high, the attention weight of the short-term volatility prediction branch is increased (weight range 0.5-0.7); when the market is in a stable trend, the attention weight of the long-term trend prediction branch is increased (weight range 0.6-0.8), achieving dynamic collaboration between the two branches. The second step is prediction result fusion: a weighted summation method is used to fuse the prediction results of the two branches to obtain the final price prediction value (multi-scale prediction combining daily and monthly data). The fusion weight is determined by the attention weight output by the collaborative attention module, and the calculation formula is as follows: ,in This represents the final predicted value. This represents the predicted value of the long-term trend branch. This represents the predicted value of the short-term fluctuation branch. Represents the collaborative attention weights (0 < <1).

[0065] Finally, model training and hyperparameter optimization are performed. The first step is dataset partitioning: the optimized fused feature set from step 7 and the corresponding historical lithium carbonate price data are divided into training, validation, and test sets in a 7:2:1 ratio. The training set is used for model parameter training, the validation set for hyperparameter tuning and model generalization verification, and the test set for final model performance evaluation. Simultaneously, data augmentation techniques (time series shifting and noise addition) are used to expand the training set data volume, further improving the model's generalization ability. The second step is loss function selection: a composite loss function combining mean squared error (MSE) and mean absolute percentage error (MAPE) is adopted, taking into account both the absolute and relative errors of the predicted values. The loss function calculation formula is: The MSE measures the absolute deviation between the predicted and actual values, while the MAPE measures the relative deviation, avoiding error assessment bias caused by a large price range. The third step involves optimizer selection and training parameter settings: the Adam optimizer is used, with an initial learning rate of 0.001. A learning rate decay strategy is employed (decreasing the learning rate to 0.9 every 10 epochs) to accelerate model convergence. The number of training epochs is set to 100, the batch size to 32, and an early stopping strategy (patience=15) is used. Training stops when the validation set loss does not decrease for 15 consecutive epochs to prevent overfitting. The fourth step is hyperparameter optimization: a grid search method is used to optimize key hyperparameters (such as the number of Transformer attention heads, the number of LSTM hidden layer units, dropout probability, and convolutional kernel size), selecting the hyperparameter combination with the minimum validation set loss as the final hyperparameters for the model. The core working principle of this step is to adapt the prediction needs of long-term trends and short-term fluctuations to a dual-branch structure, use a collaborative attention mechanism to achieve dynamic interaction between the two branches, and fuse multi-scale prediction results to solve the problem that existing models cannot simultaneously take into account both long-term trends and short-term fluctuations, thereby improving prediction accuracy and generalization ability.

[0066] Step 9: Model Prediction and Preliminary Verification of Prediction Results This step builds upon the bi-branch collaborative prediction model trained in step 8. Its core objective is to utilize this model for multi-scale prediction of lithium carbonate prices. Simultaneously, it establishes a preliminary verification mechanism to assess the reasonableness and accuracy of the prediction results, eliminating obviously unreasonable predictions and providing a foundation for subsequent optimization, thus ensuring the practicality of the predictions. The operation and working principle of this step are as follows: First, the preparation and preprocessing of the prediction input data are clearly defined. The multimodal data (structured, semi-structured, and unstructured data) collected in real-time and processed through steps 1-7 are processed into a fusion feature set (64 dimensions) that meets the model input requirements, following the standardization, denoising, feature extraction, feature fusion, and feature optimization process in steps 2-7. This ensures that the format and scale of the input data are consistent with the input data used during model training. Simultaneously, additional validation is performed on the real-time input data, focusing on its completeness, authenticity, and timeliness (e.g., ensuring the input data is the latest collected data to avoid prediction bias caused by using outdated data). Input data that fails validation is supplemented or corrected before being input into the model to prevent input data quality issues from affecting the prediction results.

[0067] Secondly, multi-scale price forecasting is performed. Based on practical application needs, the model's multi-scale forecasting function is activated: short-term forecasting (daily price forecasting for 1-7 days), calling the short-term fluctuation forecasting branch to output the daily lithium carbonate price forecast and the forecast deviation range (the deviation range is calculated based on the daily forecast error during model training and set to ±3%); medium- to long-term forecasting (price forecasting for 1-3 months), calling the long-term trend forecasting branch to output the monthly average lithium carbonate price forecast and the forecast deviation range (the deviation range is set to ±5%); simultaneously, the confidence level of the forecast results is output (calculated based on the prediction accuracy during model training, with a confidence level range of [0,1], the higher the confidence level, the more reliable the forecast result), providing a reference for subsequent verification and decision-making. For example, if the model predicts a lithium carbonate price of 200,000 yuan / ton on a certain day, with a deviation range of ±3% and a confidence level of 0.85, it indicates that the price on that day is likely between 194,000 and 206,000 yuan / ton, and the forecast result is relatively reliable.

[0068] Then, a preliminary verification mechanism for the prediction results is established to make a preliminary judgment on the rationality and accuracy of the prediction results. The design employs a three-tiered verification system, progressively strengthening each tier to ensure the reasonableness of the prediction results: The first tier is numerical range verification. Based on the historical price range and reasonable cost range of the lithium carbonate market, a reasonable range for the prediction results is set (e.g., the reasonable price range for industrial-grade lithium carbonate is 120,000-300,000 RMB / ton, and the reasonable price range for battery-grade lithium carbonate is 150,000-350,000 RMB / ton). If the prediction result exceeds this range, it is marked as an unreasonable prediction result, and the reason for the anomaly is recorded (e.g., abnormal input data, the model not yet adapted to new market changes). The second tier is trend consistency verification. The trend of the prediction result is compared with the recent actual market trend (e.g., the recent market shows a fluctuating upward trend, while the prediction result shows a sharp downward trend, without clear unforeseen factors to support it). If the trends are inconsistent, it is marked as a suspicious prediction result, and further verification is conducted in conjunction with short-term unforeseen factors in the input data (e.g., public opinion, policies). The third tier is deviation reasonableness verification. The deviation between the prediction result and recent historical prices is calculated (e.g., the deviation between the predicted value and the average price of the past 7 days). If the deviation exceeds 10%, and there is no clear unforeseen market factor to support it (e.g., major policy adjustments, major safety accidents), it is marked as an unreasonable prediction result, requiring further optimization.

[0069] Finally, the preliminary verification results are categorized. Validated results (values ​​within a reasonable range, consistent trend, reasonable deviation, confidence level ≥ 0.7) are compiled into a preliminary forecast report, clearly specifying the forecast time, forecast price, deviation range, confidence level, and core influencing factors (e.g., the predicted price increase is mainly due to the increase in power battery installations). For forecasts marked as suspicious, the input data and actual market conditions are further verified. Relevant data is supplemented and the data is re-entered into the model for prediction. If the re-prediction result is still suspicious, it is marked as a forecast to be optimized. For forecasts marked as unreasonable, they are directly discarded, and the reasons for the unreasonableness are analyzed (e.g., incorrect input data, unsuitable model hyperparameters), and timely corrections are made to provide a basis for subsequent forecast optimization steps. The core working principle of this step is that the reasonableness of the model's forecast results directly determines its practical application value. Multi-scale prediction meets the forecasting needs of different scenarios. Establishing a hierarchical preliminary verification mechanism can quickly eliminate obviously unreasonable forecast results, reduce the workload of subsequent optimization, and provide a clear direction for forecast optimization, ensuring the basic accuracy and reasonableness of the forecast results.

[0070] Step 10: Prediction Result Optimization and Dynamic Correction This step builds upon the preliminary verification results from step 9. Its core objective is to address any questionable predictions or model biases identified during the preliminary verification. By incorporating real-time market data and unforeseen factors, a mechanism for optimizing and dynamically correcting prediction results is established to rectify these biases and improve the accuracy and reliability of the predictions. Simultaneously, it addresses the lack of a dynamic adaptive adjustment mechanism in existing methods, ensuring that the predictions adapt to dynamic changes in the market environment. The operation and working principle of this step are as follows: First, we analyze the sources of prediction bias and clarify the direction for optimization. By comparing the preliminary prediction results with recent historical actual prices and real-time market data, and combining error analysis during model training, we summarize the main sources of prediction bias into four categories: First, input data bias, i.e., the real-time input multimodal data contains slight noise, lag, or missing key data, leading to model prediction bias; second, model parameter bias, i.e., insufficient adaptation of hyperparameters during model training, or differences in the distribution of model training data and real-time market data, resulting in a decrease in the model's generalization ability and prediction bias; third, feature fusion bias, i.e., in the multimodal feature fusion process in step 7, the unreasonable allocation of feature weights leads to the insufficient reflection of the impact of some key features (such as short-term sudden public opinion features), resulting in prediction bias; fourth, omission of sudden market factors, i.e., sudden market factors that occur in real time (such as sudden policy adjustments, major safety accidents) are not included in the input data in a timely manner, causing the model prediction results to fail to reflect the latest market changes.

[0071] Secondly, personalized optimization strategies are designed to address prediction biases from different sources. The first type is prediction bias caused by input data deviation: real-time input data is re-verified and supplemented. The noise filtering and outlier correction methods from step 3 are used to remove slightly noisy data and supplement missing key data (such as missing lithium mine production data). Simultaneously, a data lag correction mechanism is introduced to adjust the input data based on the data lag time (e.g., macroeconomic data lags by one month) to reduce bias caused by data lag. The second category is prediction bias caused by model parameter bias: an online fine-tuning strategy is adopted, using the latest real-time collected data (multimodal data and actual price data from the past month) to fine-tune the key hyperparameters of the model (such as the Adam optimizer learning rate, the number of Transformer attention heads, and the number of LSTM hidden layer units). Each fine-tuning only updates a small number of parameters (no more than 10% of the total parameters) to avoid model performance fluctuations. At the same time, a transfer learning method is adopted, using the trained model parameters as initial parameters, and training with new real-time data for a small number of rounds (5-10 epochs) to enable the model to quickly adapt to the new market data distribution and reduce prediction bias. The third category is prediction bias caused by feature fusion bias: recalculate the weights of multimodal features (using the weight allocation method based on information entropy and correlation coefficients in step 7), and adjust the weights of each modality feature in conjunction with real-time market conditions—for example, increase the modal weights of unstructured features when short-term sudden public opinion events have a significant impact; increase the modal weights of structured and semi-structured features when medium- and long-term supply and demand patterns change significantly; at the same time, re-optimize the fusion feature set, remove redundant features, and add new features that have a significant impact on the current market (such as quantitative features corresponding to sudden policies) to improve feature quality. The fourth category is prediction bias caused by the omission of sudden market factors: establish a real-time monitoring mechanism for sudden factors to capture sudden market factors (such as policy adjustments, safety accidents, and supply chain disruptions) in real time. Use the semi-structured data quantification method in step 6 to quickly quantify sudden factors into feature vectors, add them to the fusion feature set, re-input them into the model for prediction, correct the prediction results, and ensure that the prediction results reflect the latest market changes.

[0072] Then, a dynamic correction mechanism for the prediction results is established to achieve real-time updates and optimization of the prediction results. The first step is to set the dynamic correction time interval, with different intervals set according to the prediction scale: short-term predictions (1-7 days) are corrected once daily, combining the latest collected multimodal data and actual market prices to correct the prediction results for the following days; medium- and long-term predictions (1-3 months) are corrected twice monthly, combining the latest supply and demand data, public opinion data, and policy data for the current month to correct the prediction results for the following months. The second step is to establish a prediction error feedback mechanism, calculating the error (MSE, MAPE) between each prediction result and the actual price, recording the error magnitude and source, and feeding the error back to the model training and feature fusion steps as the basis for online model fine-tuning and feature weight adjustment—for example, if the prediction error corresponding to a certain type of feature remains large, the weight of that type of feature is reduced, and the feature fusion strategy is re-optimized; if the prediction error of a certain branch of the model remains large, the structure and parameters of that branch are fine-tuned. The third step involves introducing expert experience correction. Lithium carbonate industry experts are invited to evaluate the optimized forecast results. Combining the experts' judgments on market trends and industry experience, the forecast results are further corrected (e.g., if the experts judge that the impact of a sudden policy on prices is greater than the model's prediction, the forecast results are adjusted appropriately) to improve the practicality and reliability of the forecast results.

[0073] Finally, the optimized final forecast results are output, and a forecast report is compiled. The optimized forecast results (short-term daily forecasts, medium-to-long-term monthly forecasts), forecast deviation range, confidence level, core influencing factors, and optimization process are compiled into a standardized forecast report, clearly stating the forecast conclusions and application recommendations (e.g., suggesting companies adjust inventory based on short-term forecast results and formulate production plans based on medium-to-long-term forecast results). Simultaneously, the optimization process and error details of the forecast results are recorded to provide data support for subsequent improvements to the forecasting method and model optimization. The core working principle of this step is: dynamic changes in the market environment lead to deviations in model forecast results. By analyzing the sources of deviation, personalized optimization strategies are designed, and a dynamic correction mechanism is established to achieve real-time updates and optimization of forecast results. Combined with online fine-tuning and expert experience, forecast accuracy is further improved, addressing the problem that existing methods lack a dynamic adaptive adjustment mechanism and cannot adapt to market changes.

[0074] Step 11: Interpretability analysis and visualization of prediction results This step, based on the final prediction results optimized in step 10, aims to address the "black box" problem of existing prediction models. Through interpretability analysis, it clarifies the impact of each multimodal feature and influencing factor on the prediction results. Simultaneously, it presents the prediction results, influencing factors, and interpretability analysis conclusions in a visual manner, improving the readability and credibility of the prediction results and facilitating understanding and application by enterprise decision-makers. The operation and working principle of this step are as follows: First, an interpretability analysis of the forecast results is conducted, employing a combination of multiple interpretability methods to comprehensively analyze the formation process and influencing factors of the forecast results. First, feature importance analysis: using the SHAP (SHapley Additive ex Planations) method, the SHAP value of each feature in the fused feature set is calculated for its impact on the forecast result. The larger the absolute value of the SHAP value, the greater the influence of that feature on the forecast result. Simultaneously, the direction of influence of the features is distinguished (a positive SHAP value indicates that the feature drives up prices; a negative SHAP value indicates that the feature drives down prices). For example, if the SHAP value of the power battery installation growth rate is 0.25, and the SHAP value of the lithium ore supply-demand gap is -0.20, it indicates that the power battery installation growth rate is the main factor driving up prices, while the lithium ore supply-demand gap is the main factor inhibiting price increases. Second, modal contribution analysis: The contribution of structured, unstructured, and semi-structured modal features to the prediction results is calculated (based on a comprehensive calculation of modal weights and feature importance), clarifying the role of each modal data in the prediction process. For example, in medium- and long-term predictions, the structured modal contribution is 45%, the semi-structured modal contribution is 35%, and the unstructured modal contribution is 20%, indicating that medium- and long-term predictions mainly rely on supply and demand fundamentals and market expectations. In short-term predictions, the unstructured modal contribution is 40%, the structured modal contribution is 35%, and the semi-structured modal contribution is 25%, indicating that short-term predictions mainly rely on short-term unforeseen factors. Third, branch contribution analysis: In the dual-branch collaborative prediction model, the contribution of the long-term trend prediction branch and the short-term fluctuation prediction branch to the final prediction result is calculated, and the rationality of the branch contribution is analyzed in conjunction with the market environment. For example, when the market is in a stable trend, the long-term branch contribution is 70%, and the short-term branch contribution is 30%; when the market is in a period of sharp fluctuations, the short-term branch contribution is 65%, and the long-term branch contribution is 35%. Fourth, interpretability analysis of abnormal predictions: For abnormal prediction results with large deviations during the prediction process, the causes of the anomalies are analyzed by combining feature importance and modal contribution (e.g., a large deviation in a short-term prediction may be due to the failure to include sudden public opinion in the input data in a timely manner, resulting in a low contribution of unstructured features), clarifying the formation mechanism of abnormal predictions and providing direction for subsequent prediction optimization.

[0075] Secondly, diverse visualization methods are employed to present prediction results, interpretability analysis conclusions, and related data, enhancing readability. Multi-dimensional visualization charts are designed, balancing professionalism and ease of use: First, a price prediction trend chart, presented as a line graph, with the horizontal axis representing time (daily, monthly) and the vertical axis representing price (ten thousand yuan / ton), simultaneously marking historical actual prices, predicted prices, and prediction deviation ranges, clearly showing historical and future price trends for easy comparative analysis; Second, a feature importance bar chart, with the horizontal axis representing feature names (e.g., power battery installation growth rate, lithium mine supply and demand gap, public opinion sentiment score) and the vertical axis representing the absolute value of the SHAP value, intuitively displaying the degree of influence of each feature on the prediction results and indicating the direction of influence; Third, a modality and branch contribution pie chart, using pie charts to present the contribution of the three major modal features and the contribution of the two branches. The system offers several key features: First, a comprehensive data analysis platform that clearly demonstrates the role of various data types and branches in forecasting. Second, a multimodal data correlation heatmap that presents the correlation between different modal features and the correlation between each feature and price, visually showcasing the inherent relationships between features. Third, a forecasting error analysis histogram that presents the distribution of forecasting error (MAPE), demonstrating the model's forecasting accuracy and highlighting time points with significant errors and their causes. Fourth, a visualized interpretability analysis report that organizes the core conclusions of interpretability analysis (such as main influencing factors, forecasting logic, and causes of anomalies) into a report combining text and graphics, using flowcharts to illustrate the 12 steps of the forecasting method and the core outputs of each step, facilitating decision-makers' rapid understanding of the forecasting logic and analysis conclusions.

[0076] Next, the visualization reports were organized, clarifying the target audience and presentation format. Two types of visualization reports were designed based on the different audiences: one is a professional technical report, aimed at technical personnel, which details the 12 steps of the forecasting method, model parameters, feature extraction process, interpretability analysis details, error analysis results, and model optimization suggestions, meeting the needs of technical personnel for method optimization and model improvement; the other is a decision-making application report, aimed at enterprise decision-makers, which simplifies technical details and focuses on presenting forecast results, core influencing factors, forecast confidence levels, and application suggestions (such as inventory adjustment, production planning, and cost control suggestions), using concise and clear visualization charts and text descriptions to facilitate decision-makers' quick understanding and application of the forecast results. Simultaneously, interactive functions were provided for the visualization reports, allowing users to click on charts to view detailed data (e.g., clicking on a feature importance bar chart to view the specific value and impact analysis of that feature), enhancing the report's usability.

[0077] Finally, an update mechanism for the visualization reports is established to ensure their timeliness. Based on the dynamic correction time intervals specified in Step 10, the visualization reports are updated synchronously: short-term forecast visualization reports are updated daily, and medium- to long-term forecast visualization reports are updated monthly, supplementing the latest forecast results, actual price data, interpretability analysis conclusions, and error analysis results. Simultaneously, when significant changes occur in the market environment (such as major policy adjustments or sudden shifts in supply and demand patterns), the visualization reports are updated promptly, adding impact analysis of these significant changes to ensure the reports reflect the latest market conditions and forecast results. The core working principle of this step is that the interpretability and readability of the forecast results directly determine their practical application value. By employing various interpretability methods, the "black box model" problem is solved, making the formation process of the forecast results traceable and the influencing factors clear. Through diversified visualization presentations, the understanding cost for decision-makers is reduced, enhancing the credibility and practicality of the forecast results, and addressing the problems of poor interpretability and limited practicality of existing methods.

[0078] Step 12: Method Validation and Continuous Improvement This step serves as the concluding and enhancement step of the entire forecasting method. Its core purpose is to establish a scientific effectiveness verification mechanism, comprehensively evaluate the method's forecasting accuracy, generalization ability, and practicality, compare it with existing forecasting methods, clarify its advantages and disadvantages, and establish a continuous improvement mechanism. Based on verification results and market changes, the forecasting method will be continuously optimized to improve its adaptability and forecasting accuracy, ensuring that the method can meet the actual needs of the industry in the long term. The operation and working principle of this step are as follows: First, clearly define the evaluation metrics and dataset for validity verification to ensure the scientific rigor and objectivity of the verification results. Four core evaluation metrics are established to comprehensively cover prediction accuracy, generalization ability, and practicality: First, prediction accuracy metrics, including mean squared error (MSE), mean absolute percentage error (MAPE), and coefficient of determination (R²). The smaller the MSE and MAPE, and the closer R² is to 1, the higher the prediction accuracy—for short-term forecasts (daily), MAPE ≤ 5% and R² ≥ 0.9 are required; for medium- to long-term forecasts (monthly), MAPE ≤ 8% and R² ≥ 0.85 are required. Second, generalization ability metrics, including cross-validation error (5-fold cross-validation) and the difference in prediction accuracy under different market scenarios. The evaluation criteria are as follows: First, the smaller the cross-validation error, the smaller the difference in accuracy across different scenarios, indicating stronger generalization ability. Second, practicality indicators include prediction efficiency (prediction time per batch ≤ 10 seconds), interpretability score (using expert scoring method, full score 10 points, requiring a score ≥ 8 points), and dynamic adaptability (prediction accuracy decreases by ≤ 10% after market changes). Third, comparative indicators compare this method with existing mainstream prediction methods (such as ARIMA model, LSTM model, and simple multi-source data fusion model), calculate the evaluation indicators of each method, and clarify the advantages of this method. The evaluation dataset uses the test set divided in step 8, and is supplemented with real-time market data from the past 6 months (data not used in model training) to ensure that the evaluation dataset can cover different market scenarios (rising, falling, and fluctuating), avoiding the one-sidedness of the validation results.

[0079] Secondly, the effectiveness of the method is validated in three stages: The first stage is basic accuracy validation. Using test set data, the accuracy metrics of the method, such as MSE, MAPE, and R², are calculated to verify the method's basic predictive ability and determine whether it meets the preset accuracy requirements. If not, the reasons are analyzed (e.g., unsuitable model parameters, unreasonable feature fusion), and the corresponding steps are returned for optimization. The second stage is generalization ability validation. Five-fold cross-validation is used, randomly dividing the evaluation dataset into five parts, with four parts used as the training set and one part as the test set. Validation is repeated five times, and the cross-validation error is calculated. Simultaneously, the evaluation dataset is divided according to market scenarios (rising, falling, oscillating), and the prediction accuracy under different scenarios is calculated to verify the method's adaptability to different market environments. If the generalization ability is insufficient, the model structure (e.g., adding data augmentation, adjusting hyperparameters) and feature fusion strategies are optimized to improve generalization ability. The third stage involves comparative validation, testing this method against existing mainstream prediction methods on the same evaluation dataset. The accuracy, generalization ability, and usability metrics of each method are compared to clarify the advantages of this method—for example, its MAPE is 3-5% lower than the LSTM model, its interpretability score is 2-3 points higher than the simple multi-source data fusion model, and its prediction efficiency meets the preset requirements, thus validating the method's advancement and practicality. Simultaneously, industry experts and enterprise users are invited to evaluate the method's practicality, and user feedback (such as the reasonableness of the prediction results, the ease of use of the visualization report, and the targeted nature of application suggestions) is collected as a basis for method improvement.

[0080] Then, establish a mechanism for continuous improvement of the method, and continuously optimize the prediction method based on the results of effectiveness verification and market changes. The first step is to establish a problem collection and analysis mechanism, regularly collect problems encountered during the application of the method (such as decreased prediction accuracy, low prediction efficiency, and insufficient interpretability), deficiencies found in effectiveness verification, and suggestions from user feedback, classify and organize the problems, analyze the sources of the problems (such as decreased accuracy due to changes in the market landscape, and low prediction efficiency due to excessive feature dimensions), and clarify the direction for improvement. The second step is to develop targeted improvement plans, addressing the relevant steps of the optimization method based on the source of the problem: If the problem stems from data quality, optimize the data collection, standardization, and denoising processes in steps 1-3, expand data sources, and improve data quality; if the problem stems from feature fusion, optimize the feature weight allocation strategy and feature optimization method in step 7, and supplement with new features; if the problem stems from model performance, optimize the model structure and training parameters in step 8, introduce new deep learning techniques (such as the Transformer-LSTM hybrid model), and improve model performance; if the problem stems from interpretability, optimize the interpretability analysis method in step 11, and increase the richness of visualization charts; if the problem stems from dynamic adaptability, optimize the dynamic correction mechanism in step 10, shorten the correction interval, and improve the dynamic adaptability of the method. The third step is to establish an improvement effect verification mechanism. For each improved method, use an evaluation dataset and real-time data for verification, calculate evaluation indicators, compare the effects before and after the improvement, and ensure that the improvement can enhance the method's performance; if the improvement effect does not meet expectations, adjust the improvement plan and re-optimize until the preset requirements are met. The fourth step is to establish a method update mechanism, and to conduct a comprehensive update of the prediction method every quarter, integrating the latest technological advancements (such as deep learning models and feature extraction techniques), market data changes, and user feedback to continuously improve the method's sophistication and practicality. At the same time, a method version management mechanism should be established to record the content, improvement direction, and effect of each update, so as to facilitate traceability and subsequent optimization.

[0081] Finally, an effectiveness verification report and continuous improvement plan are compiled to form a closed-loop management system. The effectiveness verification process, evaluation indicators, comparison results, and identified shortcomings are summarized in an effectiveness verification report, clarifying the method's strengths and areas for improvement. Based on identified shortcomings and user feedback, a medium- to long-term continuous improvement plan is developed, clearly defining the responsible person, completion time, improvement goals, and verification standards for each improvement task. Simultaneously, the effectiveness verification report and continuous improvement plan are shared with technical personnel and enterprise users to solicit further suggestions, forming a closed-loop management system of "verification - problem identification - improvement - re-verification." This ensures that the method can adapt to the dynamic changes in the lithium carbonate market in the long term, continuously providing enterprises with accurate and reliable price forecasting services. The core working principle of this step is: through scientific effectiveness verification, the performance of the method is comprehensively evaluated, and its strengths and weaknesses are identified; through a continuous improvement mechanism, existing problems are addressed, the latest technologies and market data are integrated, and the prediction accuracy, generalization ability, and practicality of the method are continuously improved, forming a closed-loop optimization. This ensures that the method can meet the actual needs of the industry in the long term, solving the problem that existing methods lack a continuous optimization mechanism and are difficult to adapt to long-term market changes.

[0082] Example 2: This embodiment uses January 1, 2024 to December 31, 2024 as the implementation period, and focuses on battery-grade lithium carbonate (a core raw material for new energy vehicle power batteries, whose price fluctuations are more representative and a key category of industry focus) as the prediction object. It aims to achieve accurate predictions of its daily price for 1-7 days and its monthly price for 1-3 months, verifying the feasibility and effectiveness of the method. The implementation scenario is set as a price decision support system for a lithium carbonate production company. The core requirement is to assist the company in formulating monthly production plans, optimizing inventory management (avoiding inventory backlog or shortages), and reducing operational risks caused by price fluctuations through accurate predictions. This embodiment strictly follows the 12 steps outlined above, clearly defining the specific operations, parameter settings, data details, and actual output results for each step, as follows: I. Preliminary Preparations 1. Hardware environment: Server (CPU: Intel Xeon E5-2690 v4, GPU: NVIDIA Tesla V100, Memory: 64GB, Hard disk: 1TB) for data storage, model training and prediction calculation; 3 terminal devices for data collection and verification, prediction result viewing and report export.

[0083] 2. Software Environment: Operating System (Ubuntu 20.04 LTS), Programming Language (Python 3.9), Core Libraries (Pandas 1.5.3, NumPy 1.24.3, Scikit-learn 1.2.2, TensorFlow 2.10.0, PyTorch 1.13.1, Jieba 0.42.1, OpenCV 4.7.0), Data Collection Tools (Scrapy, API Interface Tools), Visualization Tools (Matplotlib 3.7.1, Seaborn 0.12.2), which are used to implement data processing, model construction, feature extraction and visual presentation.

[0084] 3. Data Source Confirmation: Connect with authoritative data sources in advance to ensure stable data collection. The specific data sources are as follows: Structured Data (Baichuan Information, Lonzhong Information), Semi-Structured Data (CITIC Securities, Orient Securities industry reports, Tianqi Lithium / Ganfeng Lithium corporate announcements), Unstructured Data (Baidu public opinion, industry media news, enterprise press conference recordings, images of lithium mining sites).

[0085] 4. Presetting of Basic Parameters: Set some core parameters in advance to provide support for subsequent steps, including: standardization threshold, noise filtering parameters, feature dimensions, model training parameters, prediction deviation range, etc. (specific parameters are described in detail in each step).

[0086] II. Specific Implementation Steps and Results Step 1: Determination and Data Collection of Multimodal Data Sources Combined with the requirements of the implementation scenario, determine three major categories of multimodal data sources, and adopt the combination of "automatic collection + manual verification" to complete the full-scale data collection from January 1, 2024 to December 31, 2024, specifically as follows: 1. Structured Data Collection: Collect historical price data for the past 10 years (2014 - 2023) (daily, weekly, monthly prices of battery-grade lithium carbonate, accurate to yuan / ton), annual lithium ore supply and demand data in 2024 (global and Chinese lithium ore production, inventory, import volume, export volume, collected monthly), battery industry data (new energy vehicle production, power battery installation volume, battery recycling volume, collected monthly), macroeconomic data (GDP, PPI, RMB exchange rate, comprehensive commodity index, collected quarterly), and policy quantification data (new energy subsidy intensity, number of mineral exploration approvals, collected quarterly). Automatically collect data from Baichuan Information, Lonzhong Information and the official websites of government statistical departments through API interfaces. Collect daily data at 0:00 every day, monthly data on the last day of each month, and quarterly data on the last day of each quarter; for policy quantification data that cannot be collected through API, manual entry is used, and after entry, it is cross-checked by 2 staff members.

[0087] 2. Semi-structured data collection: Using Scrapy web crawling technology, we selectively crawled lithium carbonate industry research reports (at least two per month) published by CITIC Securities and Orient Securities, as well as company announcements (quarterly announcements and capacity announcements) published by Tianqi Lithium and Ganfeng Lithium. Key information was extracted using PDF and HTML parsing technologies. An update monitoring mechanism was established to complete data collection within 24 hours of report / announcement release. In 2024, a total of 26 industry research reports and 12 company announcements were collected, with no key announcements missed.

[0088] 3. Unstructured Data Collection: News and public opinion data (collected daily, keywords: "battery-grade lithium carbonate price", "lithium mining", "power battery installation", "lithium mine safety accident", etc.), collected through Baidu's public opinion API interface, collecting the day's news at 6 PM daily, filtering advertisements and irrelevant information; Image data (collected once a week, images of lithium mining sites and lithium carbonate production lines), obtained through industry cooperation channels, selecting clear and effective images, collecting 5-8 images per week; Audio data (collected once a month, recordings of industry seminars and company press conferences), processed using professional tools after recording to ensure no background noise, collecting 12 audio data entries throughout 2024.

[0089] 4. Data Collection and Verification: Targeted verification rules were formulated. For structured data, the focus was on verifying the numerical range (a reasonable price range for battery-grade lithium carbonate is 150,000-350,000 RMB / ton; data exceeding this range was marked as abnormal) and data integrity. For semi-structured data, the reliability of the source was verified, and reports from non-authoritative institutions were removed. For unstructured data, the authenticity and relevance were verified, and fake news and blurry images were removed. A total of 32 abnormal data entries were verified during this collection (18 structured data entries, 4 semi-structured data entries, and 10 unstructured data entries). After supplementary collection and correction of 25 entries, and removal of 7 invalid data entries, the following valid data were obtained: 1260 structured data entries, 34 semi-structured data entries (reports + announcements), and 482 unstructured data entries (452 ​​news articles, 20 images, and 10 audio recordings), ensuring that the data was comprehensive, authentic, and valid.

[0090] Step 2: Multimodal data classification and standardization The collected valid data was categorized and organized, and a personalized standardization method was used to eliminate the influence of format differences and units. The specific operations and results are as follows: 1. Data Classification and Organization: Establish three major database categories. The structured database is subdivided and stored according to "price data," "supply and demand data," "macroeconomic data," and "policy quantitative data," and sorted by time. The semi-structured database is subdivided and stored according to "industry research reports" and "company announcements," and sorted by publication time and issuing institution. The unstructured database is subdivided and stored according to "news and public opinion data," "image data," and "voice data." Each data entry is given a unique identifier (e.g., "structured-price data-20240101-Baichuan Information") for easy traceability.

[0091] 2. Structured Data Standardization: The Shapiro-Wilk test was used to test the normality of each structured indicator (α=0.05). Monthly prices of battery-grade lithium carbonate and monthly installed capacity of power batteries followed a normal distribution and were standardized using Z-score (converted to a mean of 0 and a standard deviation of 1). Policy quantitative scores and lithium ore inventory did not follow a normal distribution and were standardized using Min-Max (mapped to the [0,1] interval). Missing values ​​(such as the missing lithium ore import volume in March 2024) were supplemented using linear interpolation. After supplementation, the data trend was consistent with adjacent months, with no significant deviation. After standardization, all structured data had unified dimensions and could be directly used for subsequent feature extraction.

[0092] 3. Semi-structured data standardization: For industry research reports, extract key information such as supply and demand forecasts and price trend judgments, and convert numerical information into structured data (e.g., "Estimated production of battery-grade lithium carbonate in Q4 2024 is 80,000 tons" is converted into "2024Q4, battery-grade lithium carbonate production forecast, 80,000 tons"). Remove punctuation and redundant words from text information and unify it into UTF-8 encoding. For company announcements, extract information such as production capacity and capacity adjustment, and similarly complete the numerical-to-structure and text standardization processing to ensure that it can be processed collaboratively with other data.

[0093] 4. Unstructured Data Standardization: News and public opinion text data were segmented using Jieba, stop words were removed, and the data was standardized to UTF-8 encoding and lowercase format, with special symbols removed. Image data was standardized to a size of 224×224 pixels and RGB color space, with Gaussian filtering to remove noise and normalization processing (pixel values ​​mapped to [0,1]). Audio data was standardized to a sampling rate of 16kHz and a bit depth of 16 bits, with Wiener filtering to remove noise, and processed in frames (frame length 25ms, frame shift 10ms) to convert it into standardized audio frames. After standardization, the unstructured data format is unified and can be used for subsequent feature extraction.

[0094] 5. Standardization Validation: The data format uniformity, completeness, and rationality were checked. A total of 8 non-standardized data were found (all of which were unstructured text data with encoding errors). After re-standardization, the data met the requirements, and the final output was a standardized multimodal data set.

[0095] Step 3: Personalized noise filtering and outlier correction for multimodal data Personalized processing strategies were designed to address the noise characteristics of different modalities of data, further improving data quality. The specific operations and results are as follows: 1. Structured Data: A dual strategy of "noise filtering + outlier correction" is employed. Noise Filtering: A 7-day sliding window is set for daily price data, and a 3-month sliding window is set for monthly production data. The mean and standard deviation are calculated. Data deviating from the mean by more than twice the standard deviation is marked as suspected noise. After verification, the window mean is used to replace the noise data, filtering a total of 12 noisy data points. For data on the same indicator from different sources, a weighted average is used for fusion (0.7 weight for authoritative data sources, 0.3 weight for ordinary data sources) to eliminate data source bias. Outlier Correction: The Grubbs test (α=0.05) is used to identify 6 outliers. One of these is a reasonable anomaly caused by policy adjustments (a 12% single-day increase in battery-grade lithium carbonate prices in June 2024 due to increased new energy subsidies), marked as a sudden outlier. The other 5 are data entry errors (such as errors in lithium mine production data entry), corrected using a quadratic polynomial fitting. The corrected data conforms to market trends and is unbiased.

[0096] 2. Semi-structured data: A strategy of "source verification + information filtering + contradiction correction" was adopted. Source verification: Two reports from non-authoritative institutions (published by an unknown small brokerage firm, with false data) were removed; Information filtering: Key paragraphs related to prices were selected, irrelevant background information was removed, and core content was retained; Contradiction correction: Two reports were found to have contradictory production forecasts for Q3 2024 (78,000 tons and 85,000 tons respectively). Combining the announcements of Tianqi Lithium and Ganfeng Lithium and the actual market situation, a reasonable forecast value of 82,000 tons was determined, and the contradictory data was corrected, ultimately obtaining 38 pieces of valid semi-structured key information.

[0097] 3. Unstructured data: News and public opinion texts were processed by using cosine similarity (threshold 0.8) to remove 32 duplicate news items, using a pre-trained BERT model to identify 8 fake news items (such as fabricated news about lithium mine safety accidents), removing 5 extremely false public opinion items, and retaining 407 valid news items; Image data were processed by using Gaussian filtering to remove noise points, Canny algorithm to extract core regions, CNN model to identify 3 distorted images, removing 2 images with a quality score <0.7, and retaining 15 valid images; Speech data were processed by Wiener filtering to remove background noise, MFCC feature identification to identify 2 abnormal speech items (irrelevant statements), and after removal, 8 valid speech items were retained.

[0098] 4. Overall verification: Calculate the noise removal rate and outlier correction rate for various types of data. The noise removal rate is 92% and the outlier correction rate is 96%, both meeting the preset requirements. Check the data integrity and find no data loss due to over-filtering. Output high-quality multimodal data after noise filtering and outlier correction.

[0099] Step 4: Extraction of trend and correlation features from structured data The process of extracting trend and correlation features from structured data and filtering for highly correlated features is as follows: 1. Trend Feature Extraction: Based on historical price data of battery-grade lithium carbonate, long-term trend features were extracted using MA (Moving Average) and EMA (Exponential Easing). Short-term MAs of 7 days and 30 days, and long-term MAs of 90 days and 180 days were calculated, with an EMA smoothing coefficient α=0.3. Linear regression was used to fit the price time series, and regression coefficients were calculated to quantify trend strength (e.g., the regression coefficient for January-March 2024 was 0.18, indicating a moderate upward trend). Standard deviation and variance were used to extract short-term volatility features, calculating daily and weekly volatility, and extracting extreme value features (highest price, lowest price, closing price) and skewness coefficients (skewness of 0.23 for the whole of 2024, indicating right-skewed volatility, with upward volatility greater than downward volatility). Fourier transform was used to decompose the price time series and extract periodic features, revealing a clear annual cycle in battery-grade lithium carbonate prices, closely related to increased mining difficulty in winter (November-February) and increased demand in summer (June-August).

[0100] 2. Feature Extraction: Correlation features were extracted from lithium ore supply and demand data, battery industry data, etc., including: lithium ore supply and demand gap, inventory turnover rate, import dependence, power battery installation growth rate, GDP growth rate, and policy support strength score, totaling 18 features. A combination of Pearson and Spearman correlation coefficients was used to analyze the correlation between each feature and battery-grade lithium carbonate prices (α=0.05). Eight highly correlated features with an absolute correlation coefficient > 0.5 were selected: lithium ore supply and demand gap (r=-0.75), power battery installation growth rate (r=0.83), lithium ore import dependence (r=-0.62), policy support strength score (r=0.58), battery recycling rate (r=-0.53), PPI growth rate (r=0.56), lithium ore inventory (r=-0.67), and new energy vehicle production growth rate (r=0.71). Ten low-correlation features (such as GDP growth rate r=0.32) were removed to avoid feature redundancy.

[0101] 3. Feature normalization: Min-Max standardization is adopted to map the extracted trend features (6 features) and highly correlated features (8 features), a total of 14 features, to the [0,1] interval to establish a structured feature set with 14 dimensions. The correlation coefficient of each feature is recorded to provide a basis for subsequent weight allocation.

[0102] Step 5: Semantic and Visual / Speech Feature Extraction from Unstructured Data Deep learning technology is used to transform unstructured data into quantifiable feature vectors. The specific operations and results are as follows: 1. News and Public Opinion Text Feature Extraction: For 407 valid news texts, the Jieba word segmentation + BiLSTM-CRF entity recognition model was used to identify core entities (lithium mine, lithium carbonate, power battery, etc.) and label the entity types. A fine-tuned BERT model for the lithium carbonate industry was constructed (using 1000 news articles from the lithium carbonate industry to fine-tune the general BERT model). The news texts were input into the model, and the output vector of the last layer was extracted as the semantic feature vector (768 dimensions). An LSTM sentiment analysis model was used to score the sentiment of the news texts (in the [0,1] interval) and extract sentiment features. Principal component analysis (PCA) was used to reduce the 768-dimensional semantic feature vector to 128 dimensions, which was then concatenated with the sentiment features (1 dimension) to obtain a comprehensive text feature vector with a dimension of 129. A total of 407 text feature vectors were extracted, covering short-term sudden factors (such as the news about the new energy subsidy policy in June 2024, with a sentiment score of 0.89, which drove up prices).

[0103] 2. Image Feature Extraction: For 15 valid images (7 from lithium mining sites and 8 from production lines), an improved ResNet-50 model was constructed (the last fully connected layer was removed, an adaptive average pooling layer was added, and fine-tuned using 500 images of the lithium carbonate industry chain). Data augmentation was performed on the images (random cropping, flipping, and brightness adjustment), and the images were input into the improved ResNet-50 model. The output vector of the adaptive average pooling layer was extracted as the visual feature vector (2048 dimensions). The t-SNE dimensionality reduction method was used to reduce the 2048 dimensions to 128 dimensions. An image classification model was used to classify the images according to their capacity utilization level (high=1, medium=0.5, low=0), converting them into numerical features. Visual features were then added, resulting in a comprehensive image feature vector with 129 dimensions and 15 image feature vectors (e.g., the lithium mining site image shows a high operating rate, with a capacity utilization feature value of 1.0, reflecting sufficient supply).

[0104] 3. Speech Feature Extraction: For 8 valid speech data points, the data are segmented and windowed (Hanning window). The 13-dimensional MFCC coefficients and first and second-order differences are calculated to obtain 39-dimensional MFCC features. The Wav2Vec2.0 model is used to convert the speech into text. Then, the BERT model is fine-tuned to extract the text semantic feature vector (dimensionality reduced to 128). The MFCC features (39-dimensional) and the semantic feature vector (128-dimensional) are concatenated to obtain a 167-dimensional comprehensive speech feature vector. PCA is used to reduce the dimensionality to 128, resulting in a final speech feature vector with 128 dimensions and a total of 8 speech feature vectors (e.g., when a company announces capacity expansion at a press conference, the semantic features reflect increased supply, driving down prices).

[0105] 4. Feature validity test: ANOVA was used to test the extracted unstructured features. Three invalid features (variance < 0.01) were removed, and valid features were retained. All features were normalized to establish an unstructured feature set with 128 dimensions (integrating text, image, and speech features, using feature concatenation and redundancy removal, and unifying the dimensions).

[0106] Step 6: Extraction of key information and feature quantization from semi-structured data Key information is extracted from semi-structured data and quantified into feature vectors. The specific operations and results are as follows: 1. Key Information Extraction: From 24 valid industry research reports and 12 company announcements, three main categories of key information were extracted: supply and demand forecasts, capacity adjustment information, and analysis of core influencing factors. Keyword and regular expression methods were used to locate key paragraphs. A CNN-BiLSTM paragraph classification model was then used to filter key paragraphs and remove irrelevant content. A fine-tuned BERT model was used to parse key paragraphs and extract specific information. For example, from the reports, the projected Q4 2024 production volume was 82,000 tons, demand was 85,000 tons, and prices were projected to rise by 5%-10%. From the company announcements, Tianqi Lithium's Q2 2024 capacity expansion was 20% (from 12,000 tons / month to 14,400 tons / month), and Ganfeng Lithium's Q3 2024 capacity utilization rate was 92%.

[0107] 2. Key Information Quantification: According to preset rules, the extracted key information is quantified into features: In supply and demand forecast information, the predicted output of 82,000 tons and demand of 85,000 tons are converted into numerical features; the price forecast of a 5%-10% increase is converted into a price forecast index of 1.075; In capacity adjustment information, a 20% capacity expansion is converted into 0.2, a capacity utilization rate of 92% is converted into 0.92, and the capacity adjustment time is converted into the time from the current time (e.g., if the capacity expansion is in Q2 and the announcement is released in April, it is converted into 2 months); In the core influencing factor analysis, "the demand for power batteries significantly affects the price" is converted into an impact score of 1.0 and a positive impact (+1), and is encoded using unique heat codes.

[0108] 3. Feature selection and normalization: Using the Pearson correlation coefficient, six semi-structured features with an absolute value of correlation coefficient > 0.4 with the price of battery-grade lithium carbonate were selected, and redundant features were removed. Min-Max normalization was used to map the six features to the [0,1] interval, and they were spliced ​​to form a semi-structured feature set with a dimension of 128 (adapted to the dimensions of other modal features). The credibility score of each feature source was recorded (1.0 for authoritative sources and 0.7 for ordinary sources).

[0109] Step 7: Multimodal Feature Weighted Collaborative Fusion and Feature Optimization A dual weighting strategy of "modal weight + feature weight" is adopted to fuse three types of modal features and optimize feature quality. The specific operation and results are as follows: 1. Modal weight calculation: An information entropy-based method is used to calculate the information entropy of the three modal feature sets. The probability density of each feature is calculated through kernel density estimation and then substituted into the information entropy formula. The calculated values ​​are: structured feature set H(X1) = 0.28, unstructured feature set H(X2) = 0.63, and semi-structured feature set H(X3) = 0.42. Substituting these values ​​into the modal weight formula... The modal weights were calculated as follows: structured W1≈0.43, unstructured W2≈0.22, and semi-structured W3≈0.35. The sum of the weights was 1, which met the requirements and was consistent with expectations (structured features have the largest amount of information and the highest weight).

[0110] 2. Feature Weight Calculation: A method based on correlation coefficients and confidence levels is used to calculate the feature weights within each mode. Taking the structured mode as an example, the power battery installation growth rate (|r|=0.83, C=1.0) and the lithium ore supply-demand gap (|r|=0.75, C=1.0) are substituted into the feature weight formula. The calculated weights are approximately 0.18 for the installed capacity growth rate of power batteries and approximately 0.16 for the supply-demand difference of lithium ore. The sum of the weights of all features within this mode is 1. Similarly, the weights of features within unstructured and semi-structured modes are calculated, and they all meet the requirement that the sum of the weights is 1.

[0111] 3. Multimodal Feature Fusion: First, for each modality's internal features, weighted summation is performed to fuse them into a 128-dimensional modal feature vector (V1, V2, V3); Second, modal weights are used for weighted fusion, and the results are substituted into the formula. This yields an initial fusion feature set with 128 dimensions.

[0112] 4. Feature Optimization: Analysis of variance was used to remove four invalid features with variance < 0.01; mutual information was used to remove twelve redundant features with mutual information values ​​> 0.8; principal component analysis (PCA) was used to reduce the dimensionality of the remaining features to 64 dimensions while retaining core information; finally, normalization was performed to obtain the optimized fusion feature set with 64 dimensions. This feature set covers the core information of three major modalities and can be directly used as model input, improving the efficiency of subsequent model training.

[0113] Step 8: Construction and Training of the Two-Branch Collaborative Prediction Model A dual-branch collaborative prediction model was constructed, and model training and hyperparameter optimization were completed. The specific operations and results are as follows: 1. Model Structure Construction: (1) Long-term trend prediction branch (improved Transformer model): optimize position encoding (sine and cosine position encoding + time feature encoding concatenation), simplify encoder structure, and set the number of attention heads to 8; the input layer concatenates 64-dimensional fusion feature set and structured trend features (6, 6-dimensional), with a total input dimension of 70; the output layer adopts a linear regression layer to output monthly price prediction values ​​for 1-3 months, and adds a Dropout layer (dropout=0.3) to avoid overfitting.

[0114] (2) Short-term fluctuation prediction branch (improved CNN-LSTM hybrid model): 3 convolutional layers (convolutional kernel size 3, 5, 7), max pooling layer (pooling kernel 2), 2 LSTM layers (hidden layer unit number 128, 64), channel attention mechanism (connecting CNN and LSTM); the input layer splices 64-dimensional fusion feature set and unstructured short-term related features (8, 8-dimensional), with a total input dimension of 72; the output layer adopts a linear regression layer to output the daily price prediction value for 1-7 days, and adds a BatchNorm layer to accelerate convergence.

[0115] (3) Collaborative mechanism: Add a collaborative attention module to calculate the mutual information of the features of the two branches and dynamically allocate attention weights (when the short-term fluctuations are large, the weight of the short-term branch is 0.5-0.7; when the market is stable, the weight of the long-term branch is 0.6-0.8); use weighted summation to fuse the prediction results of the two branches to obtain the final prediction value.

[0116] 2. Dataset partitioning: The 64-dimensional fused feature set and the corresponding historical price data of battery-grade lithium carbonate were divided into a training set (70%, 588 records), a validation set (20%, 168 records), and a test set (10%, 84 records) in a 7:2:1 ratio. The training set was expanded to 882 records by time series shifting and noise addition (adding 0.01 times random noise) to improve the model's generalization ability.

[0117] 3. Model training parameter settings: A composite loss function (L=0.6×MSE+0.4×MAPE) was used, along with the Adam optimizer (initial learning rate 0.001, decaying to 0.9 every 10 epochs), 100 training epochs, a batch size of 32, and an early stopping strategy (patience=15). A grid search method was used to optimize key hyperparameters, and the optimal hyperparameter combination was selected: 8 Transformer attention heads, 128 / 64 LSTM hidden layer units, dropout=0.3, and convolutional kernel size 3 / 5 / 7.

[0118] 4. Model Training Results: During training, the validation set loss reached its minimum value (L=0.023) at the 68th epoch and did not decrease for the next 15 consecutive epochs, triggering the early stopping strategy and stopping training; the training set loss was 0.018 and the validation set loss was 0.023, with no obvious overfitting (overfitting criterion: the validation set loss is more than 50% higher than the training set loss); preliminary validation on the test set showed that the daily prediction MAPE was 4.2% and the monthly prediction MAPE was 6.8%, both meeting the preset accuracy requirements (daily ≤5%, monthly ≤8%).

[0119] Step 9: Model Prediction and Preliminary Verification of Prediction Results Using the trained model, multi-scale predictions were performed, and a three-layer verification mechanism was established to initially verify the reasonableness of the prediction results. The specific operations and results are as follows: 1. Prediction Input Data Preparation: Collect real-time multimodal data from December 1 to 31, 2024. Process the data into a 64-dimensional fusion feature set according to steps 2-7. Verify the completeness, authenticity, and timeliness of the input data (ensure it is the latest data in December). Two data points were found to be outdated (lithium ore import data in November). After collecting additional data for December, the data was input into the model.

[0120] 2. Multi-scale Forecasting: The model's multi-scale forecasting function is invoked, outputting the following forecast results: Short-term forecast (January 1-7, 2025): Daily price forecast (RMB 283,000-291,000 / ton), deviation range ±3%, confidence level 0.86-0.91; Medium-to-long-term forecast (January-March, 2025): Monthly average price forecast (RMB 287,000 / ton in January, RMB 295,000 / ton in February, RMB 292,000 / ton in March), deviation range ±5%, confidence level 0.82-0.88. Example: A predicted price of RMB 285,000 / ton for January 1, 2025, with a deviation range of RMB 276,000-294,000 / ton and a confidence level of 0.89, indicates that the price on that day is highly likely to fall within this range, and the forecast is reliable.

[0121] 3. Preliminary Verification: A three-layer verification rule was adopted. The first layer was numerical range verification (the reasonable range for battery-grade lithium carbonate is 150,000-350,000 yuan / ton), and all predicted results were within the reasonable range. The second layer was trend consistency verification, comparing the actual trend in December 2024 (oscillating upward), and the predicted results showed a mild upward trend, consistent with the trend, with no abnormalities. The third layer was deviation reasonableness verification, calculating the deviation between the predicted value and the average price of the last 7 days of December 2024 (281,000 yuan / ton). The maximum deviation was 2.8%, which did not exceed 10%, and there were no sudden factors supporting the abnormal deviation. The verification was qualified.

[0122] 4. Classification and processing of forecast results: All forecast results are verified and compiled into a preliminary forecast report, which clarifies the forecast time, price, deviation range, confidence level and core influencing factors (such as the price increase from January to March, mainly due to the increase in the installed capacity of power batteries). There are no suspicious or unreasonable forecast results, and there is no need to re-forecast or remove them.

[0123] Step 10: Prediction Result Optimization and Dynamic Correction The sources of prediction bias were analyzed, a personalized optimization strategy was adopted, and a dynamic correction mechanism was established to optimize the prediction results. The specific operations and results are as follows: 1. Analysis of sources of deviation: Comparing the preliminary forecast results with the actual prices from January 1st to 7th, 2025 (RMB 284,000-293,000 / ton), the forecast deviation was calculated. It was found that the deviation mainly came from three aspects (all of which were slight deviations): First, the input data had a slight lag (some macroeconomic data lagged by 10 days); second, the feature fusion weights did not fully adapt to the changes in public opinion in December (at the end of December, there was public opinion about "the demand for power batteries breaking out ahead of schedule", and the weights of unstructured features were too low); third, the model parameters did not adapt to the recent slight adjustment in supply and demand (the import volume of lithium ore increased slightly).

[0124] 2. Personalized optimization strategy: To address the lag in input data, a trend prediction correction mechanism is adopted, which corrects macroeconomic data based on recent data trends; to address feature fusion bias, modal weights are recalculated, the unstructured modal weights are increased to 0.28, and the weights of public opinion features are adjusted; to address model parameter bias, an online fine-tuning strategy is adopted, using data from December 2024 to fine-tune model hyperparameters (learning rate 0.0008, fine-tuning parameters accounting for 8%), and a small amount of training is conducted for 5 epochs to adapt to recent supply and demand changes.

[0125] 3. Dynamic Correction: Correction intervals are set. Short-term forecasts (1-7 days) are corrected daily, adjusting the forecasts for subsequent days based on the actual price of the day. Medium- to long-term forecasts (1-3 months) are corrected twice monthly. On January 15th, the forecasts for February and March are adjusted based on the actual data from the first 15 days of the month (February adjusted to RMB 293,000 / ton, March adjusted to RMB 290,000 / ton). An error feedback mechanism is established to calculate the forecast error from January 1st to 7th (MAPE=3.7%, a decrease of 0.5% compared to before optimization), record the source of error, and feed it back to the feature fusion and model training steps. Two lithium carbonate industry experts were invited to conduct empirical correction. The experts determined that the price increase in January was slightly lower than the model's prediction, so the forecast value for late January was slightly adjusted (downwarded by RMB 2,000 / ton) to improve practicality.

[0126] 4. Final forecast output: After optimization, the short-term forecast (January 1-7) MAPE is 3.7%, and the medium-to-long-term forecast (January-March) MAPE is 6.1%. The results are compiled into a standardized forecast report with the following application recommendations: It is recommended that companies moderately increase their inventory in January (in conjunction with the moderate increase in short-term prices), maintain reasonable inventory in February and March, and formulate monthly production plans based on the medium-to-long-term forecast (with February's output slightly higher than that of January and March).

[0127] Step 11: Interpretability analysis and visualization of prediction results Multiple interpretability methods were employed to analyze the prediction results, design visualization charts, and compile reports tailored to different audiences. The specific operations and results are as follows: 1. Interpretability analysis: (1) Feature Importance Analysis (SHAP value): The SHAP value of each feature in the integrated feature set was calculated. It was found that the power battery installation growth rate (SHAP=0.27), lithium mine supply and demand difference (SHAP=-0.22), and public opinion sentiment score (SHAP=0.18) are the top three features affecting the prediction results. Among them, the power battery installation growth rate positively drives the price increase, while the lithium mine supply and demand difference negatively inhibits the price increase, which is consistent with the actual market.

[0128] (2) Modal contribution analysis: In the medium- and long-term forecast (1-3 months), the structured mode contributed 45%, the semi-structured mode 36%, and the unstructured mode 19%, indicating that the medium- and long-term forecast mainly depends on the supply and demand fundamentals and market expectations; in the short-term forecast (1-7 days), the unstructured mode contributed 42%, the structured mode 34%, and the semi-structured mode 24%, indicating that the short-term forecast mainly depends on short-term sudden factors such as public opinion.

[0129] (3) Branch contribution analysis: The market showed a moderate upward trend in January (stable trend), with a long-term branch contribution of 72% and a short-term branch contribution of 28%. If a sudden public opinion event occurs (such as a lithium mine safety accident), the short-term branch contribution will rise to 68% and the long-term branch to 32%, which is reasonable.

[0130] (4) Abnormal prediction analysis: There were no obvious abnormal prediction results, and there were slight deviations (such as the predicted value of RMB 287,000 / ton on January 3, while the actual value was RMB 289,000 / ton). This was mainly due to the untimely adjustment of the weight of public opinion features. The correction interval can be optimized in the future.

[0131] 2. Visual Presentation: Six types of visualization charts are designed: Price Forecast Trend Chart (line chart, labeling historical prices, forecast prices, and deviation range, clearly showing the historical trend for the whole of 2024 and the forecast trend for January-March 2025); Feature Importance Bar Chart (horizontal axis: feature name, vertical axis: absolute value of SHAP value, labeling the direction of influence); Modal and Branch Contribution Pie Chart (showing the contribution of three modalities and two branches respectively); Multimodal Data Correlation Heatmap (showing the correlation between features); Forecast Error Analysis Histogram (MAPE distribution, most errors are concentrated in 2%-4%, with good accuracy); Interpretability Analysis Flowchart (showing 12 steps and core outputs).

[0132] 3. Visualized Report Compilation: Two types of reports are designed: Professional Technical Reports (for technical personnel, detailing the 12 steps, model parameters, feature extraction details, error analysis, and providing model optimization suggestions); and Decision Application Reports (for enterprise decision-makers, simplifying technical details, focusing on prediction results, core influencing factors, and application suggestions, using concise charts for quick understanding). Interactive functions are provided; clicking on charts allows viewing detailed data (e.g., clicking on a feature importance bar chart to view specific SHAP values ​​and impact analysis).

[0133] 4. Report Updates: Short-term forecast visualization reports are updated daily, medium- and long-term reports are updated monthly, and forecast results and visualization charts for February and March are updated simultaneously on January 15th to ensure the timeliness of the reports.

[0134] Step 12: Method Validation and Continuous Improvement An effectiveness verification mechanism was established to comprehensively evaluate the method's performance, and a continuous improvement mechanism was established to ensure that the method remains adaptable to market demands in the long term. Specific operations and results are as follows: 1. Validation Metrics and Dataset: Four types of evaluation metrics were set: prediction accuracy metrics (MSE, MAPE, R²), generalization ability metrics (5-fold cross-validation error, accuracy differences in different scenarios), usability metrics (prediction efficiency, interpretability score, dynamic adaptation ability), and comparison metrics (comparison with ARIMA, LSTM, and simple multi-source fusion models). The evaluation dataset consisted of the test set (84 records) from step 8 plus real-time data from July to December 2024 (180 records), covering three market scenarios: rising, falling, and fluctuating (rising from January to March 2024, fluctuating from April to June, falling from July to September, and rising from October to December), ensuring comprehensive validation.

[0135] 2. Validation results: (1) Basic accuracy verification: The daily forecast MAPE=3.7% and R²=0.93 for the test set, and the monthly forecast MAPE=6.1% and R²=0.88, both of which meet the preset requirements; MSE=0.016, indicating good forecast accuracy.

[0136] (2) Generalization ability verification: 5-fold cross-validation error = 0.021. The accuracy difference is small in different scenarios (MAPE = 3.9% in the upward scenario, 4.1% in the oscillating scenario, and 3.8% in the downward scenario), indicating strong generalization ability and adaptability to different market environments.

[0137] (3) Practicality verification: Single batch prediction time = 7.2 seconds (≤10 seconds), expert interpretability score = 8.6 points (≥8 points), after the market structure changes (such as the price changing from rising to falling in July 2024), the prediction accuracy decreases by 7.3% (≤10%), the dynamic adaptation capability is good and meets the actual application needs of enterprises.

[0138] (4) Comparative Validation: Compared with existing mainstream models, the daily MAPE of this method is 3.8% lower than that of the LSTM model, 5.2% lower than that of the ARIMA model, and 4.5% lower than that of the simple multi-source fusion model; the interpretability score is 2.4 points higher than that of the simple multi-source fusion model, and the prediction efficiency meets the requirements, highlighting the advanced nature and practicality of this method. At the same time, feedback from enterprise users was collected, and 85% of users believed that the prediction results were reasonable and the application suggestions were highly targeted, and 78% of users believed that the visualization report was easy to use.

[0139] 3. Continuous Improvement Mechanism: Establish a problem collection mechanism to regularly collect problems encountered in the application of the method (such as the ability to improve short-term public opinion response speed and add mobile adaptation to the visualization report); formulate targeted improvement plans to optimize the speed of public opinion feature extraction in step 5 (using a lightweight BERT model), optimize the visualization report in step 11, and add mobile adaptation functionality; establish an improvement effect verification mechanism, after which the speed of public opinion feature extraction is improved by 30%, the mobile report adapts well, and the prediction accuracy does not decrease; establish a method update mechanism, with a comprehensive update every quarter, and the plan to introduce the latest Transformer-LSTM hybrid model in Q1 2025 to further improve prediction accuracy; compile effectiveness verification reports and continuous improvement plans to form a closed-loop management of "verification-improvement-re-verification".

[0140] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims.

Claims

1. A lithium carbonate price prediction method based on multimodal data collaborative driving, characterized in that, Includes the following steps: S1: Steps for determining multimodal data sources and acquiring data; S2: Steps for multimodal data classification and standardization; S3: Steps for personalized noise filtering and outlier correction of multimodal data; S4: Steps for extracting trend and correlation features from structured data; S5: Steps for extracting semantic and visual / speech features from unstructured data; S6: Steps for extracting key information and quantifying features from semi-structured data; S7: Steps for performing multimodal feature weighted collaborative fusion and feature optimization; S8: Steps for constructing and training a dual-branch collaborative prediction model; S9: Steps for performing model prediction and preliminary verification of prediction results; S10: Steps for optimizing and dynamically correcting prediction results; S11: Steps for interpreting and visualizing the prediction results; S12: Steps for validating the effectiveness of the method and for continuous improvement.

2. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 1, characterized in that, Step S1 specifically includes: Based on the core influencing factors of lithium carbonate price fluctuations, three major categories of multimodal data sources were identified, and the collection scope and standards for each type of data were clearly defined to avoid data redundancy or missing key data. The first category is structured data, which has a clear numerical format and statistical standards, reflecting the supply and demand fundamentals of the lithium carbonate market, the macroeconomic environment, and the operation of the upstream and downstream of the industrial chain. Specifically, it includes: historical lithium carbonate price data, i.e., daily, weekly, and monthly prices of industrial-grade and battery-grade lithium carbonate, with a collection range of the past 10 years and data accuracy to yuan / ton; lithium ore supply and demand data, i.e., lithium ore mining volume, inventory, import volume, and export volume in major global producing areas, collected monthly; battery industry data, i.e., new energy vehicle production, power battery installation volume, and battery recycling volume, collected monthly; macroeconomic data, i.e., GDP, PPI, RMB exchange rate, and commodity composite index, collected quarterly; and policy quantitative data, i.e., new energy subsidy intensity, number of mining approvals, and environmental control levels, collected annually / quarterly and converted into numerical values ​​using a quantitative scoring method; the second category is structured data. The second category is semi-structured data. This type of data has a certain structural framework but lacks a standard numerical format. It includes: industry research reports, such as lithium carbonate market analysis reports released by securities firms and industry associations, which focus on extracting key information such as supply and demand forecasts, price trend judgments, and analysis of core influencing factors; and announcements from companies in the industrial chain, such as capacity announcements, performance announcements, and capacity expansion / contraction information from lithium mining companies, lithium carbonate producers, and battery companies, collected quarterly / annually. The third category is unstructured data. This type of data has no fixed format and mainly reflects the impact of short-term unforeseen factors on prices. Specifically, it includes: news and public opinion data, such as news texts from mainstream financial and industry media regarding lithium mining, lithium carbonate production, supply chain logistics, policy regulation, and safety accidents, collected daily; image data, such as images of lithium mining sites, lithium carbonate production lines, and new energy vehicle production workshops, collected weekly to determine capacity utilization; and audio data, such as recordings of speeches about the lithium carbonate market at industry seminars and company press conferences, collected monthly to extract key viewpoints. Establish multi-source data acquisition channels to ensure data real-time performance and integrity. For structured data, connect to industry databases, government statistical department websites, and corporate websites, using API interfaces for automatic data collection. Set fixed collection times daily, weekly, and monthly to ensure timely data updates. For data that cannot be collected via API interfaces, employ a combination of manual entry and verification, establishing standardized data entry specifications to avoid errors. For semi-structured data, use web crawling technology to selectively crawl research reports and announcements from designated securities firms, industry associations, and corporate websites. Extract key information using PDF and HTML parsing techniques. The data is converted into a processable text format. Simultaneously, a monitoring mechanism for updating reports and announcements is established. Once new content is released, the data collection process is immediately initiated. For unstructured data, news and public opinion data are obtained by connecting to a public opinion monitoring platform and using keyword retrieval methods. Keywords include "lithium carbonate price," "lithium mining," "power battery installation," "lithium mine safety accident," and "lithium carbonate supply chain." Image data is obtained through industry cooperation channels, filtering for clear and effective images and removing blurry or irrelevant ones. Audio data is obtained by recording industry seminars and company press conferences, or by connecting to relevant audio platforms and using professional audio acquisition tools to ensure clear and noise-free audio. Finally, a data collection and verification mechanism is established to eliminate false, erroneous, and duplicate data. Targeted verification rules are formulated for each type of data collected: for structured data, the focus is on verifying numerical range, data integrity, and data consistency; for semi-structured data, the focus is on verifying source reliability and information integrity; and for unstructured data, the focus is on verifying authenticity, relevance, and completeness. Data that fails verification is marked as abnormal data, and the reasons for the abnormality are recorded. Through supplementary collection, correction, or removal, the comprehensiveness, authenticity, and validity of the collected multimodal data are ensured, providing a reliable foundation for subsequent data preprocessing steps.

3. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 2, characterized in that, Step S2 specifically includes: The multimodal data collected in step S1 is classified and organized to establish a categorized database. Based on the three major data types determined in step S1, the collected data is stored in a structured database, a semi-structured database, and an unstructured database. Each database is further categorized by data type: the structured database is subdivided into "price data," "supply and demand data," "macroeconomic data," and "policy quantitative data," with each sub-type sorted by time. The semi-structured database is subdivided into industry research reports and corporate announcements, with each sub-type sorted by publication time and issuing institution. The unstructured database is subdivided into news and public opinion data, image data, and audio data. News and public opinion data is sorted by publication time and keywords, image data by collection time and location, and audio data by recording time and speaker. A unique identifier is added to each data point, including data type, collection time, and source, to facilitate subsequent data traceability and management and avoid data confusion. Secondly, targeted standardization methods are adopted for different types of multimodal data to eliminate format differences and dimensional effects. For structured data, since it already has a numerical format, dimensional standardization is performed to eliminate dimensional differences between different indicators. This step uses a combination of two standardization methods: for structured data that follows a normal distribution, the Z-score standardization method is used to transform the data into standardized data with a mean of 0 and a standard deviation of 1; for structured data that does not follow a normal distribution, the Min-Max standardization method is used to map the data to the [0,1] interval. In practice, the normality test is first performed on the data of each structured indicator, and the corresponding standardization method is selected according to the test results to ensure the rationality of the standardization process. For missing values ​​in the structured data, linear interpolation is used to fill in the missing values ​​to avoid the missing values ​​affecting subsequent processing. The linear interpolation method constructs a linear function by using the effective data adjacent to the missing value to calculate the estimated value of the missing value, ensuring that the supplemented data conforms to the data change trend. For semi-structured data, the core is to convert it into a structured or standardized text format to facilitate subsequent collaborative processing. For industry research reports, PDF parsing and text extraction technologies are used to extract key information, converting numerical information into structured data and textual information into standardized text. For company announcements, key information such as production capacity, performance, and capacity adjustments is extracted; similarly, numerical information is converted into structured data and textual information into standardized text, ensuring that semi-structured data can be processed collaboratively with other types of data. For unstructured data, diverse standardization methods are employed to transform it into a format suitable for subsequent processing. For news and public opinion text data, text standardization is used: special symbols, punctuation marks, and irrelevant words are removed; Chinese word segmentation and stop word removal are performed; and the text is uniformly converted to UTF-8 encoding and lowercase format to ensure consistent format across news texts from different sources. For image data, image standardization is used: all images are standardized in size and color space; images are normalized; image pixel values ​​are mapped to the [0,1] range; and noise is removed to ensure consistency and clarity. For speech data, speech standardization is used: the speech sampling rate and bit depth are standardized; noise is removed; the speech is framed; and the speech is converted into a standardized audio frame format, laying the foundation for subsequent speech feature extraction. Establish a standardized data verification mechanism to ensure that the standardized data meets the requirements. For each type of standardized data, focus on checking the uniformity of format, data integrity, and data rationality. For data that fails verification, re-standardize it until it meets the requirements.

4. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 1, characterized in that, Step S3 specifically includes: First, we analyze the noise types and sources of different modalities of data to clarify the key points and difficulties of noise filtering. The noise of structured data mainly comes from statistical errors, data entry errors, and data source deviations, with outliers mainly manifested as numerical mutations. The noise of semi-structured data mainly comes from redundant information, vague statements, and false information in reports or announcements, with outliers mainly manifested as contradictory information. The noise sources of unstructured data are the most complex. The noise of news and public opinion data mainly comes from fake news, irrelevant information, and duplicate news, with outliers mainly being extremely negative / positive false public opinion. The noise of image data mainly comes from image blurring, noise points, and irrelevant backgrounds, with outliers mainly being image distortion and forged images. The noise of speech data includes environmental noise and recording interference, with outliers mainly being speech breaks, forged speech, and irrelevant remarks. Secondly, for structured data, a dual processing strategy of noise filtering and outlier correction is designed. The first step, noise filtering, employs an adaptive filtering method based on a sliding window. Different sliding window sizes are set according to the time characteristics of different structured indicators. The mean and standard deviation of the data within the sliding window are calculated. Data deviating from the mean by more than twice the standard deviation is marked as suspected noise data. This is verified in conjunction with the reliability of the data source. If confirmed as noise data, it is replaced with the mean within the window; if it is data with reasonable fluctuations, it is retained. For the same structured indicator data from different sources, a weighted average method is used to fuse the data, eliminating noise caused by data source bias. The second step, outlier correction, uses the Grubbs test to identify outliers. For identified outliers, a judgment is made based on the actual market situation. If the outlier is a reasonable anomaly caused by short-term sudden factors, the data is retained and marked as a sudden outlier point for subsequent targeted processing by the model. If the outlier is unreasonable, a correction method based on trend fitting is used. A quadratic polynomial fitting function is constructed using the effective data before and after the outlier to calculate the corrected value of the outlier, replacing the original outlier to ensure that the corrected data conforms to market trends. For semi-structured data, a noise filtering strategy combining source verification, information screening, and contradiction correction is designed. The first step, source verification, involves establishing an authoritative source database to verify the issuing institutions of industry research reports and company announcements, eliminating content from non-authoritative institutions to reduce false information noise. The second step, information screening, uses a combination of keyword matching and semantic analysis to filter out key information related to lithium carbonate prices, eliminating redundant information and ambiguous statements, and retaining core content. The third step, contradiction correction, compares the statements of the same indicator in different reports or announcements. If contradictions exist, they are verified in conjunction with authoritative sources and actual market conditions to confirm the correct information and correct contradictory data. If verification is not possible, the data is marked as suspicious and temporarily removed to avoid interfering with subsequent processing. For unstructured data, personalized noise filtering and outlier correction strategies are employed. For news and public opinion text data: First, duplicate news is removed using text similarity calculation; news with a similarity exceeding 0.8 is considered duplicate, and the earliest published and most authoritative source is retained. Second, a fake news identification method based on the BERT model is used; the news text is input into a pre-trained BERT model to identify and remove fake news. Third, sentiment analysis is used to mark news with extreme sentiment and no factual basis as abnormal public opinion. This is verified in conjunction with actual market conditions; if it is false extreme public opinion, it is removed; if it is a real extreme event, it is retained and marked. For image data: First, Gaussian filtering is used to remove image noise points, and edge detection algorithms are used to extract image data. The image processing steps are as follows: First, the core region is removed, and irrelevant background is eliminated, retaining only the core region. Second, a CNN-based image distortion recognition model is used to identify blurry, distorted, and forged images, treating them as anomalous images and removing them. Third, the retained images are scored based on quality, using three dimensions: clarity, completeness, and relevance, with a maximum score of 1. Images with a score > 0.7 are retained to ensure the validity of the image data. For speech data: First, Wiener filtering is used to remove environmental noise, and a speech enhancement algorithm is used to improve speech clarity. Second, a speech anomaly recognition method based on MFCC features is used to identify speech breaks, forged speech, and irrelevant remarks, treating them as anomalous speech and removing them. Third, the retained speech is framed and preprocessed to ensure that the speech data can be used for subsequent feature extraction. Finally, the multimodal data after noise filtering and outlier correction are validated as a whole. The noise removal rate and outlier correction rate of each type of data are calculated to ensure that the noise removal rate is greater than 90% and the outlier correction rate is greater than 95%. At the same time, the integrity and rationality of the data are checked to avoid data loss due to over-filtering.

5. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 1, characterized in that, Step S4 specifically includes: First, we need to clarify the scope and objectives of feature extraction from structured data. Structured data includes five categories: historical lithium carbonate price data, lithium mine supply and demand data, battery industry data, macroeconomic data, and policy quantitative data. The extraction objective of this step is to extract long-term price trends, short-term fluctuations, and cyclical characteristics from historical price data; and to extract price-related correlation features from the other four categories of data. At the same time, we will analyze the degree of correlation between various indicators and prices, screen out highly correlated features, and avoid feature redundancy. Secondly, the trend characteristics of the structured data are extracted, focusing on historical lithium carbonate price data. The first step is long-term trend feature extraction: a combination of moving average and exponential smoothing methods is used to calculate the short-term moving average, long-term moving average, and exponential smoothing value of the price. The long-term trend characteristics of the price are extracted through the cross relationship between the short-term and long-term moving averages. Simultaneously, linear regression is used to fit the price time series, and regression coefficients are calculated to quantify the strength of the price trend. The second step is short-term volatility feature extraction: daily and weekly volatility of the price are calculated. Volatility = (current price - previous price) / previous price × 100%, using standard deviation, etc. The variance quantifies the volatility amplitude; simultaneously, it extracts extreme value characteristics of prices, including daily high, low, and closing prices, weekly / monthly high, low, and average prices, as well as the symmetry of volatility, measured by a skewness coefficient. Skewness > 0 indicates right-skewed volatility, meaning upward volatility is greater than downward volatility; skewness < 0 indicates left-skewed volatility, meaning downward volatility is greater than upward volatility. This comprehensively captures the short-term volatility characteristics of prices. The third step is the extraction of periodic features: Fourier transform is used to decompose the price time series, breaking down the price data into trend, periodic, and stochastic components. By analyzing the frequency and amplitude of the periodic component, the periodic characteristics of prices are extracted, clarifying the periodic patterns of price fluctuations. Then, the correlation features of the structured data are extracted to analyze the correlation between each indicator and the price of lithium carbonate. The first step involves extracting correlation features from lithium mine supply and demand data, battery industry data, macroeconomic data, and policy quantitative data: For lithium mine supply and demand data, features such as supply-demand gap, inventory turnover rate, and import dependence are extracted; for battery industry data, features such as the growth rate of installed power battery capacity, the growth rate of new energy vehicle production, and battery recycling rate are extracted; for macroeconomic data, features such as GDP growth rate, PPI growth rate, and RMB exchange rate fluctuation range are extracted; and for policy quantitative data, features such as policy support strength score and mining control intensity score are extracted. The second step uses correlation analysis to analyze the correlation between each extracted feature and the price of lithium carbonate, selecting features with high correlation. A significance level of α=0.05 is used, and a correlation coefficient absolute value > 0.5 is considered highly correlated. Specifically, a combination of Pearson correlation coefficient and Spearman rank correlation coefficient is used: Pearson correlation coefficient is used to analyze linear correlation, while Spearman rank correlation coefficient is used to analyze nonlinear correlation, avoiding omissions caused by a single correlation coefficient. Finally, the extracted trend features and correlation features are normalized by using the Min-Max standardization method to map them to the [0,1] interval, eliminating the dimensional differences between features and establishing a structured feature set to support the subsequent multimodal feature fusion steps. The correlation coefficients between each feature and the price of lithium carbonate are recorded as an important basis for subsequent feature weight allocation.

6. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 5, characterized in that, Step S5 specifically includes: First, semantic and sentiment features are extracted from news and public opinion text data. News and public opinion text data contains a large amount of information reflecting short-term, sudden factors; its core features are semantic information and sentiment tendency. This step uses a text feature extraction method based on an improved BERT model. The specific operations are as follows: First, the standardized and denoised news text is segmented, part-of-speech tagging is performed, and entity recognition is conducted using jieba segmentation combined with a BiLSTM-CRF entity recognition model to identify core entities in the text and label entity types, laying the foundation for semantic feature extraction; Second, a pre-trained BERT model for the lithium carbonate industry is constructed; Third,… The news text is input into the finely tuned BERT model, and the output vector of the last layer of the model is extracted as the semantic feature vector of the text. The dimension is set to 768, which can comprehensively capture the semantic information of the text. In the fourth step, the sentiment analysis model based on LSTM is used to score the sentiment of the news text. The scoring range is [0,1], where 0 represents extreme negative and 1 represents extreme positive. Sentiment features are extracted to reflect the short-term impact of public opinion on the price of lithium carbonate. The extracted semantic feature vector is dimensionality reduced by using the PCA principal component analysis method to reduce the 768-dimensional vector to 128 dimensions, eliminating feature redundancy and improving the efficiency of subsequent processing. Secondly, for image data, visual features are extracted, focusing on reflecting capacity utilization and supply-demand changes. Image data, including lithium mining sites, lithium carbonate production lines, and new energy vehicle production workshops, can intuitively reflect the capacity situation of the industry chain. A visual feature extraction method based on an improved CNN model is adopted, and the specific operations are as follows: First, an improved ResNet-50 model for the lithium carbonate industry chain is constructed. Based on the ResNet-50 model, the last fully connected layer is removed, and an adaptive average pooling layer is added. Industry image data is used for fine-tuning to improve the model's feature extraction capability for industry images. Second, data augmentation processing is performed on the denoised and standardized image data to improve the model's generalization ability and avoid over-generalization. The first step is fitting; the second step is to input the enhanced image into the improved ResNet-50 model and extract the output vector of the model's adaptive average pooling layer as the visual feature vector of the image. The dimension is set to 2048, which can capture the core visual information in the image; the third step is to perform dimensionality reduction on the visual feature vector. The t-SNE dimensionality reduction method is used to reduce the 2048-dimensional vector to 128-dimensional, eliminating redundancy and ensuring that the visual features can be synergistically fused with other modal features. An image classification model is used to classify the image into three levels of energy utilization: high, medium, and low. The classification results are converted into numerical features, with high = 1, medium = 0.5, and low = 0, to supplement the quantitative information of the visual features and improve the practicality of the features. Then, for the speech data, speech features are extracted to capture key industry viewpoints. The speech data includes the judgments and predictions of industry experts and company leaders on the lithium carbonate market. Its core features are the Mel-frequency cepstral coefficients and semantic information of the speech. A dual method of speech feature extraction and semantic transformation is adopted. The specific operation is as follows: First, the standardized and denoised speech data is segmented and windowed. The MFCC features of each frame of speech are calculated, and 13-dimensional MFCC coefficients, as well as first-order and second-order differences, are extracted, for a total of 39 dimensions. MFCC features can capture the spectral characteristics of speech. The first step involves using the speech-to-text model based on Wav2Vec2.0 to convert speech data into text data. Then, the improved BERT model from step S5 is used to extract the semantic feature vector of the text, reducing the dimensionality to 128 dimensions to capture the core viewpoints in the speech. The second step involves concatenating the MFCC features and semantic feature vectors to obtain the comprehensive feature vector of the speech, with a dimension of 39 + 128 = 167 dimensions. This is then further reduced to 128 dimensions using PCA to ensure that the dimensions are consistent with other unstructured features, facilitating subsequent fusion. Finally, the extracted text semantic features, sentiment features, image visual features, and speech comprehensive features are normalized and mapped to the [0,1] interval to establish an unstructured feature set, which provides support for the subsequent multimodal feature fusion steps. The effectiveness of each type of unstructured feature is tested by using analysis of variance to test whether the features can effectively distinguish different price fluctuation scenarios, and invalid features are eliminated to ensure the practicality of the unstructured features.

7. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 1, characterized in that, Step S6 specifically includes: First, the scope of key information extraction from semi-structured data is clearly defined. Considering the factors influencing lithium carbonate price fluctuations, the key information is categorized into three main types: supply and demand forecasts, capacity adjustment information, and analysis of core influencing factors. These three types of information reflect market expectations for lithium carbonate prices and changes in the medium- to long-term supply and demand pattern, providing significant reference value for price forecasting. They also represent core content that existing methods have not fully explored. Secondly, for industry research reports, a method combining key information location, semantic analysis, and quantitative scoring is used to extract and quantify key information. The first step, key information location, employs keyword-based and regular expression-based methods to locate key paragraphs in the report related to supply and demand forecasts, capacity adjustments, and influencing factor analysis. Simultaneously, a text paragraph classification model is used to categorize report paragraphs, filtering out those containing key information and eliminating irrelevant paragraphs. The second step, semantic analysis, applies an improved BERT model (consistent with the fine-tuned BERT model in step S5) to the selected key paragraphs, enhancing the semantic understanding of industry texts; semantic analysis is then performed to extract key information. The third step, key information quantification, quantifies the extracted key information... Key information is transformed into quantifiable features, with specific quantification rules as follows: For supply and demand forecast information, forecasted output and forecasted demand are transformed into numerical features, with the unit uniformly set to 10,000 tons, and the forecasted price range is transformed into a price forecast index; For capacity adjustment information, the capacity expansion and contraction ranges are transformed into numerical features, the capacity utilization rate is transformed into a quantified value in the [0,1] range, and the capacity adjustment time is transformed into the time elapsed since the current time; For core influencing factor analysis, the description of the degree of influence is transformed into a quantitative score, with significant influence = 1.0, moderate influence = 0.7, and slight influence = 0.3, the direction of influence is transformed into a sign feature, with positive influence = +1 and negative influence = -1, and the name of the influencing factor is transformed into a unique heat-coded feature; Then, for enterprise announcements, a method of information extraction, authenticity verification, and quantification is adopted to extract and quantify key information. The first step is information extraction: PDF parsing and text extraction technologies are used to extract key information such as production capacity, performance, production capacity adjustment, and cooperation agreements from the announcements. The second step is authenticity verification: Combining past announcements of enterprises and the actual situation of the industry, the extracted information is verified for authenticity. False or questionable information is removed to ensure the authenticity of the extracted information. The third step is information quantification: The extracted true key information is converted into quantitative features. The specific rules are as follows: production capacity data is converted into numerical features, production capacity utilization rate is converted into a quantitative value in the range of [0,1], revenue and profit are converted into growth rate features, cooperation scale is converted into numerical features, and production capacity adjustment time is converted into the time from the current time. Finally, the quantified semi-structured features were integrated and optimized: First, correlation analysis was used to analyze the degree of correlation between each semi-structured feature and the price of lithium carbonate, and features with an absolute correlation coefficient greater than 0.4 were selected, while redundant features were removed; Second, the selected features were normalized using Min-Max standardization, mapped to the [0,1] interval to eliminate the influence of dimensions; Third, all quantified semi-structured features were spliced ​​together to form a semi-structured feature set with a dimension of 128 to ensure compatibility with the dimensions of the structured and unstructured feature sets, providing supplementary support for subsequent multimodal feature fusion steps. The source and credibility score of each semi-structured feature were recorded, with credibility of authoritative sources = 1.0 and credibility of ordinary sources = 0.7, serving as the basis for subsequent feature weight allocation.

8. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 6, characterized in that, Step S7 specifically includes: First, the characteristics and importance of the three types of multimodal features are analyzed to determine the basis for feature weight allocation. Structured features primarily reflect the long-term trend of lithium carbonate prices and supply and demand fundamentals, serving as the core for predicting long-term price trends and thus having the highest importance. Semi-structured features primarily reflect market expectations and medium- to long-term capacity changes, acting as a bridge between long-term trends and short-term emergencies, and thus have the next highest importance. Unstructured features primarily reflect short-term emergencies, significantly impacting short-term price fluctuations but having a weaker long-term impact, and thus have relatively lower importance. Furthermore, the importance of different features within each modality also varies. A dual weighting strategy of modality weights plus feature weights is adopted to ensure the rationality and relevance of the weight allocation. Secondly, the weights of each modality are calculated using a modality weight allocation method based on information entropy. The specific steps are as follows: First, calculate the information entropy of the feature set for each modality. Information entropy measures the uncertainty and information content of a feature. The smaller the information entropy, the greater the information content of the feature and the greater its contribution to prediction. The formula for calculating information entropy is: ; in, Information entropy represents the feature set of a certain modality; This indicates the number of features in the feature set of this modality; Indicates the first The probability density of each feature is calculated using the kernel density estimation method, which estimates the probability density of each feature based on the sample distribution of the feature. Indicates the first The higher the probability density of a feature, the higher the information content; the smaller the negative value, the lower the information entropy. The second step is to calculate the weight of each mode based on the information entropy. The formula for calculating the mode weight is as follows: ; in, Indicates the first Modal class, For structured purposes, For unstructured, The weights are semi-structured. Indicates the first Information entropy of modal feature sets; Representing three types of modes The sum of these values ​​is used for normalization to ensure that the sum of all modal weights is 1. ); The weights of features within each mode are calculated using a feature weight allocation method based on correlation coefficients and confidence levels. The first step involves calculating the correlation coefficient between each feature within each mode and the price of lithium carbonate. The Pearson correlation coefficient, with a value range of [-1, 1], is obtained by taking the absolute value of the correlation coefficient. The first step is to measure the correlation between features and prices; the second step is to combine the feature's credibility score. The confidence level for structured features is 1.0, the confidence level for semi-structured features is determined based on the source, and for unstructured features, the confidence level for authoritative sources is 1.0, and the confidence level for ordinary sources is 0.

7. The third step is to calculate the weight of each feature. The calculation formula is: ; in, Indicates the first In the class modality, the first The weights of each feature; Indicates the first The absolute value of the correlation coefficient between each characteristic and the price of lithium carbonate; Indicates the first Credibility score for each feature; Indicates the first The number of features in a class modality; Indicates the first All features in class modality The sum of these weights is used for normalization to ensure that the sum of the weights of all features within each modality is 1. Multimodal feature weighted collaborative fusion is performed as follows: First, for the features within each modality, a weighted summation method is used to fuse all features of that modality into a single modality feature vector with a dimension of 128. Modal class, fused modal feature vector = ,in Indicates the first A vector of features, The first step represents the weight of the feature; the second step is to use modal weights to weight and fuse the feature vectors of the three modalities to obtain the final fused feature set. The calculation formula is: ; in, This represents the final fused feature set, with a dimension of 128. These represent the weights of the structured, unstructured, and semi-structured modes, respectively. These represent the feature vectors after the fusion of the three modalities. Through this formula, the synergistic fusion of the three types of multimodal features is achieved, which not only highlights the core role of structured features, but also takes into account the supplementary role of semi-structured and unstructured features, thereby maximizing the value of multimodal data. Finally, the fused feature set is optimized to improve feature quality: First, analysis of variance is used to test the effectiveness of the fused features and invalid features with variance less than 0.01 are removed; second, mutual information is used to analyze the redundancy between the dimensions within the fused features and redundant features with mutual information values ​​greater than 0.8 are removed; third, principal component analysis (PCA) is used to reduce the dimensionality of the optimized fused features, adjusting the feature dimensions to 64 dimensions, thereby improving the training efficiency and generalization ability of the subsequent prediction model while retaining core information; fourth, the dimensionality-reduced fused features are normalized to ensure that all features are in the [0,1] interval, laying the foundation for the input of the subsequent prediction model.

9. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 8, characterized in that, Step S8 specifically includes: The overall structure of the dual-branch collaborative prediction model is clearly defined. The model is divided into a long-term trend prediction branch and a short-term fluctuation prediction branch. The two branches share the input of the fused feature set and realize information interaction and weight allocation through a collaborative attention mechanism. Finally, the fused price prediction result is output. The long-term trend prediction branch is used to capture the long-term trend of lithium carbonate price and adapts to the information of structured and semi-structured features. The short-term fluctuation prediction branch is mainly used to capture the short-term fluctuation of price and adapts to the information of unstructured features. The two branches work together to balance the prediction accuracy of long-term trends and short-term fluctuations. Secondly, a long-term trend prediction branch is constructed using an improved Transformer model to capture long-term price dependencies and medium- to long-term supply and demand patterns. The first step is model structure improvement: based on the standard Transformer model, a positional encoding optimization mechanism is introduced, employing a concatenation of sine and cosine positional encoding with time feature encoding. The time feature encoding combines the daily, weekly, and monthly time attributes of price data to enhance the model's sensitivity to time series data. Simultaneously, the encoder structure is simplified, reducing the number of attention heads from 12 to 8, lowering model complexity, avoiding overfitting, while retaining the encoder's core ability to capture long-distance dependencies. The second step is input layer design: the optimized fusion feature set from step S7 is concatenated with trend features from the structured feature set as input to the long-term trend prediction branch, ensuring that the input features fully reflect long-term trend information. The third step is output layer design: a linear regression layer is used as the output layer, outputting the predicted lithium carbonate price for the next 1-3 months. A Dropout layer is added with a dropout probability set to 0.3 to improve the model's generalization ability and prevent overfitting during training. Then, a short-term volatility prediction branch is constructed, employing an improved CNN-LSTM hybrid model to focus on capturing short-term sudden fluctuations and random changes in prices. The first step involves model structure improvement: a CNN model is used as the feature extraction front-end, employing three convolutional layers with kernel sizes of 3, 5, and 7 to extract local features from the fused feature set. Max pooling layers with a kernel size of 2 are used to reduce feature dimensionality and retain key information. An LSTM model is used as the temporal prediction back-end, employing two LSTM layers with 128 and 64 hidden units to capture the temporal variation patterns of local features, addressing the limitation of CNN models in capturing temporal dependencies. The integration of CNN and LSTM... The first step involves adding an attention mechanism to weight the local features extracted by the CNN, focusing on features that significantly affect short-term fluctuations. The second step is the input layer design: the semantic, visual, and speech features from the optimized fusion feature set in step 7 and the unstructured feature set are combined, and short-term relevant features are selected from the 128-dimensional features after dimensionality reduction. These features are then concatenated into a 32-dimensional array and used as input to the short-term fluctuation prediction branch, ensuring that the input features fully reflect short-term fluctuation information. The third step is the output layer design: a linear regression layer is used as the output layer to output the predicted lithium carbonate price for the next 1-7 days. A BatchNorm layer is added to accelerate model training convergence and further improve the model's generalization ability. Next, a dual-branch collaboration mechanism is designed to achieve information interaction and prediction result fusion between the two branches. The first step is the collaborative attention mechanism design: a collaborative attention module is added before the output layers of the two branches. This module extracts the feature representations of the long-term trend prediction branch and the short-term volatility prediction branch, calculates the mutual information between the features of the two branches, and assigns attention weights based on the magnitude of the mutual information. When short-term volatility is high, the attention weight of the short-term volatility prediction branch is increased, with a weight range of 0.5-0.7; when the market is in a stable trend, the attention weight of the long-term trend prediction branch is increased, with a weight range of 0.6-0.8, achieving dynamic collaboration between the two branches. The second step is prediction result fusion: a weighted summation method is used to fuse the prediction results of the two branches to obtain the final price prediction value. The fusion weight is determined by the attention weight output by the collaborative attention module, and the calculation formula is as follows: ,in This represents the final predicted value. This represents the predicted value of the long-term trend branch. α represents the predicted value of the short-term fluctuation branch, and α represents the collaborative attention weight, 0 < α < 1; Finally, model training and hyperparameter optimization are performed. The first step is dataset partitioning: the optimized fused feature set from step 7 and the corresponding historical lithium carbonate price data are divided into training, validation, and test sets in a 7:2:1 ratio. The training set is used for model parameter training, the validation set for hyperparameter tuning and model generalization verification, and the test set for final model performance evaluation. Data augmentation techniques are used to expand the training set data volume to further improve the model's generalization ability. The second step is loss function selection: a composite loss function combining mean squared error and mean absolute percentage error is adopted, taking into account both the absolute and relative errors of the predicted values. The loss function calculation formula is as follows: The first step involves MSE, which measures the absolute deviation between the predicted and actual values, and MAPE, which measures the relative deviation of the predicted values. This avoids error assessment bias caused by a large range of price values. The second step is optimizer selection and training parameter settings: the Adam optimizer is used, with an initial learning rate of 0.

001. A learning rate decay strategy is adopted, reducing the learning rate to 0.9 every 10 epochs to accelerate model convergence. The number of training epochs is set to 100, the batch size to 32, and an early stopping strategy is adopted. Training is stopped when the validation set loss does not decrease for 15 consecutive epochs to avoid model overfitting. The third step is hyperparameter optimization: a grid search method is used to optimize the key hyperparameters of the model, selecting the hyperparameter combination with the minimum validation set loss as the final hyperparameters of the model. Step S9 specifically includes: First, the preparation and preprocessing of the prediction input data are clearly defined. The multimodal data collected in real time and processed undergoes standardization, denoising, feature extraction, feature fusion, and feature optimization to create a fusion feature set that meets the model input requirements, ensuring that the format and scale of the input data are consistent with the input data used during model training. At the same time, additional validation is performed on the real-time input data, focusing on checking the completeness, authenticity, and timeliness of the data. Input data that fails validation is supplemented or corrected before being input into the model to avoid the input data quality issues affecting the prediction results. Secondly, multi-scale price forecasting is performed. Based on actual application needs, the model's multi-scale forecasting function is activated: short-term forecasting, daily price forecasting for 1-7 days, calls the short-term fluctuation forecasting branch, outputs the daily lithium carbonate price forecast value and the forecast deviation range, which is calculated based on the daily forecast error during model training and set to ±3%; medium- and long-term forecasting, price forecasting for 1-3 months, calls the long-term trend forecasting branch, outputs the monthly average lithium carbonate price forecast and the forecast deviation range, which is set to ±5%; at the same time, the confidence level corresponding to the forecast result is output, calculated based on the forecast accuracy during model training, with a confidence level range of [0,1]. The higher the confidence level, the more reliable the forecast result, providing a reference for subsequent verification and decision-making; A preliminary verification mechanism for forecast results is established to make an initial judgment on the reasonableness and accuracy of the forecast results. A three-tiered verification rule is designed, progressively strengthening the process to ensure the reasonableness of the forecast results: The first tier is numerical range verification. Based on the historical price range and reasonable cost range of the lithium carbonate market, a reasonable range for the forecast results is set. If the forecast results exceed this range, they are marked as unreasonable, and the reason for the anomaly is recorded. The second tier is trend consistency verification. The trend of the forecast results is compared with the recent actual market trend. If the trends are inconsistent, the forecast results are marked as suspicious and further verified in conjunction with short-term unforeseen factors in the input data. The third tier is deviation reasonableness verification. The deviation between the forecast results and recent historical prices is calculated. If the deviation exceeds 10% and there is no clear unforeseen market factor to support it, the forecast results are marked as unreasonable and require further optimization. Finally, the preliminary verification results are categorized. Validated results (values ​​within a reasonable range, consistent trend, reasonable deviation, and confidence level ≥ 0.7) are compiled into a preliminary forecast report, specifying the forecast time, forecast price, deviation range, confidence level, and core influencing factors. For forecasts marked as suspicious, the input data and actual market conditions are further verified. Relevant data is supplemented and re-entered into the model for forecasting. If the re-forecast result is still suspicious, it is marked as a forecast to be optimized. For forecasts marked as unreasonable, they are directly discarded, and the reasons for the unreasonableness are analyzed and corrected in a timely manner, providing a basis for subsequent forecast optimization steps.

10. The lithium carbonate price prediction method based on multimodal data collaborative driving according to claim 9, characterized in that, Step S10 specifically includes: First, the sources of prediction bias are analyzed to clarify the direction of optimization. By comparing the preliminary prediction results with recent historical actual prices and real-time market data, and combining error analysis during model training, the main sources of prediction bias are summarized into four categories: First, input data bias, i.e., the real-time input multimodal data contains slight noise, lag, or missing key data, leading to model prediction bias; second, model parameter bias, i.e., insufficient adaptation of hyperparameters during model training, or differences in the distribution of model training data and real-time market data, resulting in a decrease in model generalization ability and prediction bias; third, feature fusion bias, i.e., in the multimodal feature fusion process of step S7, the feature weight allocation is unreasonable, resulting in the incomplete reflection of the influence of some key features and prediction bias; fourth, omission of sudden market factors, i.e., sudden market factors occurring in real time are not included in the input data in a timely manner, causing the model prediction results to fail to reflect the latest market changes. Secondly, personalized optimization strategies are designed to address prediction biases from different sources. The first type, prediction bias caused by input data bias, involves re-verifying and supplementing real-time input data. The noise filtering and outlier correction methods from step S3 are used to remove slightly noisy data and supplement missing key data. Simultaneously, a data lag correction mechanism is introduced to adjust the input data based on its lag time, reducing bias caused by data lag. The second type, prediction bias caused by model parameter bias, employs an online fine-tuning strategy. The latest real-time data is used to fine-tune the model's key hyperparameters, updating only a small number of parameters each time to avoid model performance fluctuations. Simultaneously, a transfer learning method is used, employing the trained model parameters as initial parameters and performing a few training iterations with new real-time data. The first category of prediction bias is caused by the following: First, the model can quickly adapt to new market data distributions, reducing prediction bias. Second, the prediction bias caused by feature fusion bias is addressed by recalculating the weights of multimodal features and adjusting them in conjunction with real-time market conditions. Third, when there are significant changes in the medium- to long-term supply and demand patterns, the modal weights of structured and semi-structured features are increased. Simultaneously, the fusion feature set is re-optimized, redundant features are removed, and new features that significantly impact the current market are added to improve feature quality. Fourth, the prediction bias caused by the omission of sudden market factors is addressed by establishing a real-time monitoring mechanism for sudden factors, capturing these factors in real time, and using the semi-structured data quantification method in step S6 to quickly quantify them into feature vectors, adding them to the fusion feature set, re-inputting them into the model for prediction, correcting the prediction results, and ensuring that the prediction results reflect the latest market changes. Then, a dynamic correction mechanism for the forecast results is established to achieve real-time updates and optimization of the forecast results. The first step is to set dynamic correction time intervals, with different intervals set according to the forecast scale: for short-term forecasts (1-7 days), correction is performed once daily, combining the latest collected multimodal data and actual market prices to correct the forecast results for the following days; for medium- and long-term forecasts (1-3 months), correction is performed twice monthly, combining the latest supply and demand data, public opinion data, and policy data for the month to correct the forecast results for the following months. The second step is to establish a forecast error feedback mechanism, calculating the error between each forecast result and the actual price, recording the error magnitude and source, and feeding the error back to the model training and feature fusion steps as a basis for online model fine-tuning and feature weight adjustment. If the forecast error of a certain branch of the model remains large, the structure and parameters of that branch are fine-tuned. The third step is to introduce expert experience correction, inviting lithium carbonate industry experts to evaluate the optimized forecast results. Combining the experts' judgments on market trends and industry experience, the forecast results are further corrected to improve their practicality and reliability. Finally, the optimized final prediction results are output, and a prediction report is compiled. The optimized prediction results, prediction deviation range, confidence level, core influencing factors, and optimization process are compiled into a standardized prediction report, clarifying the prediction conclusions and application suggestions. At the same time, the optimization process and error situation of the prediction results are recorded to provide data support for the improvement of subsequent prediction methods and model optimization. Step S11 specifically includes: First, an interpretability analysis of the prediction results is conducted, employing a combination of multiple interpretability methods to comprehensively analyze the formation process and influencing factors of the prediction results. First, feature importance analysis: using the SHAP value method, the SHAP value of each feature in the fused feature set is calculated to influence the prediction results. The larger the absolute value of the SHAP value, the greater the influence of that feature on the prediction results. Simultaneously, the direction of influence of the feature is distinguished: a positive SHAP value indicates that the feature drives up prices, while a negative SHAP value indicates that the feature drives down prices. Second, modal contribution analysis: the contribution of structured, unstructured, and semi-structured modal features to the prediction results is calculated. Based on a comprehensive calculation of modal weights and feature importance, the role of each type of modal data in the prediction process is clarified. The role of the model in short-term forecasting: In short-term forecasting, the unstructured mode contributes 40%, the structured mode contributes 35%, and the semi-structured mode contributes 25%, indicating that short-term forecasting mainly relies on short-term sudden factors; Third, branch contribution analysis: In the dual-branch collaborative forecasting model, the contribution of the long-term trend forecasting branch and the short-term fluctuation forecasting branch to the final forecast result is calculated, and the rationality of the branch contribution is analyzed in combination with the market environment; When the market is in a state of violent fluctuation, the short-term branch contributes 65%, and the long-term branch contributes 35%; Fourth, anomaly forecast interpretability analysis: For anomaly forecast results with large deviations in the forecasting process, the causes of anomalies are analyzed in combination with feature importance and modal contribution, clarifying the formation mechanism of anomaly forecasts and providing direction for subsequent forecast optimization; Secondly, diverse visualization methods are employed to present prediction results, interpretable analysis conclusions, and related data, enhancing readability. Multi-dimensional visualization charts are designed, balancing professionalism and ease of use: First, a price prediction trend chart, presented as a line graph, with time on the horizontal axis and price on the vertical axis, simultaneously marking historical actual prices, predicted prices, and prediction deviation ranges, clearly showing historical and future price trends for easy comparative analysis; Second, a feature importance bar chart, with feature name on the horizontal axis and absolute SHAP value on the vertical axis, intuitively displaying the degree of influence of each feature on the prediction results and indicating the direction of feature influence; Third, a pie chart of modal and branch contribution is used to present the contribution of the three major modal features and the contribution of the two branches, clearly showing the role of each type of data and each branch in the prediction; Fourth, a heatmap of multimodal data correlation presents the correlation between different modal features and the correlation between each feature and price, intuitively showing the inherent relationship between features; Fifth, a histogram of prediction error analysis presents the distribution of prediction errors, shows the prediction accuracy of the model, and marks the time nodes with large errors and their causes; Sixth, a visualization of the interpretability analysis report organizes the core conclusions of the interpretability analysis into a report combining text and graphics. Next, the visualization reports were organized, clarifying the target audience and presentation format. Based on the different audiences, two types of visualization reports were designed: one is a professional technical report, aimed at technical personnel, which details the steps of the prediction method, model parameters, feature extraction process, interpretability analysis details, error analysis results, and model optimization suggestions to meet the needs of technical personnel for method optimization and model improvement; the other is a decision-making application report, aimed at enterprise decision-makers, which simplifies technical details and focuses on presenting prediction results, core influencing factors, prediction confidence levels, and application suggestions, using concise and clear visualization charts and text descriptions to facilitate decision-makers' quick understanding and application of the prediction results; at the same time, interactive functions are provided for the visualization reports, allowing users to click on charts to view detailed data; Finally, an update mechanism for the visualization reports is established to ensure their timeliness. Based on the dynamic correction time interval specified in step S10, the visualization reports are updated synchronously: short-term forecast visualization reports are updated daily, and medium- to long-term forecast visualization reports are updated monthly, supplementing them with the latest forecast results, actual price data, interpretability analysis conclusions, and error analysis results. Simultaneously, when significant changes occur in the market environment, the visualization reports are updated promptly, adding impact analysis of these significant changes to ensure the reports reflect the latest market conditions and forecast results. Step S12 specifically includes: First, the evaluation metrics and dataset for validity verification are clearly defined to ensure the scientific rigor and objectivity of the verification results. Four core evaluation metrics are set to comprehensively cover prediction accuracy, generalization ability, and practicality: First, prediction accuracy metrics, including mean squared error (MSE), mean absolute percentage error (MAPE), and coefficient of determination (R²). The smaller the MSE and MAPE, and the closer R² is to 1, the higher the prediction accuracy. For short-term predictions, MAPE ≤ 5% and R² ≥ 0.9 are required; for medium- and long-term predictions, MAPE ≤ 8% and R² ≥ 0.85 are required. Second, generalization ability metrics, including cross-validation error and differences in prediction accuracy under different market scenarios. The smaller the cross-validation error and the smaller the differences in accuracy under different scenarios, the stronger the generalization ability. Third, practicality metrics, including prediction efficiency, interpretability score, and dynamic adaptation ability. Fourth, comparative metrics, comparing this method with existing mainstream prediction methods, calculating the evaluation metrics of each method, and clarifying the advantages of this method. The evaluation dataset uses the test set divided in step S8, and is supplemented with real-time market data from the past 6 months to ensure that the evaluation dataset can cover different market scenarios and avoid the one-sidedness of the verification results. Secondly, the effectiveness of the method was validated in three stages: The first stage, basic accuracy validation, used test set data to calculate the MSE, MAPE, and R² accuracy metrics of the method, validating its basic predictive ability and determining whether it met the preset accuracy requirements. If not, the reasons were analyzed, and the corresponding steps were returned for optimization. The second stage, generalization ability validation, employed 5-fold cross-validation. The evaluation dataset was randomly divided into 5 parts, with 4 parts used as the training set and 1 part as the test set, and the validation was repeated 5 times, calculating the cross-validation error. Simultaneously, the evaluation dataset was divided according to market scenarios, and the prediction accuracy under different scenarios was calculated to verify the method's adaptability to different market environments. If the generalization ability was insufficient, the model structure and feature fusion strategy were optimized to improve generalization. The third stage, comparative validation, tested the method against existing mainstream prediction methods on the same evaluation dataset, comparing the accuracy metrics, generalization ability metrics, and practicality metrics of various methods to clarify the advantages of the method. Furthermore, industry experts and enterprise users were invited to evaluate the practicality of the method, and user feedback was collected as a basis for method improvement. Then, a continuous improvement mechanism is established to continuously optimize the prediction method based on the effectiveness verification results and market changes. The first step is to establish a problem collection and analysis mechanism, regularly collecting problems encountered during method application, shortcomings discovered during effectiveness verification, and user feedback suggestions. Problems are categorized, their sources analyzed, and improvement directions clarified. The second step is to formulate targeted improvement plans, optimizing the relevant steps of the method according to the problem's source: If the problem stems from data quality, optimize the data collection, standardization, and denoising processes in steps S1 to S3, expand data sources, and improve data quality; if the problem stems from feature fusion, optimize the feature weight allocation strategy and feature optimization method in step S7, and supplement with new features; if the problem stems from model performance, optimize the model structure and training parameters in step S8, introduce new deep learning technologies, and improve model performance; if the problem stems from interpretability... The first step is to optimize the interpretability analysis method in step S11 and increase the richness of the visualization charts. If the problem stems from dynamic adaptability, the second step is to optimize the dynamic correction mechanism in step S10, shorten the correction interval, and improve the dynamic adaptability of the method. The third step is to establish an improvement effect verification mechanism. For each improved method, the evaluation dataset and real-time data are used for verification, evaluation indicators are calculated, and the effects before and after the improvement are compared to ensure that the improvement can improve the performance of the method. If the improvement effect does not meet expectations, the improvement plan is adjusted and re-optimized until the preset requirements are met. The fourth step is to establish a method update mechanism. The prediction method is comprehensively updated every quarter, integrating the latest technological progress, market data changes, and user feedback to continuously improve the advancement and practicality of the method. A method version management mechanism is established to record the content, improvement direction, and effect of each update for easy traceability and subsequent optimization. Finally, an effectiveness verification report and a continuous improvement plan are compiled to form a closed-loop management system. The effectiveness verification process, evaluation indicators, comparison results, and shortcomings are summarized in the effectiveness verification report, clarifying the advantages and improvement directions of the method. Based on the shortcomings and user feedback, a medium- to long-term continuous improvement plan is formulated, specifying the person in charge, completion time, improvement goals, and verification standards for each improvement task. At the same time, the effectiveness verification report and continuous improvement plan are shared with technical personnel and enterprise users to solicit further suggestions, forming a closed-loop management system of verification-problem identification-improvement-re-verification. This ensures that the method can adapt to the dynamic changes in the lithium carbonate market in the long term and continuously provide enterprises with accurate and reliable price forecasting services.