Bulk commodity transaction market stabilization method based on BSL-CPFS deep learning model

Through the analysis of the BSL-CPFS deep learning model and the transmission of telephone recording, the efficient analysis and transmission of unstructured data in the commodity market is solved, real-time structured conversion of market information and price prediction, and the transparency of market pricing and transaction stability are improved.

CN120410653APending Publication Date: 2025-08-01DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510532488.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently analyze and transmit unstructured data in telephone recordings, resulting in an intensification of information asymmetry in commodity markets, lagging in price correction mechanisms, and large market fluctuations.

Method used

The deep learning model based on BSL-CPFS, including DeepSpeech2, BERT, TextCNN, NER and BiLSTM networks, is adopted to realize speech transcription, semantic classification, structured data annotation and price prediction, build a commodity price prediction model, extract key information in real time and predict price trends.

Benefits of technology

Real-time structured conversion and efficient data transmission of telephone recordings are realized, market pricing transparency is improved, market pricing is quickly responding to market changes, real-time decision-making support is provided, information asymmetry is reduced, and trading market stability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410653A_ABST
    Figure CN120410653A_ABST
Patent Text Reader

Abstract

The invention discloses a bulk commodity transaction market stabilization method based on a BSL-CPFS deep learning model. The method comprises the following steps: preprocessing a voice data set into a training set and a test set; introducing a DeepSpeech2 model as a speech transcription model, and training the speech transcription model by using a training set and a text set; constructing a BSL model based on a BERT model and a TextCNN, and inputting the speech transcription text data training set for training to obtain a BSL language processing model; an NER model is introduced, the data set is labeled, and the NER model is input for training to form structured data; constructing a CPFS model based on a BiLSTM network, and inputting structured data for training to obtain a price prediction model; a prediction result and historical price fluctuation are dynamically compared to identify an abnormal fluctuation signal and push the abnormal fluctuation signal in real time, the asymmetry of information is reduced, a supervision department continuously monitors the deviation between the market price and the prediction result and formulates a targeted intervention strategy, then the price correction is realized, the stability of a bulk commodity transaction market is improved, and the market competitiveness is improved. And the market pricing transparency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and particularly to a method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model. Background Art

[0002] In the current trend of economic globalization, bulk commodities, as the basis of industrial and agricultural production and consumption, have the characteristics of large supply and demand volume and large price fluctuations. The current bulk commodity market faces significant problems of price premium and imbalance in trading efficiency. Information asymmetry and non-transparent trading behaviors exacerbate the risk of price deviation from the reasonable range. China is a major agricultural and manufacturing country and one of the main trading countries of bulk commodities. Therefore, it is urgent to solve such market structural problems to stabilize the economic and financial systems.

[0003] Traditional data acquisition methods mainly rely on structured data sources such as market quotations and news. However, with the acceleration of globalization and the diversification of communication means, more and more bulk commodity trading information is conveyed orally. Especially in telephone conversations, these telephone recordings may contain important information such as predictions of bulk commodity prices, market trends, and trading intentions. However, telephone recordings belong to unstructured data, and traditional data analysis methods are difficult to capture the implicit unpublicized trading intentions, implicit supply and demand signals, and market sentiment fluctuations in telephone conversations, resulting in the loss of key information and the lag of the price correction mechanism.

[0004] Most of the price prediction methods in China mainly adopt single or combined machine learning models. Although no model can directly predict prices by analyzing phone recordings yet, using text data for price prediction has become a popular research direction in recent years. Bing Guiying proposed a new data-driven hybrid model K-means-KPCA-KELM for predicting international copper prices, which uses online media text, Google Trends, and traditional economic data. The effective sequence is extracted through text analysis by the CNN-VMD method. The independent variable sequence is divided into K clusters by the K-means method. For each cluster, KPCA is used to reduce the data dimension, and the low-dimensional features are used as the input matrix of KELM to predict international copper prices. Jiang Feng et al. extracted the latent features in news texts through the Word2Vec method, fused the news text word vectors and the daily rise and fall information of crude oil futures using the MTAE network topology, and integrated the text features with crude oil-related indicators such as economic development, energy, and climate environment using long short-term memory neural networks to predict crude oil futures prices. However, these methods are applicable to text data with relatively systematic information organization, prominent key points, and clear paragraph structures, but it is difficult to achieve the expected results for text data with scattered and incoherent information. Foreign scholars started earlier in using text information for price prediction. Schumaker and Chen extracted the investor sentiment in financial news texts and input it into a support vector machine model to construct a stock price prediction model, and found that the prediction effect was good. Li J, Xu J, etc. used futures news titles to measure sentiment as a predictor to improve the prediction accuracy. Naima et al. classified events through machine learning and natural language processing technologies and used the Prophet model to predict freight rates, and the effect was better than the prediction results without considering events. The existing technologies focus on the research based on structured and semi-structured text data, but limited by the high-frequency uploading and inefficient transmission of plaintext data, it is difficult to meet the real-time parsing and pricing response requirements of the commodity market for unstructured voice information (such as phone recordings).

[0005] In summary, the core challenges in using phone recordings to crack the pricing problem in the commodity market are as follows:

[0006] (1) The bottleneck in the efficient parsing and transmission of unstructured voice data. Phone recordings contain a large number of colloquial expressions (such as pauses, repetitions, and informal terms), and traditional data processing technologies are difficult to achieve low-latency transmission and high-precision semantic parsing, resulting in low efficiency in extracting market implicit signals (such as supply and demand intentions, speculative behaviors) and exacerbating information asymmetry.

[0007] (2) The problem of associated modeling between fragmented information and pricing logic. The conversation topics are highly jumpy, and price driving factors (such as sudden events) are scattered in discontinuous contexts. Traditional models cannot quickly identify structured features from redundant content, restricting the improvement of market pricing transparency.

[0008] (3) The problem of data value mining under privacy compliance constraints. The requirements for encryption and desensitization of sensitive information increase the real-time processing complexity of voice data. Existing technologies are difficult to achieve rapid decoding and cross-institutional sharing of high-frequency trading signals (such as hoarding strategies and position adjustments) while ensuring compliance, hindering the dynamic response ability of the market pricing correction mechanism. Summary of the Invention

[0009] The present invention provides a method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model to overcome the technical problems of increased market information asymmetry, lagging price correction mechanism, and large fluctuations in the bulk commodity trading market due to low efficiency of unstructured data transmission, insufficient semantic parsing accuracy, and lagging information extraction.

[0010] To achieve the above object, the technical solution of the present invention is:

[0011] A method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model, comprising:

[0012] S1: Obtain a voice data set containing bulk commodity market conversations and a corresponding text set, preprocess the voice data set, and divide the preprocessed voice data set into a training set and a test set;

[0013] S2: Introduce the DeepSpeech2 model as a voice transcription model, and use the training set and the corresponding text set to train the DeepSpeech2 model to obtain an optimized voice transcription model and a voice transcription text data training set;

[0014] S3: Construct a BSL model based on the BERT model and TextCNN, input the voice transcription text data training set into the BSL model for training to obtain a trained BSL language processing model, and the BSL language processing model is used to perform semantic classification on the content in the voice transcription text data training set to obtain a classified statement data set;

[0015] S4: Introduce the NER model, annotate the classified statement data set, and input the annotated statement data set into the NER model for training to obtain an NER annotation model. The NER annotation model is used to identify the classified statement data set to obtain the recognition results of each category in the annotated statement data set, forming structured data including date, variety, price, and output.

[0016] S5: Construct a CPFS model based on the BiLSTM network, input the structured data into the CPFS model for training to obtain a CPFS price prediction model, which is used to predict the price of each variety according to the information in the structured data and the economic data of the corresponding variety, and obtain the future price of each variety;

[0017] S6: Connect the optimized speech transcription model, BSL language processing model, NER annotation model, and CPFS price prediction model in sequence to obtain a bulk commodity price prediction model. Input the test set into the bulk commodity price prediction model to classify and extract information from the recorded data of the test set, and predict the future price of bulk commodities based on the extracted information;

[0018] S7: Dynamically compare the prediction results output by the bulk commodity price prediction model with the historical price fluctuation threshold to identify abnormal fluctuation signals and push them to the regulatory agency and the enterprise side in real time to reduce information asymmetry. At the same time, the regulatory department formulates targeted intervention strategies by continuously monitoring the deviation between the market price and the prediction results, thereby achieving price correction and improving the stability of the bulk commodity trading market.

[0019] Further, construct a BSL model based on the BERT model and the TextCNN model. The BSL model includes an input module, a BERT model, a TextCNN model, a residual connection module, and a classification module;

[0020] The input module includes an input token masking layer and a Bert encoding layer connected in sequence;

[0021] The residual connection module includes a first linear transformation layer and a feature weighting layer connected in sequence;

[0022] The classification module includes a first Dropout layer, a second linear transformation layer, and an output logic layer connected in sequence.

[0023] Further, input the speech transcription text data training set into the BSL model for training, including:

[0024] S31: Input the speech transcription text data training set into the input module, and perform word segmentation, mask filling, and Bert encoding on the speech transcription text data through the input token masking layer and the Bert encoding layer to form a Token sequence that meets the input requirements of the BERT model, as shown in formula (1),

[0025] X input ={x1,x2,…,x n}∈R n(1)

[0026] In the formula, X input represents the index representation of the Token sequence that meets the input requirements of the BERT model after word segmentation, mask filling, and Bert encoding. {x1, x2, …, x n} represents the Token sequence of length n, and R n represents that this sequence is an integer index vector of length n; n represents the sequence length; R is the integer index vector;

[0027] S32. Input the Token sequence into the BERT model for semantic feature extraction to obtain a semantic vector sequence and a global semantic feature;

[0028] S33. Input the semantic vector sequence into the TextCNN model, and perform secondary feature extraction on the semantic vector sequence through the TextCNN model to obtain local convolutional features;

[0029] S34. Input the global semantic feature and the local convolutional feature into the residual connection module. Through the residual connection module, perform a linear transformation on the global semantic feature and perform weighted fusion with the local convolutional feature to obtain a fixed-length vector, as shown in formulas (2) and (3),

[0030] h′ cls = W cls h cls + b cls ∈ R 3f (2)

[0031] h fusion = h cnn + h′ cls (3)

[0032] Formula (2) represents performing a linear transformation on the global semantic feature to convert the original dimension of the global semantic feature to the same dimension as the output dimension of the TextCNN model. h′ cls represents the global semantic feature after linear transformation. W cls ∈ R 3f×d is the weight matrix of the first linear transformation layer, d is the original dimension; h cls represents the global semantic feature, b cls represents the bias vector, and 3f represents the output dimension of the TextCNN model;

[0033] Formula (3) represents fusing the local convolutional feature output by the TextCNN model with the globally semantically feature after linear transformation. h cnn represents the local convolutional feature output by the TextCNN model; h fusion represents the fixed-length vector obtained after fusion;

[0034] S35. Input the fixed-length vector into the first Dropout layer of the classification module for random inactivation processing, as shown in formula (4):

[0035] h drop = Dropout(h fusion ·p) (4)

[0036] h drop represents the high-dimensional global semantic features after regularization; p is the dropout rate;

[0037] Input h drop into the second linear transformation layer for linear transformation to obtain the unnormalized classification score, as shown in formula (5):

[0038] y logits = W out h drop + b out ∈R c (5)

[0039] y logits represents the unnormalized classification score, c is the number of classification categories; W out represents the weight matrix of the second linear transformation layer; b out is the bias vector;

[0040] Input the unnormalized classification score into the output logic layer, and normalize the classification score through Softmax, as shown in formula (6):

[0041] y pred = Softmax(y logits ) (6)

[0042] where, y pred represents the final class probability distribution, and each value represents the probability that the input belongs to that class;

[0043] Input y pred into the argmax function to obtain the class index corresponding to the maximum probability, that is, the final classification result, as shown in formula (7):

[0044] y pred_class = argmax(y pred ) (7)

[0045] y pred_class represents the class index with the maximum probability.

[0046] Further, a CPFS model is constructed based on the BiLSTM network. The CPFS model includes an input layer, an embedding layer, a BiLSTM network, an SE module, a fully connected layer module, and a Value layer connected in sequence;

[0047] The fully connected layer module includes a first fully connected layer, a second Dropout layer, and a second fully connected layer connected in sequence.

[0048] Further, structured data is input into the CPFS model for training, including:

[0049] S51. Obtain the time series data of a certain variety of commodity. The time series data includes the structured data and economic data of a certain variety of commodity; input the time series data of a certain variety of commodity into the input layer, and then input the time series data into the embedding layer through the input layer. The input time series data is mapped to the embedding space through the embedding layer to generate a feature vector with a dimension of d e as shown in formula (8),

[0050]

[0051] In the formula, Embedding represents the embedding operation, d e is the embedding dimension, E represents the feature vector after the input time series passes through the embedding layer processing, and is a tensor with a shape of (b, n, d e ) where b is the batch size and n is the sequence length; X represents the time series data, d in is the feature input dimension of each time step;

[0052] S52. Input the feature vector into the BiLSTM network, calculate the forward and backward hidden states respectively, and splice them to obtain the hidden state feature matrix;

[0053] S53. Input the spliced hidden state feature matrix into the SE module. The SE module performs adaptive weighted adjustment on the hidden state feature matrix through the global information compression and self-excitation mechanism, as shown in formulas (9) and (10),

[0054]

[0055] s = F ex (z) = σ(W2δ(W1z)) (10)

[0056] In formula (9), z represents the average value of the hidden state feature matrix H output by the BiLSTM network in the time dimension, that is, the global feature vector of the entire time series data, F sq(H) represents the operation of performing sequence global average pooling on the input hidden state matrix H, where n is the sequence length, and H i represents the hidden state at the i-th time step; d h represents the dimension of each hidden layer in the BiLSTM network;

[0057] In formula (10), s represents the weight vector of each channel calculated through the self-excitation mechanism, and F ex (z) represents the weighted processing of the feature channels of the compressed features through the self-excitation mechanism; where W1 and W2 are weight matrices, r is the dimensionality reduction ratio; δ represents the activation function;

[0058] The hidden state feature matrix is calibrated using the calculated weight vector of each channel, as shown in formula (11),

[0059]

[0060] represents element-wise multiplication, represents the recalibrated hidden state feature matrix;

[0061] S54. Input the recalibrated hidden state feature matrix into the fully connected layer module, and perform a non-linear transformation on through the first fully connected layer, as shown in formula (12),

[0062]

[0063] FC1 represents the feature matrix after the non-linear transformation of the recalibrated hidden state feature matrix through the ReLU activation function in the first fully connected layer, b1 is the bias term of the first fully connected layer, and W3 represents the weight matrix of the first fully connected layer; represents the output dimension of the first fully connected layer;

[0064] Input the non-linearly transformed feature matrix into the second fully connected layer through the second Dropout layer for linear transformation, as shown in formulas (13) and (14),

[0065] D1 = Dropout(FC1) (13)

[0066]

[0067] FC2 represents the price prediction result output by the second fully connected layer, and D1 represents the intermediate feature matrix after being processed by the second Dropout layer; b2 is the bias term of the second fully connected layer;

[0068] Input the price prediction result output by the second fully connected layer into the Value layer to calculate the MSE loss function, as shown in Formulas (15) and (16).

[0069]

[0070] represents the price prediction result output by the second fully connected layer represents calculating the error between the model prediction value and the true value using the MSE loss function; is the predicted value of the i-th sample; y i is the true value of the i-th sample; b represents the batch size.

[0071] Beneficial effects: The present invention provides a method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model, which has the following advantages:

[0072] (1) Construct the BSL model based on the BERT model and the TextCNN model, which can automatically extract key data such as price, production, and inventory through semantic classification and feature extraction methods, realize real-time structured conversion of telephone recordings and efficient data transmission, and accurately identify complex and diverse conversation contents in the bulk commodity market;

[0073] (2) Construct the CPFS model based on BiLSTM, directly extract the dynamics of the bulk commodity market from the text converted from telephone recordings, predict price trends, and improve market pricing transparency;

[0074] (3) Through comprehensive analysis of the extracted core elements, the system can predict price fluctuations in real time through the CPFS model, and provide price prediction results based on historical data and current supply and demand conditions in the market, compare with historical data, suppress irrational price premiums, and improve the stability of the bulk commodity trading market;

[0075] (4) The whole process realizes automatic processing without manual intervention. The entire process is fast from data extraction to price prediction, ensuring rapid response to market changes and providing real-time decision support. Description of the Drawings

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0077] Figure 1Flowchart of a method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model provided by the present invention;

[0078] Figure 2 Structural diagram of the DeepSpeech2 model;

[0079] Figure 3 Structural diagram of the BSL model designed by the present invention;

[0080] Figure 4 Flowchart of NER model data processing;

[0081] Figure 5 Structural diagram of the CPFS model designed by the present invention;

[0082] Figure 6 Comparison chart of the predicted value and the actual value of the metallurgical coke price obtained by the CPFS model;

[0083] Figure 7 Variation chart of the training loss and validation loss of the CPFS model;

[0084] Figure 8 Price prediction trend chart generated according to the metallurgical coke price prediction results. Detailed implementation method

[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0086] This embodiment provides a method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model, as Figure 1 shown, including:

[0087] S1: Obtain a speech data set containing bulk commodity market conversations and the corresponding text set, preprocess the speech data set, and divide the preprocessed speech data set into a training set and a test set;

[0088] S2: Introduce the DeepSpeech2 model as a speech transcription model, use the training set and the corresponding text set to train the DeepSpeech2 model, and obtain an optimized speech transcription model and a speech transcription text data training set;

[0089] S3: Build the BSL model based on the BERT model and TextCNN, input the speech transcription text data training set into the BSL model for training to obtain a trained BSL language processing model, which is used to perform semantic classification on the content in the speech transcription text data training set to obtain a classified sentence data set;

[0090] S4: Introduce the NER model, annotate the classified sentence data set, and input the annotated sentence data set into the NER model for training to obtain an NER annotation model, which is used to identify the classified sentence data set to obtain the recognition results of each category in the annotated sentence data set, forming structured data including date, variety, price, and output;

[0091] S5: Build the CPFS model based on the BiLSTM network, input the structured data into the CPFS model for training to obtain a CPFS price prediction model, which is used to predict the price of each variety according to the information in the structured data and the economic data of the corresponding variety to obtain the future price of each variety;

[0092] S6: Connect the optimized speech transcription model, BSL language processing model, NER annotation model, and CPFS price prediction model in sequence to obtain a bulk commodity price prediction model, input the test set into the bulk commodity price prediction model, classify and extract information from the recorded data of the test set, and predict the future price of bulk commodities based on the extracted information;

[0093] S7: Dynamically compare the prediction results output by the bulk commodity price prediction model with the historical price fluctuation threshold, identify abnormal fluctuation signals, and push them to the regulatory agency and the enterprise side in real time to reduce information asymmetry. At the same time, the regulatory department formulates targeted intervention strategies by continuously monitoring the deviation between the market price and the prediction results, thereby realizing price correction and improving the stability of the bulk commodity trading market.

[0094] Specifically, first, obtain a voice dataset containing conversations in the bulk commodity market and the corresponding text set. Preprocess the voice dataset and divide the preprocessed voice dataset into a training set and a test set. The collected telephone recording data is mainly used to train a speech transcription model. Obtain the corresponding text data of the recording as label data and optimize the speech transcription model. Introduce the DeepSpeech2 model as the speech transcription model and use the training set and the corresponding text set to train the DeepSpeech2 model to obtain an optimized speech transcription model and a training set of speech transcription text data. Using the DeepSpeech2 model can greatly simplify the process of speech recognition and improve the efficiency of speech recognition;

[0095] Secondly, construct a BSL model based on the BERT model and TextCNN. Input the training set of speech transcription text data into the BSL model for training to obtain a trained BSL language processing model. The BSL language processing model is used to perform semantic classification on the content in the training set of speech transcription text data to obtain a classified sentence dataset. By adding a residual layer after the BERT and TextCNN models, an efficient fusion of global semantic features and local convolutional features can be achieved. The CLS Token of BERT provides a sentence-level global semantic representation, while TextCNN is good at capturing local context patterns. The residual connection directly fuses the two, retaining both global semantic information and enhancing the expression of local features. The residual structure can alleviate the problem of gradient disappearance during the feature fusion process, avoid the model's over-reliance on local features, and thus improve the robustness and generalization ability of the model;

[0096] Introduce a NER model, annotate the classified sentence dataset, and input the annotated sentence dataset into the NER model for training to obtain a NER annotation model. The NER annotation model is used to identify each category in the annotated sentence dataset to obtain the recognition results of each category in the annotated sentence dataset, forming structured data including date, variety, price, and output. This solution uses SpaCy to construct a custom NER model to extract key information in sentences, which can accurately extract keywords in different types of sentences and improve classification accuracy;

[0097] Thirdly, construct a CPFS model based on the BiLSTM network. Input the structured data into the CPFS model for training to obtain a CPFS price prediction model. The CPFS price prediction model is used to predict the price of each variety according to the information in the structured data and the economic data of the corresponding variety to obtain the future price of each variety. Adaptive weighting is performed on the key features of the input data based on the traditional BiLSTM, thereby enhancing the feature selectivity and non-linear expression ability of the model;

[0098] Connect the optimized speech transcription model, BSL language processing model, NER annotation model, and CPFS price prediction model in sequence to obtain a bulk commodity price prediction model. Input the test set into the bulk commodity price prediction model to classify and extract information from the recorded data in the test set, and predict the future prices of bulk commodities based on the extracted information. It can achieve fully automated information extraction and price prediction of recorded data without manual intervention, ensure rapid response to market changes, provide real-time decision support, and provide high-precision price prediction results;

[0099] Finally, dynamically compare the prediction results output by the bulk commodity price prediction model with the historical price fluctuation threshold to identify abnormal fluctuation signals and push them to the regulatory authorities and the enterprise side in real time, reducing information asymmetry. At the same time, the regulatory department formulates targeted intervention strategies by continuously monitoring the deviation between the market price and the prediction results, thereby achieving price correction and improving the stability of the bulk commodity trading market.

[0100] In a specific embodiment, the scheme for obtaining a speech data set containing bulk commodity market conversations and the corresponding text set, preprocessing the speech data set, and dividing the preprocessed speech data set into a training set and a test set is as follows:

[0101] Collect 1,150 pieces of telephone recording data provided by Shanghai Steel Union E-commerce Co., Ltd. The data is in MP3 format and mainly involves conversations about the bulk commodity market, such as coal, iron ore, etc. The recording data presents price fluctuations, supply and demand relationships, and market sentiment in the bulk commodity market. Its main content includes the quality, transaction price, and production and inventory issues of different varieties of bulk commodities;

[0102] Since the training process also requires actual text data corresponding to the audio data as label data, use iFlytek dictation software to perform audio transcription text operations. Given that there are sometimes dialects in telephone recordings, which leads to a decrease in transcription accuracy, manual proofreading and modification are combined to ensure the accuracy of the transcribed text, forming a text set corresponding to the speech data set;

[0103] In the data cleaning stage, remove recordings with too much noise, background noise, or overlapping speech. Then, transcode the recording data into a WAV file with a sampling rate of 16 kHz, a duration of t seconds, and mono using FFmpeg. FFmpeg is a powerful multimedia processing tool that can perform audio and video recording, conversion, compression, editing, and media transmission. When processing text tags, remove unnecessary punctuation marks and convert them to a unified format.

[0104] The telephone recording data collected by the present invention is mainly used to train a speech recognition model. These data are only used for model training. Once the training is completed, the system can automatically extract key information from new telephone recordings without repeatedly collecting a large amount of recording data. During the actual application process, only by inputting the new recording data into the trained model can the relevant information of bulk commodities be automatically obtained and price prediction be carried out.

[0105] In a specific embodiment, the DeepSpeech2 model is introduced as a speech transcription model. The scheme for training the DeepSpeech2 model using the training set and the corresponding text set to obtain an optimized speech transcription model and a speech transcription text data training set is as follows:

[0106] The training process is as Figure 2 shown. First, the audio of the training set is converted into features suitable for model input - Mel Frequency Cepstral Coefficients (MFCC) before being input into the model. The features can effectively compress audio information while retaining the core information of the speech.

[0107] Secondly, the converted training set and the corresponding text set are input into the DeepSpeech2 model. The structure of the DeepSpeech2 model is as Figure 2 shown, including a 2D convolutional layer (Cov2D), a bidirectional RNN layer, a fully connected layer, a loss function layer, and a decoder.

[0108] In the model, first, a 3 - layer 2D convolutional layer (Cov2D) is used to pre - process the extracted audio features to capture local features in the speech.

[0109] Subsequently, the processed features are input into the bidirectional RNN layer. The 7 RNN layers understand the long - term dependencies in the audio sequence through the context information of the previous and subsequent time steps. The bidirectional RNN simultaneously utilizes the past and future information of the speech signal to enhance the ability to capture the time - dynamic changes in the speech.

[0110] The data output from the RNN layer passes through the fully connected layer and is further processed into the probability distribution of the character sequence.

[0111] Then, the CTC loss function is used to align the model output with the target text sequence, solving the problem of the inconsistency between the speech signal and the text length.

[0112] Finally, the decoder converts the output character probability distribution into the final text sequence.

[0113] During training, the Adam optimizer was used to adjust the weights of each neuron in the neural network, minimizing the loss function and ensuring that the model's predicted text was as close to the real text as possible. The training set was further divided into a training set, a validation set, and a test set, with a ratio of 8:1:1. The model was trained for 50 epochs (epoch=50) with a learning rate (LR) of 0.001. Within each epoch, the model processed the training data in batches of 32 (batch_size=32). All other parameters remained the default values. Finally, the character error rate (CER) was calculated to evaluate the model's performance on unseen data. The training result was a CER of 0.112335.

[0114] After training, the trained DeepSpeech2 model can recognize any new recording file. Taking metallurgical coke as an example, the audio content is converted into text in txt format, as shown in Table 1.

[0115] Table 1

[0116]

[0117] DeepSpeech2 is an end-to-end speech recognition model based on a deep neural network (RNN) architecture that mimics the human auditory process to achieve highly accurate speech recognition. Compared with traditional speech recognition systems, DeepSpeech2's model structure is simpler, requiring only a single neural network model to complete speech-to-text conversion. This end-to-end speech recognition technology can greatly simplify the speech recognition process and improve its efficiency.

[0118] In a specific embodiment, a BSL model is constructed based on the BERT model and TextCNN, and the speech transcription text data training set is input into the BSL model for training to obtain a trained BSL language processing model. The BSL language processing model is used to semantically classify the content in the speech transcription text data training set. The solution for obtaining the classified sentence data set is:

[0119] like Figure 3 As shown, the BSL model includes an input module, a BERT model, a TextCNN model, a residual connection module and a classification module;

[0120] The input module includes an input token mask layer (Input Token&mask) and a Bert encoding layer (Bert Encoder) connected in sequence;

[0121] The residual module includes a linear transformation layer (Linear Transform) and a feature weighting layer (Add) connected in sequence;

[0122] The classification module includes a first Dropout layer, a linear layer, and an output logic layer (OutputLogits) connected in sequence;

[0123] Training the speech transcription text data training set by inputting it into the BSL model includes:

[0124] S31. Input the speech transcription text data training set into the input module, and perform word segmentation, mask filling, and Bert encoding on the speech transcription text data through the input token mask layer and the Bert encoding layer to form a Token sequence that meets the input requirements of the BERT model, as shown in formula (17):

[0125] X input ={x1,x2,…,x n}∈R n (17)

[0126] In the formula, X input represents the index representation of the Token sequence that meets the input requirements of the BERT model after word segmentation, mask filling, and Bert encoding. {x1,x2,…,x n} represents the Token sequence of length n, and R n indicates that this sequence is an integer index vector of length n; n represents the sequence length; R is an integer index vector;

[0127] S32. Input the Token sequence into the BERT model for semantic feature extraction to obtain a semantic vector sequence and a global semantic feature, as shown in formulas (18) and (19):

[0128] H bert =Bert(X input )∈R n×d (18)

[0129] h cls =H bert [o]∈R d (19)

[0130] Formula (18) indicates that first, the multi-layer Transformer of the BERT model extracts features from the input Token sequence to generate corresponding embedding vectors, and then the Sequence Output layer aggregates all the embedding vectors; H bertRepresents the semantic representation of the token sequence extracted by the BERT model, that is, the semantic vector sequence, where d is the dimension of the hidden layer in the Transformer; the feed-forward network layer set by the model consists of 768 hidden neurons, that is, the hidden layer dimension is d = 768; in this solution, the BERT model used is the Bert-base-Chinese model;

[0131] Formula (19) indicates that the CLS token is used to record and is fixed at the start position of each input sequence to capture the global semantic features at the sentence level, h cls Represents the first feature vector output by the Sequence Output;

[0132] S33. Input the semantic vector sequence into the TextCNN model. The structure of the TextCNN model is as Figure 3 shown, including a Reshape Layer, three parallel convolutional layer modules, and a feature fusion layer. The convolutional layer module includes a convolutional layer, an activation function layer, and a max pooling layer connected in sequence;

[0133] Perform secondary feature extraction on the semantic vector sequence through the TextCNN model to obtain local convolutional features, as shown in formulas (20)-(23),

[0134] H reshape = Reshape(H bert ) ∈ R 1×n×d (20)

[0135] C k = ReLU(Conv2D(H reshape , W k )) ∈ R f×(n-k+1)×1 (21)

[0136] P k = MaxPool(C k ) ∈ R f (22)

[0137] h CNN = Concat(P2, P3, P4) ∈ R 3f (23)

[0138] Formula (20) indicates that H bert is reshaped through the Reshape Layer in the TextCNN model to add a channel dimension so that it can be input into the two-dimensional convolutional layer; a total of 3 layers of convolution are set for parallel processing, with 256 convolutional kernels for each layer of convolution, and different convolutional kernel sizes k ∈ {2, 3, 4} are used for each layer to capture n-gram features of different lengths; H reshapeDenote the reshaped feature vector;

[0139] Formula (21) indicates that the reshaped feature vector is convolved and activated through three parallel convolutional layer modules in the TextCNN model to obtain a feature map. f is the number of convolutional kernels, and W k is the parameter matrix of the convolutional kernel, k is the window size of the convolutional kernel (k = 2, 3, 4), and n - k + 1 is the length of the sequence after convolution; the ReLU activation function is used to perform a non-linear transformation on the output of the convolution to enhance the feature expression ability; C k represents the processed feature map;

[0140] Formula (22) indicates that the feature map C output by each convolutional kernel k is subjected to max pooling to extract the most significant value of each feature map along the time dimension (n - k + 1). The result P of the pooling k is a vector of size f; k

[0141] Formula (23) indicates that the pooling outputs P of different convolutional kernel sizes k are concatenated into a vector h CNN , to achieve the comprehensive representation of features of different granularities. h CNN is a vector of size 3f;

[0142] S34. Input the global semantic feature and the local convolutional feature into the residual connection module. Through the residual connection module, the global semantic feature is linearly transformed and weighted and fused with the local convolutional feature to obtain a fixed-length vector, as shown in Formulas (24) and (25).

[0143] h′ cls = W cls h cls + b cls ∈ R 3f (24)

[0144] h fusion = h cnn + h′ cls (25)

[0145] Formula (24) indicates that the global semantic feature is linearly transformed to convert the original dimension of the global semantic feature to the same dimension as the output dimension of the TextCNN model. h′ cls represents the globally semantic feature after linear transformation. W cls ∈ R 3f×d is the weight matrix of the fully connected layer; h cls represents the globally semantic feature, and b cls represents. 3f represents the output dimension of the TextCNN model;

[0146] Equation (25) represents the feature fusion of the local convolutional features output by the TextCNN model and the globally semantic features after linear transformation. h cnn represents the local convolutional features output by the TextCNN model; h fusion represents the fixed-length vector obtained after fusion;

[0147] S35. Input the fixed-length vector into the first Dropout layer of the classification module for random inactivation processing, as shown in Equation (26),

[0148] h drop = Dropout(h fusion ·p) (26)

[0149] h drop represents the high-dimensional globally semantic features after regularization; p is the dropout rate;

[0150] Input h drop into the second linear transformation layer for linear transformation to obtain the unnormalized classification scores, as shown in Equation (27),

[0151] y logits = W out h drop + b out ∈ R c (27)

[0152] y logits represents the unnormalized classification scores, c is the number of classification categories; W out represents the weight matrix of the second linear transformation layer; b out is the bias vector;

[0153] Input the unnormalized classification scores into the output logic layer, and normalize the classification scores through Softmax, as shown in Equation (28),

[0154] y pred = Softmax(y logits ) (28)

[0155] where, y pred represents the final class probability distribution, where each value represents the probability that the input belongs to that class;

[0156] Input y pred into the argmax function to obtain the class index corresponding to the maximum probability, that is, the final classification result, as shown in Equation (29),

[0157] y pred_class = argmax(y pred ) (29)

[0158] y pred_class represents the class index with the highest probability.

[0159] In this embodiment, taking metallurgical coke as an example, 1960 text sentences transcribed by DeepSpeech are extracted and spliced into a txt file as training data. Since there are many colloquial expressions in the phone calls, it is necessary to preprocess the data first. Replace some colloquial names and inaccurate words with accurate variety words. As shown in Table 2, manually label all sentences. Label the sentences related to "variety" as the "variety" class, the sentences containing "price" data as the "price" class, the sentences containing "output" data as the "output" class, and the sentences that do not contain "variety", "price", or "output" as "irrelevant sentences". Input the data and data labels into BSL for training;

[0160] Table 2

[0161] Colloquial Name Accurate Variety Term after Replacement Prime Coking Coal Prime Coking Coal Mixed Alkali / Alkali Mixture Alkali Mixed Coal Ultra Special Ultra Special Powder PP / Bip Powder PP Powder Iron Ore Iron Ore Pellet Pellet …… ……

[0162] Use the Adam optimization algorithm with adaptive learning rate to optimize the model, and cross-entropy loss as the loss function. The classification categories used in the training process are four categories (NUM_CLASSES = 4), namely: variety, price, output, and irrelevant sentences; Divide the input data into BSL training set, BSL validation set, and BSL test set, with a ratio of 8:1:1; The maximum text length is 73, the number of training epochs Epoch = 50, the learning rate is set to 0.001, the batch size is 10 (batch_size = 10). In each epoch, the model processes the training data in batches, batch_size = 10, and other parameters are default training parameters.

[0163] After 50 rounds of training, the model achieved significant results on the test set. The classification accuracy reached 0.90, and the macro-average and weighted-average F1 scores also both reached 0.90. This indicates that the model has strong generalization ability and robustness, and can accurately classify sentences. The classification results are shown in Table 3,

[0164] Table 3

[0165]

[0166]

[0167] In this solution, by adding a residual layer after the BERT and TextCNN models, efficient fusion of global semantic features and local convolutional features can be achieved. The CLS Token of BERT provides a sentence-level global semantic representation, while TextCNN is good at capturing local context patterns. The residual connection directly fuses the two, retaining both global semantic information and enhancing the expression of local features. In addition, the residual structure can alleviate the problem of vanishing gradients during the feature fusion process, avoid the model's over-reliance on local features, and thus improve the robustness and generalization ability of the model.

[0168] In a specific embodiment, an NER model is introduced. The classified sentence dataset is annotated, and the annotated sentence dataset is input into the NER model for training to obtain an NER annotation model. The NER annotation model is used to identify the annotated sentence dataset to obtain the recognition results of each category in the annotated sentence dataset, and the solution for forming structured data including date, variety, price, and output is as follows:

[0169] Use SpaCy to build a custom NER model to extract key information in sentences. SpaCy is an open-source natural language processing library mainly used for text analysis and provides efficient text preprocessing tools such as word segmentation, named entity recognition (NER), part-of-speech tagging, dependency parsing, etc.

[0170] Create entity categories "Category" and "Num", manually annotate all training data, and provide the data to the NER model. The data is divided into an NER training set, an NER validation set, and an NER test set in a ratio of 8:1:1, and all other model parameters are default. The Token-level metric precision of the training result is 100%, indicating that the model has no errors in the word segmentation and annotation tasks. In terms of named entity recognition, the precision (P) of the model is 0.9309, the recall (R) is 0.8817, and the F1-score is 0.9057, indicating that the model can accurately identify all named entities. In the evaluation by entity type, the F1-score of the model for the variety (Category) entity is 0.9181; the F1-score of the quantifier (Num) entity is 0.8933.

[0171] Taking metallurgical coke as an example, the specific processing steps are as Figure 4As shown, first, the training data is manually labeled and can be classified into types such as variety, price, and yield. A custom NER model is constructed using the SpaCy database to identify the input data. The data obtained through the NER model still needs to be cleaned. Duplicate removal is performed on the variety, price, and yield data. The cn2an library is used to convert Chinese numerals in price and yield into Arabic numerals, and noise data less than 10 and redundant Chinese characters that cannot be converted into numerals are removed. Since a conversation between two people on the phone may involve key information of different commodities, the data is then segmented according to different varieties and arranged in chronological order. For the data of a certain date, if there are multiple similar quantifiers for the price and yield of a certain variety, the average value needs to be taken for processing. Then, all the data of a certain variety is traversed. If there is a data missing situation, linear interpolation is used to supplement it. After the above processing, structured data of "date: variety: price: yield" can be obtained, and the result is shown in Table 4, Table 4

[0172] Date Category Price / CNY Production / 1000kg 20230502 Metallurgical Coke 2200 500 20230503 Metallurgical Coke 2200 800 20230504 Metallurgical Coke 2200 1000 20230505 Metallurgical Coke 2150 1000 20230506 Metallurgical Coke 2150 1000 20230507 Metallurgical Coke 2100 1000 …… …… …… ……

[0173] In this solution, for the already classified dialogue sentences, it is necessary to extract the variety nouns from the "variety" type sentences and extract the corresponding quantifiers from the "price" and "yield" type sentences. Since ordinary word segmentation methods cannot perfectly achieve this goal, this solution uses SpaCy to construct a custom NER model to extract the key information in the sentences.

[0174] In a specific embodiment, a CPFS model is constructed based on the BiLSTM network. The structured data is input into the CPFS model for training to obtain a CPFS price prediction model. The CPFS price prediction model is used to predict the price of each variety according to the information in the structured data and the economic data of the corresponding variety, and the solution for obtaining the future price of each variety is as follows:

[0175] As Figure 5 shown, the CPFS model includes an input layer, an embedding layer, a BiLSTM network, an SE module, a fully connected layer module, and a Value layer connected in sequence;

[0176] The fully connected layer module includes a first fully connected layer, a second Dropout layer, and a second fully connected layer connected in sequence;

[0177] Inputting the structured data into the CPFS model for training includes:

[0178] S51. Obtain the time series data of a certain variety of commodity. The time series data includes the structured data and economic data of a certain variety of commodity. The structured data is specifically "date: price: output", and the economic data is the data representing macroeconomic variables, including the 1-year interest rate of the People's Bank of China, the macroeconomic prosperity index, and the US dollar-renminbi exchange rate. Therefore, the final time series data should be "date: price: output: 1-year interest rate of the People's Bank of China: macroeconomic prosperity index: US dollar-renminbi exchange rate".

[0179] Input the time series data of a certain variety into the input layer, and then input the time series data into the embedding layer through the input layer. The input time series data is mapped to the embedding space through the embedding layer to generate a feature vector with a dimension of d e as shown in formula (30).

[0180]

[0181] In the formula, Embedding represents the embedding operation, and d e is the embedding dimension, E represents the feature vector after the input time series is processed by the embedding layer, and it is a tensor with a shape of (b, n, d e ). b is the batch size, and n is the sequence length. X represents the time series data, and d in is the feature input dimension of each time step.

[0182] Input the feature vector into the BiLSTM network, calculate the forward and backward hidden states respectively, which are used to capture the forward and backward temporal information of the time series and are concatenated to obtain the hidden state feature matrix. The forward LSTM processes the information of the forward time step, as shown in formula (31).

[0183]

[0184] In the formula, represents the hidden state of the LSTM at the previous time step t - 1, and e t represents the input vector at time step t, represents the internal state of the LSTM at time step t after processing e t , which contains the historical information from time step 1 to t, and d h is the output dimension.

[0185] The backward LSTM processes the information from the backward time step, as shown in formula (32).

[0186]

[0187] Concatenate the information output by the LSTM in both directions, as shown in formula (33).

[0188]

[0189] H represents the hidden state matrix calculated by the bidirectional LSTM at all time steps t, with a shape of (b, n, 2d n );

[0190] S53. Input the concatenated hidden state feature matrix into the SE module. The SE module adaptively weights and adjusts the hidden state feature matrix through global information compression and self-excitation mechanism, as shown in Formulas (34) and (35).

[0191]

[0192] s = F ex (z) = σ(W2δ(W1z)) (35)

[0193] In Formula (34), z represents the average value of the hidden state feature matrix H output by the BiLSTM network in the time dimension, that is, the global feature vector of the entire time series data. F sq (H) represents performing a sequence global average pooling operation on the input hidden state matrix H. n is the sequence length, and H i represents the hidden state at the i-th time step; d h represents the dimension of each hidden layer in the BiLSTM network;

[0194] In Formula (31), s represents the weight vector of each channel calculated through the self-excitation mechanism. F ex (z) represents weighting the feature channels of the compressed features through the self-excitation mechanism. Among them, W1 and W2 are weight matrices, r is the dimensionality reduction ratio; δ represents the activation function;

[0195] Calibrate the hidden state feature matrix using the calculated weight vector of each channel, as shown in Formula (36).

[0196]

[0197] S54. Input the recalibrated hidden state feature matrix into the fully connected layer module. Perform a non-linear transformation on it through the first fully connected layer, as shown in Formula (37). as shown in Formula (37).

[0198]

[0199] FC1 represents the feature matrix after the non - linear transformation of the re - calibrated hidden - state feature matrix through the ReLU activation function in the first fully - connected layer, b1 is the bias term of the first fully - connected layer, and W3 represents the weight matrix of the first fully - connected layer; represents the output dimension of the first fully - connected layer;

[0200] The feature matrix after non - linear transformation is input into the second fully - connected layer through the second Dropout layer for linear transformation, as shown in formulas (38) and (39),

[0201] D1 = Dropout(FC1) (38)

[0202]

[0203] FC2 represents the price prediction result output by the second fully - connected layer, D1 represents the intermediate feature matrix after being processed by the second Dropout layer; b2 is the bias term of the second fully - connected layer;

[0204] The price prediction result output by the second fully - connected layer is input into the Value layer to calculate the MSE loss function, as shown in formulas (40) and (41),

[0205]

[0206] represents the price prediction result output by the second fully - connected layer, represents calculating the error between the model prediction value and the true value using the MSE loss function; is the predicted value of the i - th sample; y i is the true value of the i - th sample; b represents the batch size. The prediction target is the true price data of this variety, and the structured data should be "date: true price".

[0207] The input data of the CPFS model is the "date: price: output" structured data of a certain variety of bulk commodities obtained through the DeepSpeech2 model, BSL model, and Ner model, as well as the true price data of this variety within the corresponding time range, and other relevant economic data (RMB interest rate, macro - economic prosperity index, RMB exchange rate). In the training parameters of the model, the time step is set to 30, that is, the model uses 30 consecutive days of data to predict the price on the 31st day. The model uses the Adam optimizer and the mean squared error (MSE) as the loss function. The ratio of training data, validation data, and test data is 8:1:1, the batch size (batch_size) is set to 64, and the number of training epochs is set to 50;

[0208] In this solution, by introducing the Squeeze-and-Excitation (SE) module, the key features of the input data are adaptively weighted on the basis of the traditional BiLSTM, thereby enhancing the feature selectivity and non-linear expression ability of the model. The core task of this model is to perform multi-feature prediction on the price of each variety involved in the text one by one, mainly relying on the historical price, yield of the variety, and other relevant economic data (RMB interest rate, macroeconomic prosperity index, RMB exchange rate), and the obtained prediction results have high accuracy.

[0209] In a specific embodiment, the optimized speech transcription model, BSL language processing model, NER annotation model, and CPFS price prediction model are connected in sequence to obtain a bulk commodity price prediction model. The solution of inputting the test set into the bulk commodity price prediction model, classifying and extracting information from the recorded data of the test set, and predicting the future price of bulk commodities based on the extracted information is as follows:

[0210] Connecting the speech transcription model, BSL language processing model, NER annotation model, and CPFS price prediction model in sequence to obtain a bulk commodity price prediction model can achieve the information extraction and price prediction of fully automated recorded data. Inputting a large amount of recorded data in the test set into this combined network can obtain results as shown in Figure 6 and Figure 7 shown (taking the price prediction result of metallurgical coke as an example);

[0211] Figure 6 It shows the comparison between the metallurgical coke price prediction value obtained by the CPFS model and the actual value. The blue line represents the prediction value of the model on the training set. The yellow line represents the prediction value of the model on the test set. The green line represents the actual value (true value) of the price, that is, the target that the model needs to predict.

[0212] Figure 7 It shows the changes in the training loss and validation loss of the CPFS model. The model uses the mean square error (MSE) as the loss function. The blue line represents the training loss (Train Loss). The yellow line represents the validation loss (Validation Loss). Both the training loss and the validation loss gradually decrease and tend to be stable and convergent, indicating that the model has good performance.

[0213] Judging from the prediction results, both the training set loss (decreasing from 0.1747 to 0.0025) and the validation set loss (decreasing from 0.4880 to 0.0037) continue to decline, and the accuracy rate in the 50th round can reach 80%. Generally speaking, the model can fit the data well and maintain stable performance on unseen data.

[0214] Figure 8It is a price prediction trend chart generated based on price prediction results, which graphically presents historical prices and prediction data to help users quickly identify market trends and conduct commodity trading.

[0215] In a specific embodiment, the solution to dynamically compare the prediction results output by the bulk commodity price prediction model with the historical price fluctuation threshold, identify abnormal fluctuation signals, and push them to the regulatory agency and the enterprise side in real time to reduce information asymmetry. At the same time, the regulatory department formulates targeted intervention strategies by continuously monitoring the deviation between the market price and the prediction results, thereby achieving price correction and improving the stability of the bulk commodity trading market is as follows:

[0216] For the price trend output by the bulk commodity price prediction model, dynamically compare the prediction results with the historical price fluctuation threshold, identify abnormal fluctuation signals (such as speculative hoarding, regional supply-demand imbalance), and push them to the regulatory agency and the enterprise side in real time through a data interface, significantly reducing information asymmetry;

[0217] Investors optimize their trading decisions based on the pushed information, avoid irrational chasing and selling behaviors, thereby suppressing the spread of price bubbles;

[0218] The regulatory department can quickly and effectively formulate targeted intervention strategies by continuously monitoring the deviation between the market price and the predicted value. This mechanism effectively blocks the transmission path of abnormal price fluctuations by enhancing the response speed and accuracy of market correction, promotes the transformation of the bulk commodity market from passive response to active regulation, and ultimately realizes the high efficiency of the price discovery mechanism and the long-term stability of the market.

[0219] Through comprehensive analysis of the extracted core elements, the bulk commodity price prediction model enables the system to achieve real-time structured conversion of telephone recordings and efficient data transmission, predict price fluctuations in real time, and provide price prediction results based on the historical data and current supply-demand situation of the market to help investors make decisions and suppress irrational price premiums; at the same time, this model can rely on the prediction terminal to achieve automated processing without manual intervention. The entire process from data extraction to price prediction is fast, ensuring a quick response to market changes, providing real-time decision support, and enhancing market pricing transparency.

[0220] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model, characterized in that, Including: S1: Obtain a voice dataset containing conversations in the bulk commodity market and the corresponding text set, preprocess the voice dataset, and divide the preprocessed voice dataset into a training set and a test set; S2: Introduce the DeepSpeech2 model as a speech transcription model, use the training set and the corresponding text set to train the DeepSpeech2 model, and obtain an optimized speech transcription model and a speech transcription text data training set; S3: Construct a BSL model based on the BERT model and TextCNN, input the speech transcription text data training set into the BSL model for training, and obtain a trained BSL language processing model. The BSL language processing model is used to perform semantic classification on the content in the speech transcription text data training set to obtain a classified statement dataset; S4: Introduce the NER model, annotate the classified statement dataset, and input the annotated statement dataset into the NER model for training to obtain an NER annotation model. The NER annotation model is used to identify the classified statement dataset to obtain the recognition results of each category in the annotated statement dataset, forming structured data including date, variety, price, and output; S5: Construct a CPFS model based on the BiLSTM network, input the structured data into the CPFS model for training, and obtain a CPFS price prediction model. The CPFS price prediction model is used to predict the price of each variety according to the information in the structured data and the economic data of the corresponding variety to obtain the future price of each variety; S6: Connect the optimized speech transcription model, BSL language processing model, NER annotation model, and CPFS price prediction model in sequence to obtain a bulk commodity price prediction model. Input the test set into the bulk commodity price prediction model to classify and extract information from the recorded data of the test set, and predict the future price of bulk commodities based on the extracted information; S7: Dynamically compare the prediction results output by the bulk commodity price prediction model with the historical price fluctuation threshold, identify abnormal fluctuation signals, and push them to the regulatory agency and the enterprise side in real time to reduce information asymmetry. At the same time, the regulatory department formulates targeted intervention strategies by continuously monitoring the deviation between the market price and the prediction results, thereby realizing price correction and improving the stability of the bulk commodity trading market.

2. A method for stabilizing the commodity trading market based on the BSL-CPFS deep learning model according to claim 1, characterized in that Construct a BSL model based on the BERT model and TextCNN model. The BSL model includes an input module, a BERT model, a TextCNN model, a residual connection module, and a classification module; The input module includes an input token masking layer and a Bert encoding layer connected in sequence; The residual connection module includes a first linear transformation layer and a feature weighting layer connected in sequence; The classification module includes a first Dropout layer, a second linear transformation layer, and an output logic layer connected in sequence.

3. A method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model according to claim 2, characterized in that Inputting the speech transcription text data training set into the BSL model for training includes: S31. Input the speech transcription text data training set into the input module, and perform word segmentation, mask filling, and Bert encoding on the speech transcription text data through the input token mask layer and the Bert encoding layer to form a Token sequence that meets the input requirements of the BERT model, as shown in formula (1). X input = {x1, x2, …, x n} ∈ R n (1) where X input represents the index representation of the Token sequence that meets the input requirements of the BERT model after word segmentation, mask filling, and Bert encoding, and {x1, x2, …, x n} represents the Token sequence of length n, and R n represents that this sequence is an integer index vector of length n; n represents the sequence length; R is an integer index vector; S32. Input the Token sequence into the BERT model for semantic feature extraction to obtain a semantic vector sequence and global semantic features. S33. Input the semantic vector sequence into the TextCNN model, and perform secondary feature extraction on the semantic vector sequence through the TextCNN model to obtain local convolutional features. S34. Input the global semantic features and local convolutional features into the residual connection module. Through the residual connection module, perform a linear transformation on the global semantic features and perform weighted fusion with the local convolutional features to obtain a fixed-length vector, as shown in formulas (2) and (3). h′ cls = W cls h cls + b cls ∈ R 3f (2) h fusion = h cnn + h' cls (3) Equation (2) represents a linear transformation of the global semantic features, converting the original dimension of the global semantic features to the same dimension as the output dimension of the TextCNN model, h′ cls represents the global semantic features after the linear transformation, W cls ∈R 3f×d is the weight matrix of the first linear transformation layer, d is the original dimension; h cls represents the global semantic features, b cls represents the bias vector, 3f represents the output dimension of the TextCNN model; Formula (3) represents the feature fusion of the local convolutional features output by the TextCNN model and the globally semantic features after linear transformation, where h cnn represents the local convolutional features output by the TextCNN model; h fusion represents the fixed-length vector obtained after fusion; S35. Input the fixed-length vector into the first Dropout layer of the classification module for random inactivation processing, as shown in formula (4). h drop = Dropout(h fusion ·p) (4) h drop represents the high-dimensional global semantic features after regularization; p is the dropout rate; Input h drop into a second linear transformation layer for linear transformation to obtain unnormalized classification scores, as shown in formula (5). y logits = W out h drop + b out ∈ R c (5) y logits represents the unnormalized classification score, where c is the number of classification categories; W out represents the weight matrix of the second linear transformation layer; b out is the bias vector; Input the unnormalized classification scores into the output logic layer, and normalize the classification scores through Softmax, as shown in formula (6). y pred = Softmax(y logits ) (6) where y pred represents the final class probability distribution, where each value represents the probability that the input belongs to that class; Input y pred Input it into the argmax function to obtain the class index corresponding to the maximum probability, which is the final classification result, as shown in formula (7). y pred_class = argmax(y pred ) (7) y pred_class Represents the class index with the highest probability.

4. A method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model according to claim 1, characterized in that, Construct a CPFS model based on the BiLSTM network. The CPFS model includes an input layer, an embedding layer, a BiLSTM network, an SE module, a fully connected layer module, and a Value layer connected in sequence. The fully connected layer module includes a first fully connected layer, a second Dropout layer, and a second fully connected layer connected in sequence.

5. A method for stabilizing the bulk commodity trading market based on the BSL-CPFS deep learning model according to claim 4, characterized in that Input the structured data into the CPFS model for training, including: S51. Obtain the time series data of a certain variety of commodity, where the time series data includes the structured data and economic data of the certain variety of commodity; input the time series data of the certain variety of commodity into the input layer, and then input the time series data into the embedding layer through the input layer. The embedding layer maps the input time series data into the embedding space to generate a feature vector with a dimension of d e as shown in formula (8). Where, Embedding represents the embedding operation, and d e is the embedding dimension, E represents the feature vector after the input time series is processed by the embedding layer, and it is a tensor with the shape of (b, n, d e ), where b is the batch size and n is the sequence length; X represents the time series data, and d in is the feature input dimension at each time step; S52. Input the feature vector into the BiLSTM network, calculate the forward and backward hidden states respectively, and splice them to obtain a hidden state feature matrix. S53. Input the spliced hidden state feature matrix into the SE module. The SE module performs adaptive weighted adjustment on the hidden state feature matrix through the global information compression and self-excitation mechanism, as shown in formulas (9) and (10). s = F ex (z) = σ(W2δ(W1z)) (10) In Equation (9), z represents the average value of the hidden state feature matrix H output by the BiLSTM network in the time dimension, that is, the global feature vector of the entire time series data, F sq (H) represents performing a sequence global average pooling operation on the input hidden state matrix H, n is the sequence length, H i represents the hidden state at the i-th time step; d h represents the dimension of each hidden layer in the BiLSTM network; In formula (10), s represents the weight vector of each channel calculated by the self-excitation mechanism, F ex (z) represents the weighted processing of the feature channel of the compressed feature through the self-excitation mechanism; where W1 and W2 are weight matrices, r is the dimensionality reduction ratio; δ represents the activation function; Calibrate the hidden state feature matrix using the calculated weight vector for each channel, as shown in formula (11). denotes element-wise multiplication, represents the recalibrated hidden state feature matrix; S54. Input the recalibrated hidden state feature matrix into the fully connected layer module, and perform a non-linear transformation on through the first fully connected layer, as shown in formula (12). FC1 represents the feature matrix after the non-linear transformation of the re-calibrated hidden state feature matrix through the ReLU activation function in the first fully connected layer, b1 is the bias term of the first fully connected layer, and W3 represents the weight matrix of the first fully connected layer; represents the output dimension of the first fully connected layer; Input the non-linearly transformed feature matrix into the second fully connected layer through the second Dropout layer for linear transformation, as shown in formulas (13) and (14). D1 = Dropout(FC1) (13) FC2 represents the price prediction result output by the second fully connected layer, D1 represents the intermediate feature matrix after being processed by the second Dropout layer; b2 is the bias term of the second fully connected layer. Input the price prediction result output by the second fully connected layer into the Value layer to calculate the MSE loss function, as shown in formulas (15) and (16). Represents the price prediction result output by the second fully connected layer, represents the error between the model prediction value and the true value calculated using the MSE loss function; is the predicted value of the i-th sample; y i is the true value of the i-th sample; b represents the batch size.

Citation Information

Cited By

  • Information extraction method for bulk commodity market investigation voice

    CN121789686A