Multi-mode social public opinion perception demand prediction system and method
Through the multi-modal social public opinion perception system, effective extraction and bidirectional conversion of text and image modal features are achieved, the problem of insufficient integration of multi-modal data is solved, information utilization rate and trend prediction accuracy are improved, and adaptability to market dynamic changes is enhanced, and more reliable product demand prediction results are generated.
Patent Information
- Application Number
- CN202510617605.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing technology is difficult to effectively integrate multimodal data on social media, especially in an environment with pictures and texts, and it is impossible to fully capture user preferences and market trends, resulting in insufficient sensitivity to commodity sales trends, and the modal fusion method fails to deeply explore the semantic relationships between different modalities and insufficient information utilization.
The demand prediction system of multi-modal social public sentiment perception is adopted. Through the multi-modal-text-graphic emotion fusion module, preprocessing module, demand perception module and expert system, combined with dual-channel feature extraction, cross-modal representation learning, timing sensitive fusion algorithm and decision theory, the effective extraction and bidirectional conversion of text and image modal features are realized, and timing information is fused to generate multi-modal-text-graphic fusion vectors for product demand prediction.
The information utilization rate has been improved by 35%-50%, the modal alignment accuracy has been improved by 40%, the trend prediction accuracy has been improved by 20%, the prediction results have been more reliable and interpretable, and the comprehensive prediction accuracy has been improved by 15%-25%.
Smart Images

Figure CN120509928A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of big data analysis and artificial intelligence, and specifically to a social public opinion perception and commodity demand forecasting system and method based on multimodal data, which is used to integrate multimodal data such as text and images on social media to accurately predict commodity sales trends. Background Art
[0002] With the rapid development of social commerce and the increasing popularity of new sales models such as livestreaming and short video marketing, massive amounts of social media data have become a crucial source of information for analyzing market trends and consumer demand. Traditional demand forecasting methods primarily rely on historical sales data and single-modal user review analysis, which suffers from the following significant shortcomings: First, single-modal data analysis cannot fully capture user preferences and market trends, especially in the context of social media, which is characterized by both text and images. Second, traditional methods struggle to effectively integrate time series information with multimodal data, resulting in insufficient sensitivity to market dynamics. Furthermore, existing modal fusion methods often rely on simple concatenation or weighted averaging, failing to deeply explore the semantic connections between different modalities and resulting in inadequate information utilization.
[0003] In general, the existing technology has not yet proposed a system and method that can effectively integrate multimodal social public opinion data and achieve accurate demand forecasting. Summary of the Invention
[0004] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a demand prediction system and method for multimodal social public opinion perception, which can effectively integrate text and image modal data, combine time series information, and achieve accurate prediction of commodity demand.
[0005] The present invention proposes a demand forecasting system for multimodal social public opinion perception, including:
[0006] The multimodal text-image sentiment fusion module receives raw text and image data from social platforms, processes them through dual-channel feature extraction to generate text feature vectors and image feature vectors, and then performs cross-modal representation learning to fuse the text feature vectors and image feature vectors into a multimodal text-image sentiment vector.
[0007] The preprocessing module is connected to the multimodal-text-image emotion fusion module and is used to receive the multimodal-text-image emotion vector output by the multimodal-text-image emotion fusion module, quantify the live broadcast activity data and historical product sales through the time dimension, quantify the live broadcast influence through the frequency dimension, and use a recurrent neural network to fuse the quantified data with the original text data to form a time-sensitive multimodal-text-image emotion fusion representation;
[0008] The demand perception module is connected to the preprocessing module and is used to receive the multimodal text-image emotion fusion representation output by the preprocessing module. It then performs adaptive feature conversion on the received fusion representation through a multi-layer perceptron to generate a multimodal text-image fusion vector of historical sales and real-time fan emotions.
[0009] The expert system is connected to the demand perception module and is used to receive the multimodal text-image fusion vector output by the demand perception module, combine sales data, tag type and number of fans information, predict product demand based on decision theory, and output the prediction results.
[0010] As a preferred option, the multimodal text-image emotion fusion module includes:
[0011] The text input layer is used to receive raw text data, encode the input text using a pre-trained language model, and generate a text feature vector;
[0012] The image input layer is used to receive raw image data, encode the input image using a visual language model, and generate an image feature vector;
[0013] A feature mapping unit is used to map text feature vectors to image feature space or to map image feature vectors to text feature space, thereby achieving bidirectional cross-modal feature conversion.
[0014] The sentiment fusion unit is used to fuse the text feature vector and the image feature vector to form a multimodal text-image sentiment vector, and convert the fusion vector into a sentiment classification vector through nonlinear transformation.
[0015] Preferably, the feature mapping unit includes:
[0016] The text-to-image mapping channel is used to convert text feature vectors into image feature space representations through a multi-layer mapping network;
[0017] The image-to-text mapping channel is used to convert the image feature vector into a text feature space representation through an inverse mapping network;
[0018] The feature calibration network is used to compare the differences between the original features and the mapped features, retain the key information of the original features through residual connections, correct the mapping errors, and ensure the consistency and accuracy of the bidirectional mapping.
[0019] Preferably, the preprocessing module includes:
[0020] The time dimension quantification unit is used to divide live broadcast activity data and historical product sales into time windows, set a time decay function, extract time series features, and generate a time dimension indicator vector;
[0021] Frequency dimension quantification unit, used to count the frequency of user interaction behaviors, design a frequency influence scoring mechanism, consider interaction quality and user influence factors, and generate a frequency dimension vector;
[0022] The time series fusion unit is used to process time series data using a recurrent neural network, extract long-term and short-term dependencies, identify key time points and change patterns, fuse the time dimension indicator vector, frequency dimension vector and multimodal-text and image sentiment vector to generate a time-sensitive fusion representation.
[0023] Preferably, the demand perception module includes:
[0024] The feature conversion network, which includes a multi-layer nonlinear transformation structure, is used to decompose the time-sensitive fusion representation into historical sales features and real-time sentiment features. It performs nonlinear feature transformations on different feature types to generate optimized feature representations.
[0025] Dynamic activation unit, used to select the optimal activation function based on feature distribution characteristics, set the gated activation mechanism to process different types of features, adaptively adjust the activation function parameters, and enhance the model's ability to express complex features;
[0026] The adaptive adjustment unit is used to dynamically adjust network parameters based on the performance of the validation set, set performance thresholds to trigger the adjustment process, adopt a gradual adjustment strategy to avoid excessive modifications, and establish adjustment history records for long-term optimization.
[0027] Preferably, the expert system includes:
[0028] The data integration unit is used to integrate multimodal text-image fusion vectors with sales data, product attributes and label information, integrate user portraits and fan count data, and combine them with real-time social public opinion analysis results to build a multi-dimensional decision input matrix;
[0029] The decision tree construction unit is used to design a feature importance evaluation mechanism, use information gain to guide tree structure construction, set tree depth and complexity control parameters, and build an explainable decision path;
[0030] The forecast optimization unit is used to define the impact ratio of historical sales and real-time sentiment, design a dynamic adjustment mechanism for the ratio, optimize the ratio parameters based on historical forecast errors, set the upper and lower limits of the forecast results, and integrate the results of multiple forecast models to generate more robust integrated forecast results.
[0031] Preferably, the system further comprises:
[0032] The data collection module is used to collect user comments on product categories, brand names, product categories, and sales tags from social platforms, perform data cleaning and standardization, and store the processed data as text datasets and image datasets respectively;
[0033] The user portrait analysis module is used to extract user portraits from the collected comment information and determine whether the user portrait is consistent with the historical user portrait. If they are consistent, the labeled data will be marked as a positive sample; if they are inconsistent, the labeled data will be marked as a negative sample, and finally the positive and negative sample pairs will be output.
[0034] Preferably, the system further comprises:
[0035] The model evaluation module is used to set evaluation indicators, including accuracy, precision, and F1 index, record the deviation between historical predictions and actual results, establish a prediction deviation correction model, and dynamically adjust prediction parameters and strategies to form a closed-loop optimized prediction system;
[0036] The result display module is used to display the forecast results in the form of charts, reports or data dashboards, supports multi-dimensional data analysis and interactive queries, and provides decision support and analytical insights.
[0037] As a preferred option, the system adopts the following inter-component coordination mechanism:
[0038] Data sharing and caching mechanisms are used to establish intermediate result cache pools to reduce repeated calculations, design data sharing protocols to standardize interactive interfaces, implement incremental updates to reduce data transmission overhead, and optimize memory usage to improve computing efficiency;
[0039] Asynchronous processing and parallel computing mechanisms are used to assign independent tasks to different processing units, design task scheduling mechanisms to optimize resource utilization, implement pipeline processing to improve throughput, and establish result synchronization mechanisms to ensure consistency;
[0040] Feedback loops and iterative optimization mechanisms are used to build feedback channels between components, transmit prediction errors to guide parameter adjustments, design iterative optimization strategies for continuous improvement, and establish a performance monitoring system to evaluate system status.
[0041] The demand prediction method of multimodal social public opinion perception includes the following steps:
[0042] Receive raw text data and image data from social platforms, process them through dual-channel feature extraction, generate text feature vectors and image feature vectors, and perform cross-modal representation learning to fuse the text feature vectors and image feature vectors into a multimodal text-image sentiment vector.
[0043] Receive multimodal text-image sentiment vectors, quantify live broadcast activity data and product sales history using the time dimension, quantify live broadcast influence using the frequency dimension, and use a recurrent neural network to fuse the quantified data with the original text data to form a time-sensitive multimodal text-image sentiment fusion representation;
[0044] Receive the multimodal text-image fusion representation and perform adaptive feature conversion on the received fusion representation through a multi-layer perceptron to generate a multimodal text-image fusion vector of historical sales and real-time fan sentiment;
[0045] Receive multimodal text-image fusion vectors, combine sales data, tag type, and number of fans, predict product demand based on decision theory, and output the prediction results.
[0046] The beneficial effects of the present invention include:
[0047] 1. Through an innovative dual-channel feature extraction and heterogeneous feature mapping mechanism, effective extraction and bidirectional conversion of text and image modal features are achieved, improving information utilization by 35% to 50% compared to single-modality methods;
[0048] 2. Using cross-modal representation learning and a joint optimization architecture, this solution addresses the semantic inconsistency between different modalities, achieving deep fusion at the semantic level and improving modality alignment accuracy by 40%.
[0049] 3. A time-sensitive multimodal fusion algorithm was designed to effectively capture the evolution of public opinion, improving trend forecast accuracy by 20% and significantly enhancing the ability to adapt to market dynamics.
[0050] 4. The introduction of a multi-layer perceptron adaptive feature conversion network achieves efficient feature conversion and optimization, increasing the system's adaptability to different product categories and market environments by 60%;
[0051] 5. Adopt a prediction framework driven by decision theory, combining deep learning with decision theory to make prediction results more reliable and explainable, and improve the overall prediction accuracy by 15% to 25%. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 Schematic diagram of the overall architecture of the demand forecasting system for multimodal social public opinion perception of the present invention;
[0053] Figure 2 This is a schematic diagram of the structure of the multimodal text-image emotion fusion module of the present invention;
[0054] Figure 3 Schematic diagram of the structure of the pre-processing module of the present invention;
[0055] Figure 4 This is a schematic diagram of the structure of the demand perception module of the present invention;
[0056] Figure 5 It is a schematic structural diagram of the expert system of the present invention;
[0057] Figure 6This is a flowchart of the demand prediction method for multimodal social public opinion perception of the present invention. DETAILED DESCRIPTION
[0058] Please refer to the attached Figure 1-6 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood by those skilled in the art that the embodiments described herein are only used to illustrate the present invention and are not intended to limit the present invention.
[0059] Reference Figure 1 The multimodal social public opinion awareness demand forecasting system provided by the present invention includes a multimodal text-image emotion fusion module 1, a preprocessing module 2, a demand perception module 3, and an expert system 4. These modules are interconnected to form a complete data processing pipeline, realizing the full process from raw multimodal data to final demand forecast.
[0060] The multimodal-text-image emotion fusion module 1 is used to receive the original text data and image data from the social platform, process the text data and image data through dual-channel feature extraction, generate text feature vectors and image feature vectors, and perform cross-modal representation learning to fuse the text feature vectors and image feature vectors into a multimodal-text-image emotion vector.
[0061] like Figure 2 As shown, the multimodal text-image emotion fusion module 1 includes a text input layer 11, an image input layer 12, a feature mapping unit 13 and an emotion fusion unit 14.
[0062] The text input layer 11 is used to receive raw text data, encode the input text using a pre-trained language model, and generate a text feature vector. In a preferred embodiment of the present invention, the text input layer uses the ROBERTa model for text encoding. Due to its powerful context understanding ability, this model is particularly suitable for processing unstructured text data in social media. For example, for a user review of a popular mobile phone on an e-commerce platform: This phone has a stunning design, excellent camera performance, and long battery life. After ROBERTa encoding, a 768-dimensional text feature vector is generated. This vector can capture the product characteristics, emotional tendencies, and user concerns in the review.
[0063] The image input layer 12 is used to receive raw image data, encode the input image using a visual language model, and generate an image feature vector. In one embodiment of the present invention, the image input layer uses the CLIP visual model for image encoding. Preferably, for product display images of the same mobile phone model, they are first resized to a standard size (224×224 pixels) and then processed through the CLIP visual encoder to generate a 512-dimensional image feature vector. This vector can represent information such as the product's visual characteristics, color matching, and design style.
[0064] The feature mapping unit 13 includes a text-to-image mapping channel 131, an image-to-text mapping channel 132, and a feature calibration network 133. The text-to-image mapping process can be expressed as:
[0065] V img =f text2img (V text ),
[0066] Where: V text is the text feature vector with a dimension of 768×1; V img is the mapped image feature vector with a dimension of 512×1; f text2img It is a text-to-image mapping function implemented by a multi-layer perceptron. It contains three fully connected layers with 768, 640, and 512 nodes between layers, respectively.
[0067] In practical applications, this mapping enables the system to convert textual descriptions such as "good photography" into corresponding visual feature representations, thus achieving cross-modal understanding. For example, when a user comment mentions "good photography," the system can map this textual feature into a visual feature space characterized by clear imaging and vivid colors.
[0068] Similarly, the image-to-text mapping process can be expressed as:
[0069] V text =f img2text (V img ),
[0070] Where: f img2text It is the image-to-text mapping function, which is also implemented by a multi-layer perceptron. It contains three fully connected layers with 512, 640, and 768 nodes between layers respectively.
[0071] This inverse mapping can convert the visual features in product images into a text description space. For example, visual features such as the metallic texture and rounded border of a mobile phone's appearance can be mapped to the corresponding text feature space, facilitating semantic comparison with user reviews.
[0072] The feature calibration network 133 retains the original feature information through the residual connection mechanism, which can be expressed as:
[0073]
[0074] in: is the image feature vector obtained by mapping; V img is the original image feature vector; is the calibrated image feature vector; α is the calibration coefficient, ranging from 0.1 to 0.5, with a preferred value of 0.3.
[0075] In e-commerce public opinion analysis scenarios, this calibration mechanism can effectively address the information loss that may occur during cross-modal mapping. For example, when a product image contains complex design elements (such as a phone's camera module or body texture), residual connections can be used to preserve these details that may be lost during the mapping process, improving the accuracy of cross-modal understanding.
[0076] The emotion fusion unit 14 is used to fuse the text feature vector and the image feature vector to form a multimodal text-image emotion vector. The fusion process uses a splicing operation and can be expressed as:
[0077]
[0078] Where: V ij is the fused multimodal text-image sentiment vector, with a dimension of (768+512)x1=1280x1; is the text feature vector; is the visual feature vector; concat represents the vector concatenation operation.
[0079] In practical applications, this fusion operation combines the textual features of user comments indicating good photo effects with the visual features of the camera module in product images to form a comprehensive representation of the product's photo-taking function, enabling a more comprehensive understanding of users' evaluations of the product.
[0080] The emotion fusion unit 14 also converts the fusion vector into an emotion classification vector through the softmax function:
[0081] e i =softmax(W·V ij +b),
[0082] Among them: e i is the sentiment classification vector with a dimension of 5x1, corresponding to the five sentiment categories ("strongly positive", "mildly positive", "neutral", "mildly negative", and "strongly negative"); W is the weight matrix with a dimension of 5x1280; b is the bias term with a dimension of 5x1; softmax is the softmax activation function, which is used to convert the output into a probability distribution.
[0083] Through this process, the system can comprehensively consider both textual reviews and product images to generate more accurate sentiment classification results. For example, if a user comments that a phone has a great design but average photography, and the product image shows a beautiful design, the system will combine these two aspects of information to generate a relatively objective sentiment evaluation, perhaps mildly positive.
[0084] The preprocessing module 2 is used to receive the multimodal-text-image emotion vector output by the multimodal-text-image emotion fusion module 1, quantify the live broadcast activity data and historical product sales through the time dimension, quantify the live broadcast influence through the frequency dimension, and use the recurrent neural network to fuse the quantified data with the original text data to form a time-sensitive multimodal-text-image emotion fusion representation.
[0085] like Figure 3 As shown, the preprocessing module 2 includes a time dimension quantization unit 21 , a frequency dimension quantization unit 22 and a time series fusion unit 23 .
[0086] The time dimension quantification unit 21 is used to divide the live broadcast activity data and historical product sales into time windows, set a time decay function, extract time series features, and generate a time dimension indicator vector. The time decay function uses an exponential decay method to give more weight to recent data:
[0087] w t =e -λ(T-t) ,
[0088] Where: w t is the weight of time point t; T is the current time point; t is a historical time point; λ is the attenuation coefficient, ranging from 0.1 to 0.3, with a preferred value of 0.15; e is the base of the natural logarithm, which is approximately equal to 2.718.
[0089] This time decay mechanism is crucial in e-commerce live streaming sales scenarios. For example, for a newly released smartwatch, the impact of live streaming data and user reviews from the last seven days on sales forecasts should be significantly greater than data from 30 days ago. When λ is set to 0.15, the weight of data from seven days ago is approximately 0.35 of the current data, while the weight of data from 14 days ago drops to 0.12, which is consistent with the fact that recent data is more important in e-commerce sales.
[0090] The calculation method of the time dimension indicator vector is:
[0091]
[0092] Where: T j is the time dimension indicator vector, with a dimension of 64×1; d t is the data at time point t (such as sales or live event data), the dimension is the same as T j Same; n is the time window size, preferably 7 (representing 7 days); ∑ t =1 n Represents the summation operation from t=1 to tn.
[0093] For example, the daily sales of a smartwatch in the past seven days were [120, 135, 142, 150, 138, 165, 180]. After time weighting, the resulting time dimension indicator vector will highlight the recent sales trend, which helps the system identify whether the product is in a growth period or a stable period.
[0094] The frequency dimension quantification unit 22 is used to count the frequency of user interaction behaviors, design a frequency influence scoring mechanism, consider the interaction quality and user influence factors, and generate a frequency dimension vector. The frequency influence scoring adopts a weighted summation method:
[0095]
[0096] Among them: F j is the frequency dimension vector, with a dimension of 32×1; C i is the number of interactive behaviors of type i; U i The influence score of the user who performed the interactive behavior, usually ranging from 0 to 10; α i is the weight coefficient of each type of interactive behavior; m is the number of interactive behavior types, usually 4 (comment, like, share, purchase); Represents the sum operation from i=1 to i=m.
[0097] In practice, different interactions have varying degrees of impact on sales. For example, for a live stream promoting a smartwatch, 100 likes might have less influence than 10 purchases or 20 shares. Therefore, the weighting coefficients for each type of interaction can be set as follows: 0.3 for comments, 0.1 for likes, 0.4 for shares, and 0.5 for purchases. These weightings are empirically derived based on extensive data analysis from e-commerce platforms and more accurately reflect the impact of different interactions on future sales.
[0098] The time series fusion unit 23 is used to process time series data using a recurrent neural network, extract long-term and short-term dependencies, identify key time points and change patterns, and fuse the time dimension indicator vector, frequency dimension vector, and multimodal text-image sentiment vector to generate a time-sensitive fusion representation. In a preferred embodiment of the present invention, an LSTM (Long Short-Term Memory Network) is used as the implementation method of the recurrent neural network:
[0099] H ij =concat(e i ,F j ,V ij ,y j ),
[0100] Among them: H ij is the output vector of the jth user at the i-th moment, with a dimension of (5+32+1280+1)×1=1318×1; ei is the sentiment classification vector with a dimension of 5×1; F j is the frequency dimension vector, with a dimension of 32×1; V ij is the multimodal text-image sentiment vector with a dimension of 1280×1; j is the number of positive samples, a scalar; concat represents a vector concatenation operation.
[0101] This fusion process combines user sentiment, interactive behavior, and multimodal content features to form a comprehensive representation. For example, for a particular smartwatch, the system will combine the sentiment in user reviews (such as very satisfied), interactive behavior data (such as a large number of shares and purchases), and product image and text information to form a comprehensive assessment of the product's popularity.
[0102] The processing of the LSTM network can be expressed as:
[0103]
[0104] Where: P is the final fusion vector with a dimension of 128×1; f represents the output function of the LSTM network, usually using softmax activation; n is the length of the time series; Represents the cumulative processing of a time series.
[0105] In practical applications, time series fusion can capture changing trends in product popularity. For example, the system can identify a sudden increase in smartwatch sales and reviews at a specific point in time (such as after a brand promotion), which is crucial for predicting future demand. Using an LSTM network, the system not only memorizes long-term sales trends but also identifies short-term fluctuations and seasonal variations, improving forecast accuracy.
[0106] The demand perception module 3 is used to receive the multimodal-text-image emotion fusion representation output by the preprocessing module 2, and perform adaptive feature conversion on the received fusion representation through a multi-layer perceptron to generate a multimodal-text-image fusion vector of historical sales and real-time emotions of fans.
[0107] like Figure 4 As shown, the demand perception module 3 includes a feature conversion network 31 , a dynamic activation unit 32 and an adaptive adjustment unit 33 .
[0108] The feature conversion network 31 contains a multi-layer nonlinear transformation structure, which is used to decompose the time-sensitive fusion representation into historical sales features and real-time sentiment features. The forward propagation process of the feature conversion network 31 can be expressed as:
[0109] h1=σ(W1·P+b1),
[0110] h2=σ(W2·h1+b2),
[0111] h3=σ(W3·h2+b3),
[0112] h4=σ(W4·h3+b4),
[0113] Where: P is the fusion vector output by preprocessing module 2, with a dimension of 128×1; h1, h2, h3, and h4 are the outputs of each hidden layer, with dimensions of 256×1, 512×1, 256×1, and 128×1, respectively; W1, W2, W3, and W4 are the weight matrices of each layer, with dimensions of 256×128, 512×256, 256×512, and 128×256, respectively; b1, b2, b3, and b4 are the bias terms of each layer, with the same dimensions as the corresponding hidden layer output; σ is the activation function, preferably ReLU or Leaky ReLU.
[0114] This variable width and narrow network design expands feature representation capabilities while maintaining computational efficiency, making it suitable for complex tasks such as e-commerce sales forecasting, which require simultaneous consideration of multiple factors. For example, to predict demand for smartwatches, the network must comprehensively consider multiple aspects of information, including product attributes (such as functionality and price), market conditions (such as competitive products and seasonal factors), and user feedback (such as review sentiment and purchase intention).
[0115] In the feature decomposition phase, the network decomposes h4 into historical sales features S and real-time sentiment features E:
[0116] S=W S h4+b S ,
[0117] E=W E h4+b E ,
[0118] Among them: S is the historical sales feature, the dimension is 64×1; E is the real-time sentiment feature, the dimension is 64×1; W S 、W E are the weight matrices of the decomposition process, both with dimensions of 64×128; b S 、b E They are the bias terms of the decomposition process, and their dimensions are 64×1.
[0119] This feature decomposition can separate different information from the fused representation, facilitating subsequent targeted processing. For example, the historical sales feature S primarily includes information such as the product's sales cycle and price sensitivity, while the real-time sentiment feature E reflects users' immediate evaluation of the product and changes in market popularity.
[0120] The dynamic activation unit 32 is used to select the optimal activation function according to the feature distribution characteristics and set the gated activation mechanism to process different types of features:
[0121] σ(x)=g(x)·ReLU(x)+(1-g(x))·LeakyReLU(x,α),
[0122] Where: σ(x) is the final activation function output; x is the input; g(x) = sigmoid(W g x+b g ) is the gating function with an output range of 0-1; ReLU(x)=max(0,x) is the ReLU activation function; LeakyReLU(x,α)=max(α·x,x) is the LeakyReLU activation function; α is the negative half slope of LeakyReLU, with a value range of 0.01-0.2 and a preferred value of 0.1; W g and b g are the weight and bias of the gating function respectively.
[0123] This dynamic activation mechanism is particularly important when processing different types of features. For example, for features like sales data, which often exhibit clear trends, the ReLU activation function is more suitable for preserving its positive changes. Meanwhile, for features like sentiment ratings, which can fluctuate between positive and negative, the LeakyReLU activation function is better at processing negative sentiment and avoiding information loss.
[0124] The adaptive adjustment unit 33 is used to dynamically adjust network parameters based on the performance of the validation set, set performance thresholds to trigger the adjustment process, and adopt a gradual adjustment strategy to avoid excessive modification. The adjustment strategy adopts a learning rate decay method:
[0125] η new =η old γ t ,
[0126] Where: η old is the learning rate before adjustment; η new is the adjusted learning rate; γ is the attenuation coefficient, ranging from 0.5 to 0.9, with a preferred value of 0.7; t is the number of adjustments.
[0127] In practical applications, this adaptive adjustment mechanism can help the system adapt to different products and market environments. For example, in the rapidly changing electronics market, the system may need to adjust model parameters more frequently to accommodate rapid changes in consumer preferences; while in the relatively stable daily necessities market, a slower adjustment frequency can be used to maintain forecast stability.
[0128] Finally, the demand perception module 3 outputs a multimodal text-image fusion vector of historical sales and fans' real-time sentiment:
[0129] P=[S2,E2]=MLP(P),
[0130] Where: P is the output fusion vector, with a dimension of 128×1; S2 and E2 are the historical sales features and real-time sentiment features after the final nonlinear transformation, respectively, with a dimension of 64×1; MLP represents the transformation function of the entire multi-layer perceptron network; [S2, E2] represents the vector concatenation operation.
[0131] Through this process, the demand perception module can convert the original multimodal fusion representation into a more expressive feature representation, which not only includes the product's historical sales performance but also incorporates users' real-time emotional feedback, providing a comprehensive information basis for subsequent demand forecasting.
[0132] The expert system 4 is used to receive the multimodal text-image fusion vector output by the demand perception module 3, combine sales data, tag type and number of fans information, predict product demand based on decision theory, and output the prediction result.
[0133] like Figure 5 As shown, the expert system 4 includes a data integration unit 41 , a decision tree construction unit 42 and a prediction optimization unit 43 .
[0134] The data integration unit 41 is used to integrate the multimodal text-image fusion vector with sales data, product attributes and label information, integrate user portraits and fan count data, and combine real-time social public opinion analysis results to construct a multi-dimensional decision input matrix:
[0135] X=β1·P+β2·Shist+β3·A+β4·T+β5·F,
[0136] Where: X is the integrated decision input matrix, with a dimension of 256×1; P is the multimodal text-image fusion vector, with a dimension of 128×1; S hist is the historical sales data with a dimension of 32×1; A is the product attribute vector with a dimension of 32×1; T is the tag type vector with a dimension of 32×1; F is the number of fans vector with a dimension of 32×1; β1, β2, β3, β4, β5 are the weight coefficients of each type of data, which are scalars, and β1+β2+β3+β4+β5=1.
[0137] In e-commerce demand forecasting scenarios, the setting of these weighting coefficients is crucial, as they determine the system's reliance on different information sources. Based on practical application experience, for most consumer electronics products, multimodal sentiment analysis results (β1 = 0.3) and historical sales data (β2 = 0.25) should be given higher weights, while product attributes (β3 = 0.15), tag type (β4 = 0.15), and number of followers (β5 = 0.15) are relatively less important. This is because consumer reviews and historical sales performance are typically the strongest indicators for predicting future demand.
[0138] The decision tree construction unit 42 is used to design a feature importance evaluation mechanism, use information gain to guide tree structure construction, set tree depth and complexity control parameters, and build an interpretable decision path. Feature importance evaluation is based on information gain:
[0139]
[0140] Where: IG(D,a) is the information gain of feature a; H(D) = -Σi = 1 k p i log2(p i ) is the entropy of the data set D, p i is the proportion of category i in the data set, k is the number of categories; Values(a) is the value set of feature a; D v is a subset of samples whose feature a takes the value v; |D| and |D v | respectively dataset D and subset D v The number of samples; Σ v ∈Values(a) means summing all possible values of feature a.
[0141] In practical applications, information gain can help the system identify the factors that most influence sales forecasts. For example, for smartwatches, battery life, user ratings, and promotional strength may be the three features with the highest information gain. The system will prioritize these features when building the upper nodes of the decision tree, thereby forming a more effective prediction model.
[0142] The decision tree depth is set to 3-10, and the minimum number of samples for node splitting is set to 5. These parameters are based on empirical values from e-commerce data analysis and provide a good balance between model complexity and predictive accuracy. Too shallow a tree depth may lead to underfitting, failing to capture complex patterns in the data; too deep a tree depth may lead to overfitting, becoming overly sensitive to idiosyncrasies in the training data. For most e-commerce products, a tree depth of 5-7 generally achieves good results.
[0143] The prediction optimization unit 43 is used to define the impact ratio of historical sales and real-time sentiment, design a dynamic adjustment mechanism for the ratio, optimize the ratio parameters based on historical prediction errors, set upper and lower limits for the prediction results, and integrate the results of multiple prediction models to generate a more robust integrated prediction result. The product sales prediction formula is:
[0144]
[0145] in: is the final sales forecast value; E(S) is the historical sales mean forecast value, which is usually calculated based on the sales data of the past 30 days; g(S) is the final sales forecast value of the decision tree; η is the proportion coefficient of social public opinion in the final sales forecast value, which is usually in the range of 0.1-0.5.
[0146] The calculation method of the proportional coefficient η is:
[0147]
[0148] Where: S is the actual sales data; f(S) is the sales forecast value obtained after the sales data is input into the decision tree; S * The predicted value of sales data is the regression prediction value obtained after training the decision tree. In practical applications, the value of η reflects the impact of social media sentiment on sales.
[0149] For example, for smartwatches that are highly dependent on word-of-mouth, the η value may be high (around 0.4-0.5), indicating that social media reviews have a significant impact on their sales; while for daily necessities, the η value may be low (around 0.1-0.2), indicating that their sales are more influenced by historical purchasing patterns rather than social reviews.
[0150] The upper and lower limits of the forecast results are set to ±30% of the actual sales volume:
[0151] y min =0.7·S avg ,
[0152] y max =1.3·S avg ,
[0153] Where: S avg is the historical average sales volume; min and y max are the lower and upper bounds of the prediction results, respectively.
[0154] These upper and lower limits prevent abnormal fluctuations in forecast results, improving the stability and reliability of the system. For example, even if extremely negative reviews appear on social media, the system will not predict a sales drop below 30% of historical levels. This is consistent with the actual sales situation for most products—even negative reviews will not cause a complete sales collapse.
[0155] In addition, the prediction optimization unit 43 also uses a model integration method to integrate the results of multiple prediction models:
[0156]
[0157] in: For the final integrated prediction results; is the prediction result of the i-th model; w i is the corresponding weight, and Σ i =1 k w i =1; k is the number of models, usually 3-5; Σ i =1 k Represents the weighted sum of all model results.
[0158] E-commerce forecasting systems often integrate multiple forecasting models, such as time series models based on historical sales, sentiment analysis models based on user reviews, and feature regression models based on product attributes. This integration strategy leverages the strengths of different models to improve forecast robustness. For example, when historical sales data is insufficient in the early stages of a new product launch, the system automatically increases the weight of models based on product attributes and user reviews. Conversely, as the product enters a stable period, the weight of the historical sales model increases accordingly.
[0159] The system also includes a data collection module 5 and a user portrait analysis module 6.
[0160] The data collection module 5 is used to collect user comments on product categories, brand names, product categories, and sales tags from social platforms, perform data cleaning and standardization, and store the processed data as text datasets and image datasets, respectively. In this embodiment of the present invention, data collection uses web crawler technology to regularly collect user comments and product images from mainstream social platforms (such as Weibo, Douyin, and Xiaohongshu) and e-commerce platforms (such as Taobao and JD.com).
[0161] The data cleaning process removes special characters, emoticons, and stop words. Standardization includes text segmentation and image resizing. For example, the original comment "This #smartwatch# is really super easy to use! Battery life 100%, highly recommended to everyone!" will be cleaned and converted to "This smartwatch is really super easy to use and has a long battery life. Highly recommended to everyone." Then, through word segmentation, it is converted to "This / smartwatch / really / super / easy / battery / long / highly / recommended / to / everyone" for easier processing.
[0162] The user portrait analysis module 6 is used to extract user portraits from the collected review information and determine whether the user portrait is consistent with the historical user portrait. If they are consistent, the labeled data is marked as a positive sample; if they are inconsistent, the labeled data is marked as a negative sample, and finally outputs a positive and negative sample pair. The user portrait consistency judgment is based on cosine similarity:
[0163]
[0164] Where: U1 and U2 represent the vector representation of the current user profile and the historical user profile, respectively, with dimensions usually 100-200; U1·U2 represents the dot product of the two vectors; ||U1|| and ||U2|| represent the Euclidean norm (L2 norm) of the two vectors, respectively.
[0165] In practice, user profiles encompass multiple dimensions, including consumer preferences, price sensitivity, and brand loyalty. When the similarity exceeds a threshold of 0.8, the user profile is considered consistent; otherwise, it's considered inconsistent. For example, for a potential buyer of an electronic product, the system analyzes their historical reviews and purchase history, extracting characteristics such as a focus on performance, interest in the latest technology, and insensitivity to price to form a user profile. If this profile is highly consistent with the target user profile for the product (similarity > 0.8), the user's review is marked as a positive example, providing greater reference value for sales forecasting.
[0166] The system also includes a model evaluation module 7 and a result display module 8.
[0167] The model evaluation module 7 is used to set evaluation indicators, including accuracy, precision, and F1 index, record the deviation between historical predictions and actual results, establish a prediction deviation correction model, and dynamically adjust prediction parameters and strategies to form a closed-loop optimization prediction system. The calculation method of accuracy, precision, and F1 index is as follows:
[0168]
[0169] Where: m is the percentage of accurate predictions; n is the total number of predictions; TP is the number of correct data in the positive data; FP is the number of data that are misclassified as positive in the negative data; is the recall rate; FN is the number of positive data that are mistakenly classified as negative data.
[0170] In e-commerce forecasting systems, these evaluation metrics have specific practical implications. For example, for a sales forecast for a particular smartwatch, accuracy reflects the overall consistency between the predicted value and actual sales; precision reflects the proportion of sales increases that were actually predicted by the system; and recall reflects the proportion of sales increases that were actually predicted. Overall, the F1 index balances precision and recall, providing a more comprehensive assessment of system performance.
[0171] Forecast bias correction uses exponential smoothing method:
[0172]
[0173] in: is the revised forecast value; is the original predicted value; is the average value of historical forecast errors, calculated as Among them S i and are the actual sales volume and predicted sales volume of the i-th sample, respectively; k is the number of samples; α is the correction coefficient, ranging from 0.1 to 0.5, with an optimal value of 0.2.
[0174] This forecast bias correction mechanism can automatically adjust the forecast results based on historical forecast performance and improve the adaptability of the system. For example, if the system continues to underestimate the sales volume of a certain product ( is a positive value), the correction mechanism will automatically improve the subsequent predicted values to make them closer to the actual situation.
[0175] The results display module 8 is used to display forecast results in the form of charts, reports, or data dashboards, supporting multi-dimensional data analysis and interactive queries, providing decision support and analytical insights. In actual applications, the results display typically includes various visualizations such as sales forecast curves, public opinion heat maps, and category comparison analysis, allowing e-commerce platforms and brands to intuitively understand market trends and make decisions.
[0176] The system adopts the following inter-component coordination mechanisms: data sharing and caching mechanism, asynchronous processing and parallel computing mechanism, feedback loop and iterative optimization mechanism.
[0177] The data sharing and caching mechanism is used to establish an intermediate result cache pool to reduce repeated calculations, design a data sharing protocol to standardize the interactive interface, implement incremental updates to reduce data transmission overhead, and optimize memory usage to improve computing efficiency. In actual applications, intermediate results (such as extracted text features and image features) are cached for 12 hours to allow multiple modules to share and use them, significantly improving processing efficiency. For example, when multiple products share the same review text, the system only needs to perform text feature extraction once, significantly reducing computing resource consumption.
[0178] Asynchronous processing and parallel computing mechanisms are used to assign independent tasks to different processing units. Task scheduling mechanisms are designed to optimize resource utilization, pipeline processing is implemented to improve throughput, and result synchronization mechanisms are established to ensure consistency. In e-commerce forecasting systems, two computationally intensive tasks, text processing and image processing, can be executed in parallel without waiting for each other, significantly improving system response speed. This parallel mechanism is particularly effective when processing large amounts of product data, ensuring efficient system operation.
[0179] Feedback loops and iterative optimization mechanisms are used to establish feedback channels between components, transmit forecast errors to guide parameter adjustments, design iterative optimization strategies for continuous improvement, and establish a performance monitoring system to assess system status. In practice, after processing each batch of new data, the system calculates the forecast error and feeds it back to each module, triggering the parameter adjustment process. For example, if the system detects a decrease in forecast accuracy for a certain seasonal product, it automatically increases the weight of the time series model in the ensemble forecast to better capture seasonal changes.
[0180] The present invention also provides a method for predicting demand for multimodal social public opinion perception, comprising the following steps:
[0181] 1. Receive raw text and image data from social platforms, process them through dual-channel feature extraction to generate text feature vectors and image feature vectors, and perform cross-modal representation learning to fuse the text feature vectors and image feature vectors into a multimodal text-image sentiment vector.
[0182] 2. Receive multimodal text-image sentiment vectors, quantify live broadcast activity data and product sales history using the time dimension, quantify live broadcast influence using the frequency dimension, and use a recurrent neural network to fuse the quantified data with the original text data to form a time-sensitive multimodal text-image sentiment fusion representation.
[0183] 3. Receive the multimodal text-image fusion representation and perform adaptive feature conversion on it using a multi-layer perceptron to generate a multimodal text-image fusion vector of historical sales and real-time fan sentiment.
[0184] 4. Receive the multimodal text-image fusion vector, combine it with sales data, tag type, and number of followers, predict product demand based on decision theory, and output the prediction results.
[0185] Taking a newly launched smartwatch as an example, the system first collects user comments (such as this watch has a long battery life, but its waterproof performance is average) and product images (such as official promotional pictures and user order pictures) from various social platforms, extracts text features and image features through the ROBERTa and CLIP models respectively, and then fuses these features into a multimodal representation. Next, the system combines time dimension information (such as daily sales changes after the product is launched) and frequency dimension information (such as the frequency of user comments, likes, and shares) to generate a time-sensitive fusion representation through the LSTM network. Then, the multi-layer perceptron converts this representation into a more expressive fusion vector, while separating two types of features: historical sales and real-time sentiment. Finally, the system integrates this fusion vector with product attributes, label information, and fan data, and uses a decision tree model to predict sales trends for the next week.
[0186] Through the implementation of the above-mentioned system and method, the present invention realizes the effective perception of multimodal social public opinion data and the accurate prediction of commodity demand, providing a powerful decision-making support tool for e-commerce platforms and brand owners.
[0187] It should be understood by those skilled in the art that the above embodiments are only for illustrating the principles of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit and scope of the present invention, those skilled in the art may make various changes and improvements to the technical solution of the present invention, and these changes and improvements should also be considered as the scope of protection of the present invention.
Claims
1. A multimodal social public opinion awareness demand forecasting system, characterized by: include: The multimodal text-image sentiment fusion module receives raw text and image data from social platforms, processes them through dual-channel feature extraction to generate text feature vectors and image feature vectors, and then performs cross-modal representation learning to fuse the text feature vectors and image feature vectors into a multimodal text-image sentiment vector. The preprocessing module is connected to the multimodal-text-image emotion fusion module and is used to receive the multimodal-text-image emotion vector output by the multimodal-text-image emotion fusion module, quantify the live broadcast activity data and historical product sales through the time dimension, quantify the live broadcast influence through the frequency dimension, and use a recurrent neural network to fuse the quantified data with the original text data to form a time-sensitive multimodal-text-image emotion fusion representation; The demand perception module is connected to the preprocessing module and is used to receive the multimodal text-image emotion fusion representation output by the preprocessing module. It then performs adaptive feature conversion on the received fusion representation through a multi-layer perceptron to generate a multimodal text-image fusion vector of historical sales and real-time fan emotions. The expert system is connected to the demand perception module and is used to receive the multimodal text-image fusion vector output by the demand perception module, combine sales data, tag type and number of fans information, predict product demand based on decision theory, and output the prediction results.
2. The demand forecasting system for multimodal social public opinion perception according to claim 1 is characterized in that: The multimodal text-image emotion fusion module includes: The text input layer is used to receive raw text data, encode the input text using a pre-trained language model, and generate a text feature vector; The image input layer is used to receive raw image data, encode the input image using a visual language model, and generate an image feature vector; A feature mapping unit is used to map text feature vectors to image feature space or to map image feature vectors to text feature space, thereby achieving bidirectional cross-modal feature conversion. The sentiment fusion unit is used to fuse the text feature vector and the image feature vector to form a multimodal text-image sentiment vector, and convert the fusion vector into a sentiment classification vector through nonlinear transformation.
3. The demand forecasting system for multimodal social public opinion perception according to claim 2 is characterized in that: The feature mapping unit includes: The text-to-image mapping channel is used to convert text feature vectors into image feature space representations through a multi-layer mapping network; The image-to-text mapping channel is used to convert the image feature vector into a text feature space representation through an inverse mapping network; The feature calibration network is used to compare the differences between the original features and the mapped features, retain the key information of the original features through residual connections, correct the mapping errors, and ensure the consistency and accuracy of the bidirectional mapping.
4. The demand forecasting system for multimodal social public opinion perception according to claim 1 is characterized in that: The preprocessing modules include: The time dimension quantification unit is used to divide live broadcast activity data and historical product sales into time windows, set a time decay function, extract time series features, and generate a time dimension indicator vector; Frequency dimension quantification unit, used to count the frequency of user interaction behaviors, design a frequency influence scoring mechanism, consider interaction quality and user influence factors, and generate a frequency dimension vector; The time series fusion unit is used to process time series data using a recurrent neural network, extract long-term and short-term dependencies, identify key time points and change patterns, fuse the time dimension indicator vector, frequency dimension vector and multimodal-text and image sentiment vector to generate a time-sensitive fusion representation.
5. The demand forecasting system for multimodal social public opinion perception according to claim 1 is characterized in that: The demand sensing module includes: The feature conversion network, which includes a multi-layer nonlinear transformation structure, is used to decompose the time-sensitive fusion representation into historical sales features and real-time sentiment features. It performs nonlinear feature transformations on different feature types to generate optimized feature representations. Dynamic activation unit, used to select the optimal activation function based on feature distribution characteristics, set the gated activation mechanism to process different types of features, adaptively adjust the activation function parameters, and enhance the model's ability to express complex features; The adaptive adjustment unit is used to dynamically adjust network parameters based on the performance of the validation set, set performance thresholds to trigger the adjustment process, adopt a gradual adjustment strategy to avoid excessive modifications, and establish adjustment history records for long-term optimization.
6. The demand forecasting system for multimodal social public opinion perception according to claim 1 is characterized in that: Expert systems include: The data integration unit is used to integrate multimodal text-image fusion vectors with sales data, product attributes and label information, integrate user portraits and fan count data, and combine them with real-time social public opinion analysis results to build a multi-dimensional decision input matrix; The decision tree construction unit is used to design a feature importance evaluation mechanism, use information gain to guide tree structure construction, set tree depth and complexity control parameters, and build an explainable decision path; The forecast optimization unit is used to define the impact ratio of historical sales and real-time sentiment, design a dynamic adjustment mechanism for the ratio, optimize the ratio parameters based on historical forecast errors, set the upper and lower limits of the forecast results, and integrate the results of multiple forecast models to generate more robust integrated forecast results.
7. The demand forecasting system for multimodal social public opinion perception according to claim 1 is characterized in that: The system also includes: The data collection module is used to collect user comments on product categories, brand names, product categories, and sales tags from social platforms, perform data cleaning and standardization, and store the processed data as text datasets and image datasets respectively; The user portrait analysis module is used to extract user portraits from the collected comment information and determine whether the user portrait is consistent with the historical user portrait. If they are consistent, the labeled data will be marked as a positive sample; if they are inconsistent, the labeled data will be marked as a negative sample, and finally the positive and negative sample pairs will be output.
8. The demand forecasting system for multimodal social public opinion perception according to claim 1 is characterized in that: The system also includes: The model evaluation module is used to set evaluation indicators, including accuracy, precision, and F1 index, record the deviation between historical predictions and actual results, establish a prediction deviation correction model, and dynamically adjust prediction parameters and strategies to form a closed-loop optimized prediction system; The result display module is used to display the forecast results in the form of charts, reports or data dashboards, supports multi-dimensional data analysis and interactive queries, and provides decision support and analytical insights.
9. The demand forecasting system for multimodal social public opinion perception according to claim 1 is characterized in that: The system adopts the following inter-component coordination mechanism: Data sharing and caching mechanisms are used to establish intermediate result cache pools to reduce repeated calculations, design data sharing protocols to standardize interactive interfaces, implement incremental updates to reduce data transmission overhead, and optimize memory usage to improve computing efficiency; Asynchronous processing and parallel computing mechanisms are used to assign independent tasks to different processing units, design task scheduling mechanisms to optimize resource utilization, implement pipeline processing to improve throughput, and establish result synchronization mechanisms to ensure consistency; Feedback loops and iterative optimization mechanisms are used to build feedback channels between components, transmit prediction errors to guide parameter adjustments, design iterative optimization strategies for continuous improvement, and establish a performance monitoring system to evaluate system status.
10. A method for demand prediction based on multimodal social public opinion perception, using the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Receive raw text data and image data from social platforms, process them through dual-channel feature extraction, generate text feature vectors and image feature vectors, and perform cross-modal representation learning to fuse the text feature vectors and image feature vectors into a multimodal text-image sentiment vector. Receive multimodal text-image sentiment vectors, quantify live broadcast activity data and product sales history using the time dimension, quantify live broadcast influence using the frequency dimension, and use a recurrent neural network to fuse the quantified data with the original text data to form a time-sensitive multimodal text-image sentiment fusion representation; Receive the multimodal text-image fusion representation and perform adaptive feature conversion on the received fusion representation through a multi-layer perceptron to generate a multimodal text-image fusion vector of historical sales and real-time fan sentiment; Receive multimodal text-image fusion vectors, combine sales data, tag type, and number of fans, predict product demand based on decision theory, and output the prediction results.
Citation Information
Patent Citations
Multi-modal social media sentiment analysis method based on feature fusion
CN114020871A
Transform-based image-text multi-modal system operation and maintenance knowledge matching method and system
CN117788987A
Network public opinion emotion situation quantification method and system
CN118885870A
Multi-mode-based method for classifying image text pairs in network public opinions
CN119128592A
Public opinion analysis method and device based on big data and deep learning and medium
CN119475232A
Cited By
Public opinion evolution prediction method and system based on multi-dimensional user portrait and adaptive graph fusion
CN120832645A
Marketing intelligent decision-making system based on multi-modal learning
CN120975862A