Market situation awareness and decision-making method based on space-time diagram convolution and causal inference
By constructing a unified data acquisition framework and a spatiotemporal graph convolutional network, combined with causal inference and twin networks, the problem of enterprises being unable to accurately perceive and make decisions in complex market environments has been solved, achieving efficient and low-cost market situation perception and decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-03-10
AI Technical Summary
Enterprises are unable to accurately perceive market changes in a complex and ever-changing market environment and make timely and correct decisions. Existing technologies suffer from problems such as limited data collection dimensions, insufficient data processing capabilities, limited accuracy in market situation perception and trend prediction, and high cost of decision verification.
We construct a unified data acquisition framework, integrate multimodal data, conduct data quality assessment and cleaning through machine learning and deep learning, build a cross-departmental interconnected data network, use spatiotemporal graph convolutional networks to capture dynamic changes in the market, and combine causal inference and twin networks to conduct virtual decision verification, thereby achieving reinforcement learning to optimize decision-making.
It improves the accuracy of market situation awareness and the scientific nature of decision-making, reduces the cost of decision verification, and enhances the competitiveness of enterprises in complex market environments.
Smart Images

Figure CN121638873A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis, and in particular to a market situation perception and decision-making method based on spatiotemporal graph convolution and causal inference. Background Technology
[0002] In today's complex and ever-changing market environment, enterprises face the dual challenges of explosive growth in data and rapid evolution of market dynamics. How to accurately perceive market changes and make scientific decisions has become a core issue for the survival and development of enterprises. Traditional enterprise market analysis and decision-making models have significant limitations: First, data collection dimensions are limited, relying heavily on structured sales and financial data while underutilizing unstructured data such as images, audio, and video, leading to blind spots in market perception. For example, multimodal data rich in market information, such as user behavior images captured by store cameras and product review audio on social media, are often not effectively integrated and analyzed. Second, data processing capabilities are insufficient. Faced with massive amounts of multimodal data, traditional methods struggle to achieve high-quality cleaning and correlation mining. Issues such as missing, conflicting, and outdated data, as well as semantic gaps between different modalities, pose significant obstacles to building comprehensive market perception models. Third, the accuracy of market situation perception and trend prediction is limited. Existing models are mostly based on single time series or static feature analysis, lacking the ability to capture spatiotemporal relationships and struggling to cope with nonlinear and sudden market changes. Furthermore, decision-making often relies on experience-based judgment, lacking in-depth analysis of the causal relationships of market changes, resulting in insufficient scientific rigor and foresight in decision-making. Fourth, decision verification is costly and risky. Directly testing different decision-making schemes in the real market may result in huge cost losses, and it is difficult to verify the effectiveness of multiple strategies in a short period of time. To address these issues, there is an urgent need for an intelligent method that can integrate multimodal data, accurately perceive market dynamics, uncover causal relationships, and validate decisions through virtual scenarios, thereby improving enterprises' market response speed and decision-making quality. Summary of the Invention
[0003] The main objective of this invention is to provide a market situation perception and decision-making method based on spatiotemporal graph convolution and causal inference, aiming to solve the technical problem in the prior art that enterprises cannot accurately perceive market dynamics and cannot make correct market decisions in a timely manner.
[0004] To achieve the above-mentioned objectives, this invention proposes a market situation perception and decision-making method based on spatiotemporal graph convolution and causal inference, comprising the following steps: Construct a unified data acquisition framework to integrate the acquisition of traditional structured data as well as multimodal data including images, audio, and video; A data quality assessment model is pre-built using machine learning algorithms, and the multimodal data is processed using a deep learning-based data cleaning algorithm to obtain reliable multimodal data. Cluster analysis is performed on different reliable multimodal data at multiple levels to mine the similarity of data attributes and the semantic association and potential relationship between different modal data, and to construct a cross-departmental associated data network that includes clustering results and data association relationships; Based on the clustering results and data associations in the cross-departmental data network, the different types of clustered data are mapped onto a spatiotemporal graph structure. Through graph convolution and temporal convolution operations, a situational awareness model based on a spatiotemporal graph convolutional network is constructed to capture dynamic changes in the market and generate a profile of the enterprise's operational status. In the spatiotemporal graph structure, the data nodes correspond to the various types of clustered data, and the edges correspond to the associations between the data. The market trend prediction results are obtained by analyzing the operational status profile of the enterprises. If the forecast results trigger adjustments to market decisions, the causal relationships of market changes will be explored using a pre-defined causal inference method. Based on historical market data, enterprise operational status profiles, and the causal relationships of market changes discovered, a virtual decision-making scenario that matches the real environment is constructed through twin networks to simulate virtual market feedback and virtual enterprise performance corresponding to different decision-making behaviors. A reinforcement learning environment is constructed, in which corporate decision-making behavior in a virtual scenario is taken as actions, and virtual market feedback and virtual corporate performance output by the twin network are taken as reward signals. The reinforcement learning iteration is completed in the virtual environment to generate the optimal decision solution that has been verified in the virtual environment.
[0005] Beneficial effects: This application presents a market situation perception and decision-making method based on spatiotemporal graph convolution and causal inference. By constructing a unified data acquisition framework, it integrates traditional structured data with unstructured data such as images, audio, and video, breaking down data silos. Combining machine learning data quality assessment and deep learning cleaning algorithms enhances data credibility, providing comprehensive and high-quality foundational data for subsequent analysis and addressing the problems of single data dimensions and insufficient utilization in traditional methods. Through multi-level clustering analysis, a cross-departmental data network is constructed, mapping the data onto a spatiotemporal graph structure. The spatiotemporal graph convolutional network captures the spatiotemporal correlations of market dynamics, generating a comprehensive enterprise operational status profile that reflects the enterprise's operational status. This overcomes the shortcomings of traditional models in perceiving market nonlinearity and sudden changes, improving the accuracy of market situation perception. The market trend prediction results obtained from the operational status profile analysis are combined with causal inference to uncover the causal relationships of market changes, enabling decision-making to move beyond experience-based judgment and instead rely on data-driven causal cognition. This enhances the scientific rigor and foresight of decision-making, reducing the risks associated with blind decision-making. By constructing virtual decision-making scenarios using twin networks, the market feedback and corporate performance of different decisions are simulated. Reinforcement learning is used to iteratively optimize decision-making solutions within this virtual environment, eliminating the need for repeated trials in the real market. This significantly reduces decision verification costs and allows for rapid selection of the optimal decision-making solution, improving optimization efficiency. This method integrates multimodal data processing, artificial intelligence, and virtual simulation technologies, forming a complete intelligent process from data collection, processing, and analysis to decision generation and verification. It helps companies move away from traditional, extensive decision-making models, achieving precise decision-making based on data and intelligent algorithms, and enhancing their competitiveness in complex market environments. Attached Figure Description
[0006] Figure 1 A flowchart illustrating a market situation perception and decision-making method based on spatiotemporal graph convolution and causal inference, according to an embodiment of the invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0007] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0008] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any modules and all combinations of one or more associated listed items.
[0009] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0010] Reference Figure 1 This invention provides a market situation perception and decision-making method based on spatiotemporal graph convolution and causal inference, comprising the following steps: S1. Construct a unified data acquisition framework to integrate the acquisition of traditional structured data as well as multimodal data including images, audio, and video.
[0011] In this step, the unified data acquisition framework described above is a systematic architecture that can integrate multiple data sources, adapt to different data transmission protocols and formats, and achieve centralized acquisition of multiple types of data. Multimodal data refers to data with different data representation forms, such as traditional tabular structured data, as well as unstructured data such as images (e.g., product promotional images, store scene images), audio (e.g., customer service recordings, product introduction audio), and video (e.g., advertising videos, corporate event videos).
[0012] In this embodiment, a data collection architecture compatible with multiple data types and transmission protocols is established. Data is collected from multiple internal and external sources, such as Enterprise Resource Planning (ERP) systems, Customer Relationship Management (CRM) systems, social media platforms, and IoT devices. This includes financial statement data, user review texts, product images, customer service call recordings, and marketing videos, breaking down data silos and laying a data foundation for subsequent comprehensive market analysis. This step enables the aggregation of all market-related data for the enterprise, allowing it to observe the market from a richer perspective, avoiding market misjudgments due to data gaps, and providing complete "raw materials" for accurately perceiving market changes.
[0013] S2. A data quality assessment model is pre-built using machine learning algorithms, and the multimodal data is processed using a deep learning-based data cleaning algorithm to obtain reliable multimodal data.
[0014] In this step, machine learning algorithms (data quality assessment models) are used to learn and model the accuracy, completeness, and consistency of data. For example, algorithms such as decision trees and random forests are used to build models that can automatically judge the quality of data. Deep learning (data cleaning algorithms) refers to using deep neural networks, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to identify and repair noise, missing data, and errors. Examples include using generative adversarial networks (GANs) to fill in missing parts of image data and using the Transformer model to correct errors in text data.
[0015] In this embodiment, a data quality assessment model is first trained using machine learning algorithms. Historical multimodal data and its quality labels (such as accurate / inaccurate, complete / incomplete, etc.) are input, allowing the model to learn data quality judgment rules. Then, a deep learning-driven cleaning algorithm is used to clean the assessed low-quality data, performing tasks such as image denoising, audio noise removal, text typos correction, and filling in missing values in structured data, outputting reliable data. This step improves data quality, addresses the issue of inconsistent quality caused by the complex sources of multimodal data, provides reliable data support for subsequent analysis, avoids "garbage in, garbage out," and ensures the accuracy of market perception and decision-making.
[0016] S3. Perform cluster analysis on different reliable multimodal data at multiple levels to mine the similarity of data attributes and the semantic association and potential relationship between different modal data, and construct a cross-departmental associated data network that includes clustering results and data association relationships.
[0017] In this step, cluster analysis is a data analysis method that divides a dataset into different clusters (classes) based on the inherent similarity of data attributes, such as numerical and semantic features, allowing similar data to be aggregated and different data to be distinguished. The cross-departmental data network is a network structure that uses data relationships as edges and data clustering results as nodes, presenting the cross-departmental data relationships within an enterprise and demonstrating the connections between business data from different departments.
[0018] In this implementation, the cleaned multimodal data is grouped using clustering algorithms such as K-means and DBSCAN at multiple levels, including the feature level (numerical features of structured data, visual features of image data, and semantic features of text data) and the semantic level (the correlation between the meanings expressed by different data). Simultaneously, connections between different modalities are explored, such as analyzing the correlation between product image style and user evaluation sentiment, and the correlation between customer service recording topics and sales data fluctuations. This constructs a network covering data connections across various departments of the enterprise, clearly presenting the relationships between data. In this embodiment, the inherent patterns and correlations of the data are mined, and the cross-departmental business data connections within the enterprise are clarified, allowing the enterprise to understand the mutual influence of data from different business stages. This provides a structured data correlation view for perceiving the market from a holistic perspective and exploring potential market relationships.
[0019] S4. Based on the clustering results and data association relationships in the cross-departmental data network, the different types of clustered data are mapped to a spatiotemporal graph structure. Through graph convolution and temporal convolution operations, a situational awareness model based on a spatiotemporal graph convolutional network is constructed to capture dynamic changes in the market and generate a profile of the enterprise's operational status. In this model, the data nodes in the spatiotemporal graph structure correspond to the various types of clustered data, and the edges correspond to the association relationships between the data.
[0020] In this step, the spatiotemporal graph structure organizes data into a graph format according to time and space dimensions. Nodes represent data objects (such as sales data for a certain type of product or user profile data for a certain region), and edges represent the spatiotemporal relationships between data (such as the correlation between product sales data in different regions over time). The spatiotemporal graph convolutional network is a neural network model that integrates graph convolution (processing graph-structured data and capturing relationships between nodes) and temporal convolution (processing time-series data and capturing changes in the time dimension), effectively mining spatiotemporally related data features. The enterprise operational status profile is a comprehensive digital presentation of the enterprise's operational status in the market, covering key indicators and characteristics across multiple dimensions such as sales, production, users, and market competition, including sales trends, user activity distribution, and product market share.
[0021] In this embodiment, based on the clustering and association results of the cross-departmental data network, various types of data are mapped to a spatiotemporal graph structure to determine nodes (data classes after clustering) and edges (data relationships). A spatiotemporal graph convolutional network is used to extract spatiotemporal features from the graph structure data. Graph convolution operations capture the spatial relationships between data nodes, and temporal convolution operations capture the dynamic changes of data over time. A model is built to perceive market trends, ultimately outputting a state profile that reflects the overall operational situation of the enterprise. This step can accurately capture the spatiotemporal characteristics of dynamic market changes, and the generated operational profile allows enterprises to intuitively and comprehensively understand their own operational position and performance in the market, providing a clear basis for judging market trends and identifying operational problems. S5. Analyze the enterprise operation status profile to obtain market trend prediction results.
[0022] This step utilizes data analysis and prediction algorithms, such as time series analysis (ARIMA, LSTM, etc.), regression analysis, and machine learning classification prediction, to process multi-dimensional data in the operational status profile. This process uncovers data trends (such as sales growth / decline trends, user number changes), periodicity (such as seasonal product sales cycles), and correlations (such as the impact of marketing investment and market share on trends). The aim is to predict future market trends, including changes in market demand, competitor activities, and industry development trends. This provides businesses with a basis for anticipating market changes, allowing them time to prepare response strategies and avoid reactive responses to sudden market shifts, thus enhancing their market responsiveness. S6. If the prediction results trigger adjustments to market decisions, the causal relationships of market changes will be explored using a pre-defined causal inference method.
[0023] In this step, the causal inference method is a set of analytical methods for identifying causal relationships between variables, distinguishing between correlation and causation, and clarifying the driving factors and results of market changes. Commonly used methods include propensity score matching, instrumental variable method, and causal graph model.
[0024] In this embodiment, when trend forecasts indicate that market changes necessitate adjustments to corporate decisions, a causal inference method is employed to analyze the causes and consequences of these changes from massive amounts of market data. This includes determining whether a decline in market share is due to product quality issues or competitor strategies, and uncovering the key causal chains influencing market changes. By identifying the root causes of market changes, companies can make targeted adjustments, avoiding blind changes and improving the effectiveness and relevance of their decisions. S7. Based on historical market data, enterprise operational status profiles, and the causal relationships of market changes discovered, virtual decision-making scenarios that match the real environment are constructed through twin networks to simulate virtual market feedback and virtual enterprise performance corresponding to different decision-making behaviors.
[0025] In this step, the twin network consists of two structurally identical or similar sub-networks: one processes real data, and the other processes virtually generated data. Through mutual learning and feedback, the virtual scenario approximates the real environment. The virtual decision-making scenario utilizes virtual simulation technology, combining real market data and enterprise operational data to construct a simulated market environment, which can simulate the market response after different decisions are implemented.
[0026] In this embodiment, historical market data, enterprise operational status profiles, and mined causal relationships are input to drive the construction of a twin network in a virtual environment consistent with the operating logic of the real market. Different decision variables (such as price adjustments, marketing campaign planning, and product feature optimization) are set in the virtual environment to simulate the feedback in the virtual market after the decisions are implemented, including changes in user purchasing behavior, potential competitor reactions, market share fluctuations, and changes in enterprise performance such as sales revenue, profits, and costs. This allows enterprises to preview the effects of different decisions with low risk and low cost, avoiding losses from direct trial and error in the real market, and providing a simulated verification environment for selecting high-quality decisions. S8. Construct a reinforcement learning environment, take the enterprise decision-making behavior in the virtual scenario as actions, and use the virtual market feedback and virtual enterprise performance output by the twin network as reward signals to complete the reinforcement learning iteration in the virtual environment and generate the optimal decision solution that has been verified in the virtual environment.
[0027] In this step, the reinforcement learning environment is an environment in which the intelligent agent (which can be understood here as the subject of decision-making strategy learning) makes decision attempts and obtains feedback rewards in order to optimize the decision-making strategy. It includes elements such as state (such as market environment state, enterprise operation state), action (such as decision-making behavior that the enterprise can take), and reward (such as market feedback and performance resulting from the decision).
[0028] In this embodiment, a virtual decision-making scenario is used as a reinforcement learning environment. A state space (based on enterprise operation and market state data) and an action space (the set of all possible decision-making behaviors of the enterprise) are defined within the environment. Virtual market feedback and enterprise performance are transformed into reward signals. The agent continuously tries different decision-making actions in the environment, adjusting its decision-making strategy based on the reward signals. Through multiple rounds of iterative learning, the optimal decision-making solution in the virtual environment is found. Leveraging the self-optimization capability of reinforcement learning, the optimal solution is selected from numerous possible decisions, ensuring the scientific validity and effectiveness of the decision-making process and improving the quality and competitiveness of the enterprise's market decisions.
[0029] In a specific embodiment, taking a chain retail enterprise as an example, the process of applying this method is as follows: Data Acquisition: Through a unified framework, structured data on sales and inventory of each store are collected from ERP; user review text, product display images, and advertising videos are collected from online platforms; multimodal data such as store scene videos are collected from store cameras and call audio are collected from customer service systems are collected.
[0030] Data processing: A quality assessment model is built using random forest, and combined with GAN-based image inpainting and Transformer-based text correction algorithms to process multimodal data and obtain reliable multimodal data.
[0031] Clustering and Network Construction: Clustering reliable multimodal data, such as clustering sales data from different stores and user review data from corresponding regions, to uncover associations such as "high-selling stores - concentration of positive reviews", and to build cross-departmental (such as sales department, marketing department, customer service department) related data networks.
[0032] Situational awareness and profile generation: Mapping data to a spatiotemporal graph structure (nodes are various clustered data, edges are relationships), using spatiotemporal graph convolutional networks for analysis, generating operational status profiles covering sales trends, user satisfaction, market competition, etc. for each store.
[0033] Trend prediction: Analysis of customer profiles revealed a continuous decline in sales and an increase in negative customer reviews in a certain region, predicting that the market share in that region may be further lost.
[0034] Causal inference: It was determined that the decline in service quality of stores in the region (customer service recordings reflected service attitude problems, and user reviews mentioned poor service) led to customer loss and a drop in sales.
[0035] Virtual decision-making scenario: A virtual market is built using twin networks to simulate decisions such as "strengthening employee service training" and "launching promotional activities". After the "service training" decision is made, user reviews improve and sales rebound in the virtual scenario.
[0036] Reinforcement learning optimization: In a virtual environment, decision-making behavior is the action, and market feedback and performance are the rewards. After iterative learning, the optimal decision-making plan of "first strengthening service training, and then coordinating with small-scale promotional activities" is generated. After application, the market share in the region gradually recovers and sales increase.
[0037] This embodiment of the market situation perception and decision-making method based on spatiotemporal graph convolution and causal inference achieves deep integration and high-quality processing of multimodal data at the data level, providing a comprehensive and reliable data foundation for market analysis and solving the problems of single and poor-quality traditional data. At the perception and prediction level, it accurately captures dynamic changes in the spatiotemporal landscape of the market, generates comprehensive operational profiles, improves the accuracy of market trend prediction, allows enterprises to plan ahead, and enhances their foresight in responding to the market. At the decision-making level, it uncovers causal relationships to make targeted decisions, and virtual scenarios and reinforcement learning help select the optimal decisions, reduce the cost of trial and error, improve the scientific nature and effectiveness of decisions, enhance the market competitiveness and operational management level of enterprises, and help enterprises make accurate decisions and operate efficiently in complex market environments.
[0038] In one embodiment, after generating the virtually validated optimal decision scheme, the following steps are included: The situation awareness model is continuously optimized based on a preset online learning module and a meta-learning module. The online learning module captures new data features and market change information in real time to incrementally update the situation awareness model. The meta-learning module automatically adjusts the hyperparameters of the learning algorithm and the model structure by analyzing the historical learning process and the performance of the situation awareness model. At the same time, it dynamically adjusts the decision-making scheme based on the model optimization results and real-time market feedback.
[0039] In this implementation, the online learning module can receive and process newly generated data in real time, and is an algorithm module that incrementally updates the existing situational awareness model without retraining the entire model. For example, it can continuously optimize model parameters using new data based on online learning algorithms such as stochastic gradient descent. Incremental updates differ from traditional full-data retraining; they only target newly acquired data features (such as emerging user behavior patterns, sudden policy impacts on the market, etc.) to locally and gradually adjust and optimize the model's parameters, ensuring that the model can adapt to new data in a timely manner while saving computational resources. The online learning module constantly monitors new data generated during enterprise operations, including newly collected multimodal data (such as newly generated user purchase behavior images, the latest market research audio feedback, etc.), extracting new features from this data (such as new user consumption preference features, new market competition trends). Then, using online learning algorithms, such as online random forests and online gradient descent, these new features are integrated into the existing situational awareness model to fine-tune the model parameters. For example, when new user feedback characteristics on products are discovered on social media platforms (such as product feedback in the form of short videos), the online learning module updates the parameters corresponding to these characteristics in the model, enabling the model to recognize and process this new form of market feedback. This allows the situational awareness model to keep pace with dynamic market changes, promptly incorporating market information brought by new data, and avoiding misperceptions of the market due to data lag. For instance, in the event of a sudden market crisis, new data can quickly allow the model to adjust its assessment of market risk, enabling corporate decisions to be based on the latest market understanding, improving the timeliness and accuracy of decision-making.
[0040] The aforementioned meta-learning module is an algorithmic module that "learns how to learn better." It takes the model's learning process (such as changes in the loss function and accuracy fluctuations during model training) and model performance (such as the accuracy of the situational awareness model's market trend predictions and the comprehensiveness of its operational status profile) as its learning objects. It automatically optimizes the model's hyperparameters (such as the number of convolutional kernels and learning rate in spatiotemporal graph convolutional networks, and the reward function weights in reinforcement learning) and model structure (such as adjusting the number of layers and neuron connections in neural networks). Hyperparameters are parameters that need to be set before model training. Unlike the parameters learned during model training, they control the model's learning process and structure, significantly impacting model performance; examples include the depth of decision trees and the learning rate of neural networks. The meta-learning module continuously tracks the training and application process of the situational awareness model, collecting various data from its historical learning, including the loss value at each model update, the error between the prediction results and the actual market situation, and the model training time. By analyzing this data, it determines whether the current model's hyperparameters and structure are suitable for the market data characteristics and task requirements. For example, when a situational awareness model is found to have low accuracy in predicting a certain type of market trend (such as the market expansion trend of emerging industries), the meta-learning module analyzes that this is because the model's temporal convolution operation parameters are set unreasonably, or the structure of the graph convolutional network cannot effectively capture the spatiotemporal correlation of this type of trend. It then automatically adjusts the corresponding hyperparameters (such as increasing the size of the temporal convolution kernel) or optimizes the model structure (such as adding layers in the graph convolutional network to capture long-distance correlations). Simultaneously, the meta-learning module summarizes historical optimization experience, forming "knowledge" of the optimal model settings for different market scenarios, enabling faster and more accurate model optimization in subsequent market changes. This allows the situational awareness model to continuously evolve, shifting from "passively" adapting to market data to "actively" optimizing itself to better process market data. For instance, when market data characteristics shift from primarily traditional sales data to primarily multimodal social data, the meta-learning module can adjust the model structure and hyperparameters to maintain good market awareness performance, ensuring that corporate decisions are always based on high-quality situational awareness results.
[0041] After optimization through online and meta-learning modules, the situational awareness model outputs new market situational awareness results (such as a more accurate profile of enterprise operational status and more accurate market trend predictions). Simultaneously, real-time market feedback (such as actual market sales changes after decision implementation, immediate reactions from competitors, and genuine user feedback) is continuously generated. Enterprises combine these optimized model outputs with real-time market feedback to re-evaluate existing decision-making plans. For example, a decision based on the old model to "increase advertising in a certain region" might be adjusted to "conduct product experience activities in that region" based on the new situational awareness results and market feedback, after model optimization reveals that the market trend in that region is that users are more focused on product experience than advertising, and real-time market feedback also shows that sales growth after advertising is not significant. This creates a closed loop of "data collection - model awareness - decision-making - market feedback - model optimization - decision adjustment," allowing enterprise decision-making plans to continuously adapt to market changes. This avoids rigidly implementing decisions once they are made, enabling continuous optimization based on dynamic market evolution and the model's more accurate understanding of the market, improving the scientific nature and effectiveness of decisions, and helping enterprises maintain a competitive advantage in a complex and ever-changing market.
[0042] In this embodiment, optimization enables the situational awareness model to better process collected multimodal data and uncover more valuable market information. The optimized model can more accurately capture market dynamics, generating more realistic profiles of enterprise operations and market trend predictions. Dynamic adjustments to decision-making schemes allow enterprises to continuously optimize their decisions in the real market, building upon the constructed virtual decision verification. Ultimately, this makes the market situational awareness and decision-making method based on spatiotemporal graph convolution and causal inference a self-evolving, continuously adaptable intelligent system, significantly improving enterprises' ability to cope with complex market environments, enhancing their market competitiveness and the scientific nature of their decisions, and helping them maintain an advantage in long-term market operations, accurately grasping market opportunities and mitigating risks. For example, after the aforementioned retail enterprise generates a decision scheme of "enhanced service training + small discounts," new multimodal data, such as user review audio and store operation videos, are continuously generated as the market changes. The online learning module captures the characteristic in the new data that "young users pay more attention to online service experiences (such as online customer service response speed and online product display effects)," and updates the parameters of the situational awareness model. The meta-learning module discovered that the original situational awareness model was insufficient in predicting the trend of online services' impact on market sales. Adjustments were made to the model's hyperparameters (e.g., increasing the weight of online service features in the graph convolutional network) and structure (e.g., adding a temporal convolutional layer specifically for capturing the temporal features of online service feedback). Based on the optimized model, the company found that the original decision-making plan's "small discounts" had limited appeal to young users. Combining real-time market feedback (no significant increase in online orders after the promotional activity), the decision was dynamically adjusted to "strengthening online service training + launching exclusive online discount packages." After implementation, online sales in the region increased significantly, and the company's market competitiveness was further enhanced, demonstrating the synergistic value of the optimized overall approach.
[0043] In one specific embodiment, the above-described unified data acquisition framework, which integrates the acquisition of multimodal data S1 including traditional structured data and images, audio, and video, includes: S11. Establish a distributed data acquisition hub, which includes a protocol adaptation layer, a data conversion layer, and a unified storage layer. The protocol adaptation layer supports multiple data transmission protocols such as SQL, FTP, HTTP, and MQTT, enabling integration with internal enterprise systems (ERP, CRM) and external data sources (social media APIs, IoT devices). The data conversion layer uses preset modal conversion rules to uniformly convert data from different sources into a compatible intermediate format. The unified storage layer adopts a hybrid database architecture (relational database stores structured data, and distributed file system stores unstructured data), and uses a data indexing mechanism to achieve multimodal data association queries.
[0044] In this step, the distributed data acquisition hub is a system hub with a layered architecture (protocol adaptation layer, data conversion layer, and unified storage layer). It can be deployed in a distributed manner and work collaboratively to achieve efficient acquisition, conversion, storage, and correlation management of multi-source and multi-type data. The protocol adaptation layer is responsible for adapting different data transmission protocols, allowing the hub to interface with various data sources (such as SQL protocol used by internal ERP systems and MQTT protocol used by IoT devices), acting as a "translator" to ensure that data transmitted using different protocols can be recognized and received by the hub. The data conversion layer, based on preset rules, converts data of different formats and modalities (such as tabular data from ERP systems and image data from cameras) into a unified and processable intermediate format, solving the problem of data "language incompatibility." The unified storage layer adopts a hybrid database architecture, using a relational database (such as MySQL) to store structured data and a distributed file system (such as HDFS) to store unstructured data (such as video and audio). It also uses a data indexing mechanism (such as Lucene-based indexes) to correlate different modalities of data, facilitating subsequent queries.
[0045] In this embodiment, the hierarchical architecture and deployment method of the planning hub are as follows: At the protocol adaptation layer, interfaces and conversion logic supporting protocols such as SQL, FTP, HTTP, and MQTT are configured; at the data conversion layer, modal conversion rules are written or trained (e.g., using Python scripts to convert data formats, or using deep learning models to convert image semantics); and at the unified storage layer, a hybrid database environment is built, and a data indexing mechanism is configured. Once completed, the hub can connect to internal and external data sources within the enterprise, receiving, converting, and storing multimodal data, providing fundamental support for subsequent data processing. This breaks down the "island" state of data collection, allowing internal and external, multi-protocol, and multimodal data to converge in an orderly manner, providing a unified data entry point for comprehensive market analysis, ensuring the smooth flow of subsequent multimodal data processing, and avoiding analytical biases caused by scattered data collection and inconsistent formats. S12. By connecting to the Enterprise Resource Planning (ERP) system, Customer Relationship Management (CRM) system and transaction database through the Structured Query Language (SQL) interface, sales data, inventory data, customer information and financial indicators are collected. Data validation rules (such as field format validation and range validation) are used to perform preliminary screening to obtain traditional structured data. The traditional structured data is then synchronized to the relational database of the unified storage layer on a regular basis through ETL tools.
[0046] In this step, the Structured Query Language (SQL) interface is used to interact with relational databases and perform data query, insert, update, and delete operations. Databases in enterprise ERP and CRM systems often provide data access externally through the SQL interface. Enterprise Resource Planning (ERP) systems are management systems that integrate internal business processes such as finance, procurement, production, and sales, storing enterprise operational information in structured data format. Customer Relationship Management (CRM) systems focus on enterprise customer management, storing structured data such as customer information, sales leads, and service records. ETL tools, which are Extract, Transform, and Load tools, are used to extract data from source databases, transform and clean it, and then load it into target databases, such as Kettle and Informatica.
[0047] Connect to ERP, CRM, and transaction databases using an SQL interface, and write SQL queries (e.g., "SELECT * FROM sales_data WHERE time > '2024-01-01'") to extract structured data such as sales, inventory, customer, and financial data. Set data validation rules (e.g., check if the amount field in sales data is numeric and if the date field is within a reasonable range) to filter out compliant data. Then, use an ETL tool to synchronize this data to a relational database in a unified storage layer at a set time (e.g., every morning) to ensure timely data updates. Stable collection of structured data from core enterprise operations, validated and synchronized to ensure data quality and timeliness. This allows enterprises to understand basic operational conditions such as sales and inventory based on standardized structured data, providing quantifiable and statistically relevant data for market analysis, such as financial indicator analysis and sales trend statistics. S13. By deploying edge acquisition nodes (such as store cameras and product display terminals), key visual features (such as product appearance and user behavior) are extracted in real time using an attention-based target detection algorithm to obtain image data, which is then compressed and uploaded to a distributed file system. Audio streams of user inquiries and meeting records are acquired through voice acquisition devices, and feature transformation is performed using Mel-frequency cepstral coefficients (MFCC). After filtering invalid audio segments using a voice activity detection (VAD) algorithm, the audio data is stored. A segmented acquisition strategy is adopted, and video data is obtained by separating intra-frame spatial features and inter-frame temporal features through a spatiotemporal feature extraction network (such as I3D), which is then stored in a distributed file system according to time slices. In this step, the aforementioned edge acquisition nodes refer to devices or modules deployed at the network edge (such as stores or production workshops), close to data sources (such as cameras or display terminals), and capable of data acquisition and preliminary processing (such as extracting image features). This reduces the pressure on the central server and allows for rapid response to data acquisition needs. Attention-based target detection algorithms, in image recognition, mimic human attention, focusing on key targets in an image (such as product appearance or user behavior) to accurately detect and extract features. For example, YOLOv5, combined with an attention mechanism, can more accurately identify important elements in an image. Mel-frequency cepstral coefficients (MFCCs) are coefficients used in audio processing that simulate human hearing characteristics to extract audio features (such as timbre and pitch), commonly used in audio recognition and sentiment analysis. Voice activity detection (VAD) algorithms are used to detect whether an audio stream contains speech, filtering out silent, noisy, or other invalid segments. For example, in customer service recordings, only segments containing user inquiries and customer service responses are retained. Spatiotemporal feature extraction networks (such as I3D) are deep learning networks that can simultaneously extract intra-frame spatial features (such as image content) and inter-frame temporal features (such as the order of action changes) from video data, enabling them to understand the spatiotemporal dynamic information of videos.
[0048] Specifically, image data is collected by deploying edge acquisition nodes (such as smart cameras) in stores and product display areas, running attention-based target detection algorithms (such as improved Faster R-CNN) to identify key visual features such as products and user behavior in the images in real time, compressing the processed image data (such as using JPEG-XL compression), and uploading it to a distributed file system for storage via the network.
[0049] Audio data is collected using voice acquisition devices (such as microphone arrays) to capture audio streams such as user inquiries and meeting records. First, the audio features are extracted using the MFCC algorithm, and then the VAD algorithm is used to detect and filter silent and noisy segments. After obtaining valid audio data, it is stored in a designated location (such as the audio storage directory of the server).
[0050] Video data is acquired using a segmented acquisition strategy (e.g., by hour). Videos are captured using devices such as cameras and processed by a spatiotemporal feature extraction network (e.g., I3D) to separate intra-frame spatial features (e.g., the content of a particular frame) and inter-frame temporal features (e.g., the order of a user's actions from browsing to purchasing in the video). The data is then stored in a distributed file system in time slices (e.g., one slice per hour) for easy analysis by time dimension.
[0051] This embodiment achieves effective collection and preliminary processing of unstructured data, extracting key features (such as user behavior and audio emotion), providing a foundation for subsequent mining of market information from unstructured data (such as user visual feedback on products and needs expressed in audio). It enables businesses to obtain market information from images, audio, and video that cannot be covered by structured data; for example, by analyzing user behavior images in stores, store layout can be optimized; and by analyzing customer service audio, service scripts can be improved. S14. Based on a preset multimodal data weighting model, the data collection priority is dynamically adjusted according to the data's contribution to market perception. The contribution calculation formula is as follows:
[0052] in, The weights for the m-th modal data (including the weights corresponding to traditional structured data). This is the integrity coefficient of the modal data at time t (with a value of 0-1). Let M be the market relevance of this modality at time t (obtained through training on historical decision influence), M be the total number of modalities, and T be the data collection time window. The scheduling module is based on... Allocate network bandwidth and computing resources to ensure that high-value data (such as user review texts directly related to sales and core sales data) are collected first.
[0053] In this step, the multimodal data weighting model calculates the weights of different modal data (such as structured sales data and image user behavior data) based on the data's contribution to market perception, and guides the allocation of data collection resources (bandwidth, computing power).
[0054] Integrity coefficient A coefficient that measures the completeness of a certain type of modal data (m) at a certain time (t), such as whether the collected sales data includes transaction records of all stores and all products. The value ranges from 0 to 1, and the more complete it is, the closer it is to 1.
[0055] Market Relevance By training with historical data, we can determine the degree of correlation between a certain type of modal data and market decisions and changes at a certain point in time. For example, user review text data has a high market correlation because it has a significant impact on product improvement decisions.
[0056] In this embodiment, a multimodal data weighting model is first trained by inputting the integrity coefficient, market relevance, and corresponding market decision-making effect data of historical multimodal data, allowing the model to learn to calculate weights. The formula is used to calculate the current modal data in real time during actual data acquisition. (e.g., using data validation tools to ensure statistical integrity) and (e.g., querying through a historical association model), substitute into the formula to calculate The scheduling module (such as Kubernetes-based resource scheduling) is based on... Allocate network bandwidth (e.g., allocate more bandwidth to high-weight data) and computing resources (e.g., allocate more computing power to image data processing), prioritizing the collection of high-value data (e.g., user review texts strongly correlated with sales). This ensures data collection resources are targeted and avoids wasting resources by blindly collecting low-value data. When the market changes, such as a sudden increase in the contribution of a certain type of data (e.g., user feedback images during a new product launch) to market perception, the system can automatically adjust resource allocation, prioritizing the collection of this type of data to ensure that companies can obtain key market information in a timely manner and improve their market perception acumen.
[0057] In one embodiment, the data quality assessment model pre-built using machine learning algorithms, combined with a deep learning-based data cleaning algorithm, is used to process the multimodal data to obtain reliable multimodal data, including: S21. Construct a data quality assessment model from four dimensions: accuracy, completeness, consistency, and timeliness. The formula for calculating the comprehensive quality score is as follows:
[0058] Where A is the accuracy score (calculated by comparing with the benchmark data), B is the completeness score (the complement of the proportion of missing values), D is the consistency score (the complement of the logical conflict rate of cross-modal data), E is the timeliness score (the normalized value of the time difference between data generation and collection), and μ, ν, ξ and ζ are the weights of each dimension (determined by the analytic hierarchy process, satisfying μ+ν+ξ+ζ=1).
[0059] In this step, accuracy (A) measures how well the data matches reality or benchmark data, such as the match between actual sales revenue and system-recorded sales revenue. It is calculated by comparing the data to known, accurate benchmark data (such as audited financial data). The smaller the difference, the higher the accuracy score. Completeness (B) reflects whether the data contains all the necessary information, expressed as the complement of the percentage of missing values. The more complete the data, the closer B is to 1. Consistency (D) focuses on whether the logic between cross-modal data is consistent, such as whether the product price in structured sales data is consistent with the product price in image advertisements. This is measured by the complement of the cross-modal data logical conflict rate. The fewer the conflicts, the higher the D. Timeliness (E) reflects the freshness of the data, calculated as a normalized value of the time difference between data generation and collection. For example, if the time difference is t, the maximum time difference is... ,but The smaller the time difference, the closer E is to 1. The Analytic Hierarchy Process (AHP) is an analytical method that decomposes complex decision problems into multiple levels and factors, determining the weights of factors through pairwise comparisons. It is used to determine μ, ν, ξ, and ζ, making the weight allocation more scientific.
[0060] This embodiment first clarifies the dimensions of multimodal data quality that enterprises focus on (accuracy, completeness, consistency, and timeliness). Then, using the analytic hierarchy process (AHP), it compares the importance of each dimension with other dimensions (e.g., judging the importance of accuracy and completeness to market analysis), constructs a judgment matrix, and calculates the weights μ, ν, ξ, and ζ for each dimension. Next, for the collected multimodal data, it calculates the accuracy score A (compared to benchmark data), the completeness score B (calculating the complement of missing values), the consistency score D (calculating the complement of conflict rates), and the timeliness score E (time difference normalization). Finally, it substitutes these values into the formula to calculate the overall quality score Q. This embodiment establishes a scientific and comprehensive data quality assessment standard, quantifies multimodal data quality, allows enterprises to clearly understand the "health" of their data, provides a basis for subsequent data cleaning and use, and avoids the impact of low-quality data on market perception and decision-making. S22. For low-quality data with a quality score Q < 0.6 (quality threshold), a generative adversarial network (GAN) is used to repair it and obtain reliable multimodal data. The GAN uses high-quality data as training samples, generates supplementary data through a generator, and a discriminator distinguishes between real and generated data. The process is iteratively optimized until the difference in distribution between generated data and real data is less than a preset threshold (e.g., JS divergence < 0.1).
[0061] In this step, the Generative Adversarial Network (GAN) is a deep learning model consisting of a generator and a discriminator. The generator is responsible for generating new data that resembles real data, while the discriminator is responsible for distinguishing between real and generated data. The two are trained adversarially, making the data generated by the generator increasingly closer to real data. The Jensen-Shannon Divergence (JS Divergence) measures the difference between two probability distributions; the smaller the value, the closer the two distributions are. Here, it serves as an indicator of whether the generated data is similar to the real data.
[0062] Specifically, low-quality data with a comprehensive quality score (Q < 0.6) is selected, while high-quality data (Q ≥ 0.6) is used as training samples and input into the generator and discriminator of the GAN. The generator learns the distribution characteristics of high-quality data and generates supplementary data (such as missing product details in image data or missing field values in structured data); the discriminator learns to distinguish between real high-quality data and data generated by the generator. Training is iterated continuously until the JS divergence between the generated data and real data is less than 0.1. At this point, the generated data can be used to supplement missing or erroneous parts of low-quality data, improving the quality of the repaired data. This embodiment can effectively repair low-quality multimodal data, solving quality problems caused by complex acquisition environments (such as incomplete image transmission due to network fluctuations or data loss due to sensor malfunctions). It allows previously "unusable" or "low-value" data to play a role again, enriching enterprise data assets, providing more complete and reliable data for subsequent market analysis, and ensuring the accuracy of market perception and decision-making.
[0063] In one specific embodiment, a retail enterprise, after collecting multimodal data, constructs an evaluation system: using the analytic hierarchy process (AHP), accuracy μ=0.3, completeness ν=0.25, consistency ξ=0.25, and timeliness ζ=0.2. After collecting a batch of store sales data (structured), product display images, and customer service call audio, scores for each dimension are calculated: Accuracy score A=0.8 for comparing sales data with audit data; completeness score B=0.9 for image data with 10% missing values; consistency score D=0.95 for cross-modal data (e.g., sales data price versus image price) with a 5% conflict rate; and timeliness score E=0.8 for a 1-hour time difference between data generation and collection, with a maximum time difference of 5 hours. Substituting into the formula, the overall quality score Q = 0.3 × 0.8 + 0.25 × 0.9 + 0.25 × 0.95 + 0.2 × 0.8 = 0.84, which is acceptable. However, another batch of image data collected due to network failure has poor integrity (B = 0.5) and low consistency (D = 0.6), resulting in a calculated Q = 0.3 × 0.7 + 0.25 × 0.5 + 0.25 × 0.6 + 0.2 × 0.6 = 0.58 < 0.6, classifying it as low-quality data. To repair the low-quality data, high-quality image data is selected as training samples for the GAN. The generator learns its features and generates missing product image details; the discriminator distinguishes between real and generated images. After multiple iterations, the JS divergence between the generated and real images is 0.08 < 0.1. The generated images are used to repair the low-quality image data, allowing the repaired data to be used for subsequent market analysis, such as analyzing user feedback on product presentations through the repaired images.
[0064] In one embodiment, the above-mentioned clustering analysis of different credible multimodal data at multiple levels is performed to mine the similarity of data attributes and the semantic associations and potential relationships between different modal data, and to construct a cross-departmental associated data network containing clustering results and data association relationships, including: S31. Density clustering (DBSCAN) is used for data of the same modality, where the neighborhood radius ϵ is dynamically adjusted according to the standard deviation of the data (ϵ=1.5×σ, where σ is the standard deviation of the feature data of that modality). Density-based clustering (DBSCAN) is a clustering algorithm based on data density. It can group densely connected data points into clusters, discover clusters of arbitrary shapes, and identify noise points. It is suitable for mining the inherent clustering structure in data of the same modality. The neighborhood radius ϵ is a radius parameter in the DBSCAN algorithm used to determine whether a data point is a core point, determining how far away other points around the data point can be considered as belonging to the same density region. The standard deviation σ reflects the statistical measure of the dispersion of the modal characteristic data; the larger the standard deviation, the more dispersed the data.
[0065] For single-modality data (such as all image data or all structured sales data), the features of that modality are first extracted (e.g., visual features are extracted using ResNet for image data, and numerical normalization is performed for structured data). The standard deviation σ of the feature data is calculated, and the neighborhood radius ϵ = 1.5 × σ is dynamically set. Then, the DBSCAN algorithm is run to cluster densely connected data points into a single class, forming clusters of data within the same modality, while simultaneously identifying noise points (e.g., outliers far from other data points). This step can uncover the similarity structure within the same modality data, aggregating similar data and distinguishing different data, laying the foundation for subsequent cross-modal association analysis. Dynamically adjusting the neighborhood radius makes the clustering more adaptable to the data distribution characteristics, avoiding poor clustering results due to a fixed radius (e.g., a fixed small radius cannot form reasonable clusters when data is scattered), and improving the accuracy of clustering data within the same modality. S32. Calculate the semantic similarity of clustering results across different modalities using a cross-modal attention mechanism. The similarity calculation formula is as follows:
[0066] in, For modality Clusters With mode Clusters The average cosine similarity; , Clusters , Feature vectors of internal data (if it is image data), or (This is the visual feature vector extracted by ResNet; if it is structured data, it is the normalized vector of numerical features). This represents the vector dot product operation. , The Euclidean norm of a vector; , Clusters , Number of data samples included; when When the similarity threshold is >0.7, a weighted edge is established between two clusters in the cross-departmental data network, with a weight value of [value missing]. This enables semantic association modeling of multimodal data (including images and other modal data) to obtain the cross-departmental association data network.
[0067] In this step, the cross-modal attention mechanism is a mechanism that can focus on key associations between data from different modalities and mine semantic connections. It can highlight data pairs that are important for semantic associations, suppress irrelevant data pairs, and improve the accuracy of cross-modal semantic analysis. The feature vector *f* is the vector representation obtained after feature extraction from the data. For example, the visual feature vector extracted from image data by ResNet can reflect the semantic content of the image; the normalized vector of numerical features of structured data can reflect the numerical distribution characteristics of the data, etc. Cosine similarity is an index that measures the degree of similarity between two vectors based on the cosine value of the angle between them. The closer the value is to 1, the more similar the vectors are. Here, it is used to calculate the semantic similarity between cross-modal clusters.
[0068] After clustering different modalities (such as image modality and structured data modality), clusters for each modality are obtained (such as image clusters). Sales data cluster Extract the feature vector of data within each cluster. (i belongs to) ), (j belongs to) Using a cross-modal attention mechanism, feature vector pairs that are important for semantic association are selected. Then, these pairs are substituted into the similarity calculation formula to calculate the mean cosine similarity between the two modal clusters. This step measures the semantic association between cross-modal clusters. Breaking through the limitations of a single modality, it mines semantic associations between data from different modalities, discovering the association between user purchasing behavior clusters in images and sales growth clusters in structured data. This provides a semantic connection basis for building a cross-departmental data network, enabling enterprises to understand market data relationships from a multimodal fusion perspective and improve the comprehensiveness of market perception. Furthermore, compared to existing formulas (such as those that only calculate the similarity of partial data pairs or do not perform averaging), the above similarity calculation formula can more comprehensively and evenly reflect the overall semantic association between two clusters of data, avoiding misjudgments of similarity due to biases in local data pairs. Through multimodal and cross-cluster calculations, it naturally supports cross-departmental data association (e.g., image modality belongs to the marketing department, structured data belongs to the sales department). The constructed cross-departmental association network can break down "departmental walls," allowing enterprises to mine market relationships from a global perspective (e.g., the association between the advertising creativity cluster in the marketing department and the performance cluster in the sales department). Existing formulas are insufficient to support the collaborative analysis needs of cross-departmental and multimodal data. This embodiment integrates multimodal data clustering results and semantic associations to form a cross-departmental and cross-modal data association view for the enterprise. Data from various departments within an enterprise (such as image advertising data from the marketing department and structured sales data from the sales department) are interconnected through a network, enabling the enterprise to uncover potential market relationships from a holistic perspective (such as the relationship between advertising creative clusters and sales performance clusters). This provides a comprehensive data-driven foundation for subsequent situational awareness model building and market decision-making, helping enterprises break down departmental silos and achieve data-driven collaborative decision-making.
[0069] Take a retail company's application of the above method to process multimodal data as an example: Low-level clustering: For image modal data (such as product display images of various stores), visual features are extracted using ResNet, the standard deviation of the feature data is calculated as σ=0.2, and the neighborhood radius is set as ϵ=1.5×0.2=0.3. The DBSCAN algorithm is run to cluster images displaying similar products and similar layouts into one class, forming image modal clusters (such as "high-end product display cluster" and "promotional area display cluster"). For structured sales data (such as sales amount and sales volume of each store), after numerical normalization, the standard deviation is calculated as σ=0.3, ϵ=0.45, and DBSCAN is run to cluster them into "high sales volume cluster", "low sales volume cluster", etc.
[0070] High-level correlation: "High-end product display cluster" for image modalities ( ) and structured data "high-volume sales clusters" ( Intra-cluster feature vectors are extracted, with image feature vectors derived from ResNet and structured data as numerically normalized vectors. Key data pairs are selected using a cross-modal attention mechanism and then calculated using the similarity formula: assuming... =140, =20, =10, then (Due to potential decimals in actual calculations, if the value is greater than 0.7), the two clusters are considered to have a strong semantic association. A relational network is constructed: In the cross-departmental relational data network, weighted edges are established for the "high-end product display cluster" and the "high-sales cluster," with a weight of 0.7 (assuming a weight greater than 0.7). Other modal clusters are also associated in the network (such as the audio modal customer service feedback cluster and the sales cluster), forming a cross-departmental, cross-modal data relational network for the enterprise. This network can be used to analyze market relationships such as "how high-end product displays promote high sales."
[0071] In one embodiment, based on the clustering results and data relationships in the cross-departmental data network, the different clustered data are mapped onto a spatiotemporal graph structure. Through graph convolution and temporal convolution operations, a situational awareness model based on a spatiotemporal graph convolutional network is constructed to capture dynamic market changes and generate a profile of the enterprise's operational status, including: S41. In the spatiotemporal graph structure, the node feature vectors are integrated with the statistical characteristics (mean, variance) of the clusters and the timestamp information, and the edge weights are dynamically updated based on the association strength in the cross-departmental data network.
[0072] In this step, the spatiotemporal graph structure integrates temporal and spatial (data association) dimensions. Nodes represent various data clusters after clustering, and edges represent the relationships between data clusters. It also incorporates temporal information to characterize the spatiotemporal dynamics of market data. Node feature vectors describe the features of nodes (data clusters), integrating statistical characteristics of clusters (such as the mean within a cluster reflecting the overall level and variance reflecting the degree of dispersion) and timestamp information (such as the time of data generation and the time of related events), allowing nodes to reflect the multi-dimensional attributes of the data. Edge weights reflect the numerical value of the strength of the association between data clusters, dynamically updated based on the association relationships in the cross-departmental data network; the stronger the association, the greater the weight.
[0073] Based on the clustering results of the cross-departmental data network, the nodes of the spatiotemporal graph structure are determined (one node for each cluster). For each node, the statistical features of the cluster are extracted (mean, variance, etc. are calculated), and combined with the timestamp information of the data generation (such as the transaction time corresponding to the sales data cluster), a node feature vector is constructed. For the edges between nodes, the edge weights are dynamically set according to the association strength of the corresponding clusters in the cross-departmental data network (such as the semantic similarity calculated above). Higher association strength results in higher edge weights, and vice versa, thus completing the construction of the spatiotemporal graph structure. This step transforms multimodal, cross-departmental data associations from a static network into a dynamic graph structure that integrates spatiotemporal information, providing inputs that fit the characteristics of market data for subsequent spatiotemporal graph convolution operations. This allows the model to simultaneously capture the spatial associations of data (such as the associations between data clusters from different departments) and temporal dynamics (such as changes in data over time), improving the comprehensiveness of market situation perception. S42. The graph convolution operation uses a modified Chebyshev polynomial approximation, and the calculation formula is as follows:
[0074] in, For the normalized Laplace matrix, It is a Chebyshev polynomial of order k. K is a learnable parameter, and K is the order of the polynomial (set to 3 to balance computational cost and feature capture capability).
[0075] In this step, graph convolution is used to process convolution operations on graph-structured data. It can extract neighborhood features of nodes in the graph and capture the relationships between nodes, similar to the role of traditional convolution in images, but adapted to the non-Euclidean nature of graph structures. Chebyshev polynomial approximation is a method that approximates the eigenvalue decomposition of the graph Laplacian matrix using Chebyshev polynomials. This reduces the computational complexity of graph convolution operations, allowing graph convolution to run efficiently on large-scale graph structures. Normalized Laplacian matrix. It is obtained by normalizing the graph Laplacian matrix and is used to describe the Laplacian properties of the graph structure, reflecting the local structural information of the nodes in the graph. Learnable parameters During graph convolution, adjustable parameters are trained to allow the model to learn the contribution of different Chebyshev polynomial terms to feature extraction.
[0076] In this embodiment, the normalized Laplace matrix is calculated for the constructed spatiotemporal graph structure. Then, using an improved Chebyshev polynomial approximation method, the polynomial order K is set to 3 (to balance computational cost and feature capture capability), and Chebyshev polynomials of each order are calculated. These polynomials are then combined with learnable parameters. Multiply and sum the vectors, then perform graph convolution on the node feature vectors X of the spatiotemporal graph structure to obtain the vector after feature extraction via graph convolution. This embodiment achieves the capture of spatial correlation features between nodes in a graph structure. It overcomes the limitation of traditional convolution being only applicable to Euclidean structures (such as images), efficiently processing non-Euclidean data with spatiotemporal graph structures and extracting spatial correlation features between nodes (such as the collaborative relationships between data clusters from different departments). The improved Chebyshev polynomial approximation reduces computational complexity while ensuring effective capture of graph structure features, allowing the model to mine spatial correlation patterns of data clusters at a reasonable computational cost, providing spatial dimension feature support for perceiving market trends.
[0077] S43. The temporal convolution operation uses a 1D convolutional layer, and the kernel size is dynamically adjusted according to the time granularity (7×1 kernel for daily granularity and 4×1 kernel for weekly granularity). Finally, a fully connected layer is used to output a profile of the enterprise's operational status containing 128-dimensional features.
[0078] In this step, the temporal convolution operation is performed on time-series data to extract features in the time dimension and capture trends in data over time, such as monthly fluctuations in sales data and time-based changes in user behavior. The 1D convolutional layer processes one-dimensional time-series data, using sliding convolution kernels to extract features. The temporal attention mechanism focuses on time segments important to market trends during temporal convolution, improving the extraction of key temporal features and suppressing interference from irrelevant time segments. Temporal granularity refers to the precision of time division, such as daily granularity (dividing time by day) or weekly granularity (dividing time by week). Different granularities reflect changes in market data over different time scales.
[0079] The node feature vectors after graph convolution are expanded along the time dimension (e.g., node features at different time points are arranged in chronological order). A 1D convolutional layer combined with a temporal attention mechanism is used, dynamically adjusting the convolution kernel size according to the time granularity: a 7×1 convolution kernel is used for daily granularity data (capturing daily changes within a week), and a 4×1 convolution kernel is used for weekly granularity data (capturing weekly changes within a month). After the convolution operation, the features are mapped to a 128-dimensional vector through a fully connected layer to generate a profile of the enterprise's operational status. In this embodiment, the dynamic changes in market data over time (such as seasonal fluctuations and trend changes) are captured, and the impact of key time segments is highlighted (such as changes in sales data during promotional activities) by combining the temporal attention mechanism. Dynamically adjusting the convolution kernel size adapts to data of different time granularities, allowing the model to accurately extract temporal features at multiple time scales, injecting dynamic time information into the profile of the enterprise's operational status, and improving the situational awareness model's ability to capture dynamic changes in the market.
[0080] In one embodiment, the analysis of the enterprise operational status profile described above to obtain market trend prediction results includes: S51. Extract core indicator time series (such as sales revenue, customer traffic, and conversion rate) from the enterprise operation status profile, and decompose them into trend items using the STL (Seasonal-Trend Decomposition Using Loess) algorithm. Seasonal items and residuals ,Right now For user behavior clustering features mined from image modal data, an attention mechanism is used to map them into auxiliary indicators of customer traffic trends. For user feedback emotional features from audio modal data, an emotional analysis model is used to map them into factors influencing sales conversion rates, and time series decomposition logic is incorporated.
[0081] In this embodiment, the core indicator time series refers to the sequence data of key enterprise operation indicators (such as sales revenue and customer traffic) changing over time, reflecting the historical evolution and trend of the indicators, and serving as the foundational data for market trend prediction. The STL algorithm is a time series decomposition algorithm based on Loess regression, which can decompose the time series into a trend term (long-term trend), a seasonal term (periodic repetitive pattern), and a residual term (random fluctuation), clearly breaking down the components of the time series. In multimodal data fusion, the attention mechanism focuses on image features (such as user behavior clustering features) that have a significant impact on core indicators (such as customer traffic), mapping them to auxiliary indicators to improve decomposition accuracy. The sentiment analysis model (audio modality) is a model that analyzes the sentiment tendency (positive, negative) of audio data (such as user consultation recordings), extracting user feedback sentiment characteristics to assess their impact on sales conversion rates. Step Description From the enterprise operational status profile, core indicators such as sales revenue, customer traffic, and conversion rate are selected, and their time-series data is extracted. Using the STL algorithm, the time series data of each indicator is decomposed into trend terms. Seasonal items residuals Simultaneously, for user behavior clustering features mined from image modality data (such as "users gathering to try out products"), the correlation between these features and customer traffic is calculated using an attention mechanism, mapped to a customer traffic trend auxiliary indicator, and integrated into trend decomposition. For user feedback sentiment features from audio modality (such as "users complaining about products"), a sentiment analysis model is used to determine sentiment tendencies, mapped to a sales conversion rate influencing factor, and integrated into conversion rate time series decomposition, making the decomposition results more closely reflect the actual market situation reflected by multimodal data. In this embodiment, the limitation of traditional time series decomposition relying solely on single indicator data is broken, and multimodal data features are incorporated, allowing the decomposed trend, seasonality, and residual terms to more accurately reflect actual market changes (such as image user behavior assisting in correcting customer traffic trends), laying the foundation for subsequent accurate prediction and improving the comprehensiveness and accuracy of market trend decomposition.
[0082] S52. A quadratic exponential smoothing method is used to predict the trend term, combined with a Long Short-Term Memory (LSTM) network to capture non-linear trends. The formula is as follows:
[0083]
[0084] in, This represents the trend prediction value at time t+1. The trend change rate is represented by α and β, which are smoothing coefficients (ranging from 0.3 to 0.5; α focuses on the impact of recent data on the trend, while β focuses on the continuity of the trend change rate. These coefficients can be dynamically adjusted through reinforcement learning by combining the market correlation of multimodal data). This complements the LSTM model's prediction of recent sequences.
[0085] In this step, the aforementioned double exponential smoothing refers to smoothing the rate of change of the trend again on top of the first exponential smoothing. It is suitable for time series forecasting with linear trends and can capture the continuity of the trend. Long Short-Term Memory (LSTM) networks are a variant of recurrent neural networks (RNNs) that excel at handling long-sequence data and can capture non-linear trends in time series (such as trend reversals caused by sudden market changes), compensating for the shortcomings of double exponential smoothing in handling non-linearity. The smoothing coefficients α and β are parameters in double exponential smoothing. α controls the degree of influence of recent data on trend prediction (a larger α results in higher weight for recent data), and β controls the continuity of the rate of change of the trend (a larger β means the trend change is more influenced by historical rates of change).
[0086] Trend terms obtained from STL decomposition First, quadratic exponential smoothing is used for prediction: initialization (Initial trend value) =0 (initial trend change rate), then iteratively calculate Update trend forecast values. Update the trend change rate. Simultaneously, extract recent sequences of the trend term (e.g., the last n time points), input them into the LSTM model, learn non-linear trend characteristics (e.g., trend abrupt changes caused by sudden market promotions), and output LSTM predicted supplementary values. The result is then combined with the quadratic exponential smoothing result to obtain the final prediction of the trend term. In this model, α and β can be dynamically adjusted using reinforcement learning, taking into account the market relevance of multimodal data (e.g., α is increased during promotional activities when relevance is high), to adapt to market changes. This embodiment combines quadratic exponential smoothing for trend continuation capture with LSTM for nonlinear trend learning, improving the accuracy of trend prediction. Dynamically adjusting the smoothing coefficient makes the prediction more adaptable to market dynamics (e.g., α is increased when market volatility is high). (Focusing on recent data) This addresses the problem of poor adaptability in traditional fixed-coefficient forecasting, providing more reliable trend judgments for market trend prediction. S53. The seasonal term is predicted using a periodic replication method combined with the Transformer model, and the residual term is predicted using the LSTM model. Finally, the three results are fused to obtain the market trend forecast value. The forecast error is evaluated by the mean absolute percentage error (MAPE). When MAPE < 10% (preset error threshold), it is considered a valid forecast. At the same time, ensemble learning is used to fuse the forecast results of multiple models to improve the forecast robustness.
[0087] In this step, the aforementioned periodic replication method (seasonal term) utilizes the periodicity of seasonal terms (such as annual or quarterly cycles) to replicate historical periodic patterns into the future. Combined with the long-sequence modeling capabilities of the Transformer model, it predicts the future cycle of the seasonal term. The Transformer model is a deep learning model based on a self-attention mechanism, adept at handling long sequences and capturing global dependencies, effectively predicting the periodic changes of seasonal terms. Ensemble learning combines the prediction results of multiple prediction models (such as quadratic exponential smoothing + LSTM for trend terms, and Transformer for seasonal terms) to reduce the error of a single model and improve prediction robustness. MAPE (Mean Absolute Percentage Error) is an indicator for evaluating prediction error; it calculates the mean percentage error between the predicted and actual values. A smaller value indicates a more accurate prediction and is used to determine prediction effectiveness.
[0088] In this embodiment, seasonal term prediction is performed by analyzing the seasonal terms obtained from STL decomposition. For periodicity (e.g., quarterly cycle), a periodic replication method is used to replicate historical periodic patterns to the prediction period. Simultaneously, a Transformer model is used to learn the long-sequence dependencies of the seasonal term (e.g., cross-year seasonal pattern changes) to correct the periodic replication results, yielding the predicted seasonal term values. Residual term prediction, residual term... To reflect random fluctuations, an LSTM model is used to learn the patterns of these fluctuations (such as residuals caused by sudden minor events) and predict future residual terms. Results are fused and evaluated by adding the predictions of the trend, seasonal, and residual terms to obtain a market trend forecast. The MAPE (Marginal Performance Equation) of the predicted and actual values is calculated; if the MAPE is less than 10%, the prediction is considered valid. Simultaneously, ensemble learning is used to fuse the results of multiple prediction models for the trend, seasonal, and residual terms (e.g., fusing models initialized with different LSTMs) to reduce the error of a single model and improve prediction robustness. In this embodiment, the cyclical replication of the seasonal term combined with Transformer prediction accurately captures cyclical market changes (such as holiday sales peaks); the LSTM prediction of the residual term covers the impact of random fluctuations; and the fusion and ensemble learning of results ensure prediction accuracy and robustness, providing enterprises with reliable market trend judgments and assisting in decision-making.
[0089] In one implementation, the above-mentioned method of uncovering causal relationships in market changes through pre-defined causal inference includes: Propensity score matching (PSM) combined with a causal graph model was used to eliminate the influence of confounding variables, and the propensity score for each sample was calculated: Where Act is the processing variable (such as whether a promotional activity is implemented), and X is the set of confounding variables (such as holidays, competitor prices, and also includes multimodal confounding factors such as visual effect clustering of promotional posters in the image modality and emotional features of promotional slogans in the audio modality). p(x) is estimated by combining logistic regression with a neural network model. Nearest neighbor matching (caliper 0.05) was used to match three control group samples to each treatment group sample. Covariate balance tests were used to ensure matching effectiveness. The mean treatment effect (ATE) was calculated. Where Y(1) represents the results of the treatment group and Y(0) represents the results of the control group; The robustness of causal relationships is verified by combining the instrumental variable method and the difference-in-differences method. A causal discovery algorithm is used to automatically select variables with ATE>0 and p-value<0.05 as the core driving factors of market changes. A causal relationship network diagram is constructed and optimized by graph neural network. The p-value is used to test the significance of causal relationships and is a significance test index for model parameters. It is used to determine whether the association between the confounding variable and the treatment variable (Act) "truly exists" rather than random noise.
[0090] In this embodiment, the variable definition is the processing variable Act, which indicates whether a promotional activity was implemented; Act=1 indicates implementation, and Act=0 indicates no implementation. The result variable Y is the core indicator of the promotional activity's impact, here "weekly sales," measuring the promotion's effect on market performance. The set of confounding variables X includes traditional confounding factors: holidays (e.g., whether it's a weekend or a public holiday) and competitor prices (average price of summer clothing from surrounding competitors). Multimodal confounding factors include: image modality, and clustering of promotional poster visual effects (such as the clustering method in the above embodiment, clustering posters according to color matching and layout style, such as "vibrant bright color group" and "minimalist style group"), as confounding factors. Audio modality and emotional characteristics of the promotional slogan (using the sentiment analysis model in the above embodiments to extract the emotional tendency of the promotional audio, such as "positive sentiment group" and "neutral sentiment group"), as a confounding factor. The system collects the company's promotional activity records (Act) and weekly sales data (Y) for the past 12 months; it also collects corresponding holiday calendars (marking holidays) and competitor price monitoring data (weekly surveys of prices from three nearby competitors); it extracts promotional poster images (using in-store cameras and downloading from online platforms), and obtains visual effect clustering labels through clustering processing; it collects promotional audio (such as in-store broadcasts and online promotional audio), obtains sentiment feature labels through sentiment analysis, and constructs a set of confounding variables. The model uses logistic regression combined with a neural network to estimate p(x). Logistic regression handles linear relationships (such as the impact of holidays and competitor prices on Act), while neural networks (such as simple MLPs) learn the nonlinear relationships of multimodal confounding factors (such as the complex influence of poster visual effects and advertising sentiment on promotional decisions). The model input is the set of confounding variables X, and the output is the propensity score. denoted by , represents the probability of implementing a promotional activity given a confounding variable. In this embodiment, historical data (X and its corresponding Act) are divided into a training set and a test set in a 7:3 ratio to train a logistic regression-neural network hybrid model. After training, the current confounding variable X to be analyzed is input, and the propensity score p(x) for each sample is output.
[0091] The nearest neighbor matching method was used, with the caliper set to 0.05 (to control for differences in propensity score among the matched samples). For each treatment group sample (Act=1), three control group samples (Act=0) were matched. For example, if the p(x) of the treatment group sample A was 0.7, samples with p(x) between 0.65 and 0.75 were selected from the control group and matched with the three closest samples B, C, and D to ensure a balance of confounding variables.
[0092] To test the difference in the mean values of confounding variables between the matched treatment group and the control group, such as calculating the mean values of holidays, competitor prices, and multimodal factors, if the p-value of the difference is >0.05 (no statistical difference), it indicates that the matching effect is good and the confounding variables have been balanced.
[0093] Calculate the average treatment effect Where Y(1) is the weekly sales of the treatment group (implementing promotions) and Y(0) is the weekly sales of the control group (not implementing promotions). For example, if the mean of Y(1) for the 10 samples in the treatment group is 200,000 yuan and the mean of Y(0) for the 30 matched control group samples is 150,000 yuan, then ATE = 200,000 - 150,000 = 50,000 yuan, indicating that implementing promotional activities can increase weekly sales by an average of 50,000 yuan.
[0094] The causal relationship was verified using a combination of instrumental variable and difference-in-differences methods. The instrumental variable method used "approval time for promotional activities" as an instrumental variable (correlated with Act, but not with residuals) to test the significance of ATE. If the ATE regressed by the instrumental variable was consistent with the original ATE, the causal relationship was robust. The difference-in-differences method used stores that did not implement promotions as the control group and stores that implemented promotions as the treatment group. The sales differences before and after the promotion (e.g., 2 weeks before the promotion and 2 weeks during the promotion) were compared. If the difference was significant and consistent with the direction of ATE, the causal relationship was verified.
[0095] Using a causal discovery algorithm, variables with ATE > 0 and p-value < 0.05 are automatically filtered. For example, "Promotional Activities (Act)" and "Poster Vibrant Color Group (...)" are filtered out. ")" is the core driving factor. =50,000 =30,000), construct a causal relationship network graph, where nodes are variables and edges represent causal relationships, and use a graph neural network to optimize the network structure (e.g., strengthen "Act-Y"). (the edge weight of "-Y").
[0096] Through the above steps, the company verified the causal relationship between "implementing summer apparel promotions" and "sales growth" (ATE=50,000, MAPE<10%), clarified the positive impact of poster visual effects (vibrant bright colors group) on sales (ATE=30,000), and identified that the impact of confounding factors such as holidays and competitor prices has been effectively controlled (covariate equilibrium after matching). The causal network diagram clearly presents the core paths of "promotional activities → sales growth" and "poster visual effects → sales growth." For value decision optimization, based on the causal relationship, the company can expand the use of "vibrant bright colors" style posters, optimize promotional strategies, and simultaneously pay attention to the dynamics of holidays and competitor prices to adjust the promotional pace. In terms of resource allocation, more resources should be invested in promotional execution and poster design to strengthen the positive impact of core driving factors. Regarding risk control, after identifying confounding factors, the company can proactively address peak holiday traffic and the impact of low competitor prices, reducing market uncertainty.
[0097] In one embodiment, the aforementioned virtual decision-making scenario, which is consistent with the real environment and constructed through a twin network based on historical market data, enterprise operational status profiles, and the mined causal relationships of market changes, simulates virtual market feedback and virtual enterprise performance corresponding to different decision-making behaviors, including: S71. The twin network includes a real-scene branch and a virtual-scene branch with shared weights. The inputs are historical market data and enterprise operation status profile features. The network adopts the Transformer architecture to enhance feature capture capabilities.
[0098] In this step, the aforementioned Siamese network consists of two branches (a real-world branch and a virtual-world branch) with identical structures and shared weights. It can compare the features of real and virtual scenes to verify the consistency between the virtual and real environments, and is commonly used in simulation and modeling tasks. The Transformer architecture, a deep learning architecture based on a self-attention mechanism, excels at capturing global correlation features of long sequences and multimodal data, enhancing its ability to model complex relationships in market data.
[0099] The design employs a dual-branch structure for the twin network, with the real-world and virtual-world branches sharing initial weights. Historical market data (such as past sales and competitive data) and enterprise operational status profile features (generated in claim 6) are organized by time series and feature dimensions and then input into the two branches. Utilizing the multi-head self-attention mechanism of the Transformer architecture, long-range dependency features of the data (such as cross-quarter market trend correlations and multi-department data collaboration relationships) are extracted, laying the feature foundation for subsequent virtual-world simulations and comparisons with the real environment. By leveraging the twin network's dual-branch structure and the Transformer architecture, the feature input and processing logic for both real and virtual scenarios are unified, ensuring feature consistency in subsequent virtual-world simulations; enhancing feature capture capabilities allows even subtle correlations in market data (such as seasonal fluctuations in niche markets) to be uncovered, improving the accuracy of virtual decision-making scenarios in replicating the real market.
[0100] S72. The virtual scenario branch simulates the impact of decisions through a generation module guided by causal relationships. The generation module introduces a causal attention mechanism. For price adjustment decisions, the formula for calculating virtual sales changes is:
[0101] in, For virtual sales, Based on benchmark sales, η is the price change rate, and η is the price elasticity coefficient (obtained by regression of historical data combined with deep neural networks, combined with the competitive product price display effect of image mode and the emotional feedback correction of price broadcast in audio mode). For the first k related factors (such as changes in competitor prices). Factor weights (set based on causal inference results); In this step, the causal attention mechanism, within the generation module, focuses on the core associations uncovered through causal inference (such as the causal relationship between promotional activities and sales volume), guiding the generation module to prioritize simulating the impact of key factors on decision-making, thus improving the relevance of the simulation. The price elasticity coefficient η is an indicator that measures the degree to which price changes affect sales volume; η > 0 indicates a negative correlation between price and sales volume (price increases lead to decreased sales volume), and the larger the absolute value of η, the stronger the elasticity. By learning non-linear relationships through historical data regression (such as linear regression) combined with deep neural networks (such as LSTM), and further incorporating modal data corrections such as image (competitor price display effects, such as poster comparisons) and audio (price broadcast sentiment, such as user voice feedback on prices), the coefficients are made more closely aligned with the actual market. Related Factors With weight , Other market factors that influence decision-making effectiveness (such as changes in competitor prices and policy adjustments). The factor weights are set based on the results of causal inference. The stronger the correlation, the greater the weight, reflecting the importance of the factor to the decision.
[0102] For price adjustment decisions, a causal attention mechanism is introduced into the virtual scenario branch generation module to activate causal relationship features strongly associated with price decisions (such as the price-sales causal path). Substituting these features into the virtual sales change calculation formula, a baseline sales volume is first determined. (Sales volume during the same period in history without price adjustments), price change rate (For example, if the planned price increase is 10%, then it's 0.1). The price elasticity coefficient η is obtained through historical data regression, deep neural networks, and multimodal correction (e.g., considering the visual impact of competitor pricing posters, η is corrected to -2.5, meaning a 1% price increase leads to a 2.5% decrease in sales). List the related factors. (If a competitor lowers its price by 5% during the same period, it is recorded as) =-0.05), factor weights are set based on causal inference results. (e.g., weighting of competitor price changes) =0.3). Substitute these parameters into the formula to calculate the virtual sales volume. This simulation model can simulate changes in sales volume after price adjustments. For other decision types (such as channel adjustments and product iterations), corresponding formulas can be constructed (e.g., formulas for changes in market coverage after channel adjustments) and simulated similarly. This embodiment can accurately simulate the impact of decisions on market feedback (e.g., sales volume) and corporate performance (e.g., revenue). The causal attention mechanism focuses on core correlations, avoiding interference from irrelevant factors. Multimodal corrections to the price elasticity coefficient and the weights of related factors make the simulation results more closely resemble the complex interactions of the real market (e.g., users are more sensitive to price due to poster visual differences), providing quantitative and reliable feedback data for virtual decision verification.
[0103] S73. The similarity between the output features of the two branches of the Siamese network (cosine similarity > 0.9) is compared to verify the fit between the virtual scene and the real environment. Adversarial training is used to continuously optimize the parameters of the generation module until the fit requirement is met. At the same time, domain adaptation technology is introduced to improve the generalization ability of the virtual scene.
[0104] In this step, the fit between the virtual scene and the real environment is verified and optimized. Cosine similarity is an indicator that measures the cosine of the angle between two feature vectors. It is used to judge the similarity of the output features of the two branches of the Siamese network. The closer the value is to 1, the more similar the features are, and the better the virtual scene fits the real environment. Adversarial training involves the generation module (virtual scene) and the discriminator module (implicit in the Siamese network comparison) playing against each other. The generation module optimizes parameters to make the virtual features closer to reality, while the discriminator module distinguishes between real and virtual features, improving the realism of the virtual scene. Domain adaptation technology allows the virtual scene model to be adapted to different market domains (such as expanding from first-tier cities to third-tier cities), learning the common and different features of different domains, and enhancing generalization ability.
[0105] The cosine similarity is calculated by comparing the output features of the real-world branch (outputting real market features) and the virtual-world branch (outputting features after virtual decisions) of the twin network. If the similarity is <0.9, adversarial training is initiated, and the generation module adjusts its parameters (such as modifying the multimodal correction logic of the price elasticity coefficient) to make the virtual features closer to reality. Simultaneously, domain adaptation techniques (such as transfer learning) are introduced, inputting historical data from different regions and market cycles to allow the virtual scenario to learn market patterns from multiple samples, improving its adaptability to new market environments. This process is iterated until the cosine similarity is greater than or equal to 0.9, ensuring a high degree of consistency between the virtual scenario and the real environment. In this embodiment, cosine similarity is used to quantify the consistency between the virtual and real environments, providing a clear optimization objective. Adversarial training dynamically optimizes the generation module, addressing the problem of "assumed" simulations in the virtual scenario (such as overestimating promotional effects). Domain adaptation techniques allow the virtual scenario to break through the limitations of a single market, making it suitable for cross-regional and cross-business line decision verification, enhancing the practical value and generalization ability of the virtual decision-making scenario.
[0106] Taking a retail company's plan to adjust summer clothing prices as an example, a twin network is constructed: A two-branch twin network is designed, inputting historical market data (sales, competitor prices) for summer clothing over the past three years and features reflecting the company's operational status (such as "high-volume sales clusters" and "high-end product display clusters"). A Transformer architecture is used to capture cross-quarter sales trend correlations (such as the sales peak from June to August each year) and multi-department data collaboration (such as the correlation between marketing department posters and sales department performance), initializing network weights. A price adjustment decision is simulated: a baseline sales volume is used. : Calculate the average weekly sales volume of 1000 units during the same period last year before price adjustments. Price Change Rate The planned price increase is 8%, or 0.08. The price elasticity coefficient η: Historical data regression yielded η = -2.2. After adjustments based on the image modality (competitors' posters have monotonous colors, while our company's posters have better visual appeal) and the audio modality (few user complaints about price in inquiries), η is adjusted to -2.0. Related factors. Competitors plan to reduce prices by 5% during the same period. =-0.05; Factor weight Based on causal inference, the weight of the impact of competitor price changes on sales volume. =0.35. Substitute into the formula: =1000×(1+(-2.0)×0.08+0.35×(-0.05))=1000×(1-0.16-0.0175)=822.5 units, simulating sales of approximately 823 units after the price increase. The twin network's real-scene branch outputs the real feature vector of the current market (e.g., including features such as "high demand in summer, low price impact from competitors"), while the virtual-scene branch outputs the virtual feature vector after the decision (e.g., including features such as "expected sales decrease after price increase, advantage of our company's posters"). The calculated cosine similarity is 0.85<0.9, so adversarial training is initiated: the generation module adjusts parameters (e.g., strengthening the positive correction of sales by poster advantages), and after resimulation, the cosine similarity increases to 0.92, and the virtual scene matches the real environment; simultaneously, historical data from second-tier cities is input using domain adaptation technology to optimize the model's adaptation to different markets, ensuring the generalization of the virtual scene.
[0107] In one embodiment, the aforementioned reinforcement learning environment constructs a virtual environment where corporate decision-making behavior in a virtual scenario is taken as actions, and virtual market feedback and virtual corporate performance output by a twin network are taken as reward signals. Reinforcement learning iterations are completed in the virtual environment to generate the optimal decision scheme verified virtually. The state space of the reinforcement learning environment is a fusion vector of corporate operational status profile features and market trend prediction results, and the state representation is optimized through a feature selection algorithm. The action space includes discrete decision options such as price adjustment, advertising, and inventory optimization, and hierarchical reinforcement learning is used to decompose the actions.
[0108] A reinforcement learning environment is a framework that allows an intelligent agent (corporate decision-making logic) to interact with the environment (virtual market scenario), learning the optimal strategy by executing actions (decision-making) and receiving rewards (market feedback). It comprises three key elements: state space, action space, and reward mechanism. The state space is the set of environmental information perceived by the agent; it is a fusion vector of the enterprise's operational status profile and market trend predictions, reflecting the enterprise's current market situation. The action space is the set of decision options the agent can execute, such as discrete decisions like price adjustments and advertising. Hierarchical reinforcement learning decomposes complex actions (such as "omnichannel marketing") into sub-actions ("online advertising" and "offline store promotion"), reducing decision complexity.
[0109] Multimodal features (such as image clustering features and structured data trend features) of the enterprise's operational status profile are extracted and fused with market trend prediction results (such as sales growth prediction and competitive landscape prediction) to form a high-dimensional state vector. Feature selection algorithms (such as mutual information-based selection) are used to filter features that significantly impact decision-making (such as removing redundant historical data features) to optimize the state space representation. The types of executable decisions for the enterprise (pricing, advertising, inventory, etc.) are identified and broken down into discrete action options. A hierarchical reinforcement learning framework is adopted to decompose complex decisions into a sequence of sub-actions: "strategy layer (such as whether to adjust prices) - execution layer (such as the magnitude of price adjustment)," constructing an action space. This step precisely defines the interaction basis of reinforcement learning. The state space integrates core market perception information, allowing the agent to accurately "understand" the enterprise's situation. The hierarchical decomposition of the action space reduces decision complexity and adapts to the enterprise's multi-dimensional and refined decision-making needs. Feature selection optimizes the state representation, improving learning efficiency and avoiding the "curse of dimensionality."
[0110] The above-mentioned reward signals employ a weighted fusion mechanism combined with a dynamic reward strategy, and the calculation formula is as follows: in, For virtual market feedback scores (such as standardized values of sales growth rate). Assigning performance scores to virtual enterprises For dynamically rewarding shaping items, and These are the weighting coefficients for virtual market feedback and virtual enterprise performance, respectively, and γ is the weighting coefficient for the reward shaping item.
[0111] In this step, the reward signal is the feedback received by the agent after performing an action in reinforcement learning, guiding the agent to learn the "optimal decision" (prioritizing actions with high rewards). This includes market feedback, performance feedback, and dynamic shaping. The virtual market feedback score is also included. In virtual scenarios, the market response to decisions (such as standardized sales growth rates) is quantified, reflecting the market's direct feedback to those decisions. Virtual enterprise performance score. The score, which combines indicators such as corporate profit, cost, and market share, reflects the impact of decisions on the company's long-term performance. Dynamic reward shaping item. This serves as a supplementary reward signal, guiding the agent to explore potential optimal strategies (such as encouraging the trying of new market strategies) and compensating for the shortcomings of sparse rewards (such as long-term performance feedback lag). Weighting coefficients , γ, to balance the influence of different reward items. , It focuses on the trade-off between the current market and corporate performance, and γ adjusts the guiding strength of dynamic shaping.
[0112] Quantitative calculation of each reward item, Collect data on changes in sales volume and market share after a decision is made in a virtual scenario, and standardize the data into a score of 0-1 (e.g., a 20% increase in sales corresponds to a score of 0.8). The scores are calculated by combining the profits (such as revenue minus costs after price adjustment) and market share changes of virtual enterprises using the analytic hierarchy process (such as profit weighting 0.6 and market share weighting 0.4). Design dynamic rules (such as an extra reward of 0.2 for the first attempt at a new advertising format) to encourage agents to explore innovative decisions; or set trigger conditions and reward values for shaping items based on expert experience (such as "quick inventory turnover reward"). (Follow the formula) Weighted fusion, through reinforcement learning training (such as the PPO algorithm), dynamically adjusts the weight coefficients (initially). High emphasis on market feedback, later (High performance focus) and optimized reward guidance logic. Multi-dimensional reward integration balances short-term market feedback with long-term corporate performance, avoiding short-sighted decision-making (such as blindly pursuing sales at the expense of profits); dynamic shaping items compensate for reward sparsity, encourage agents to explore potential optimal strategies (such as developing new market channels), improve the adaptability of reinforcement learning to complex market decisions, and make the learned strategies more comprehensive and resilient.
[0113] The agent continuously samples in the state space, performing actions, acquiring rewards, and updating its policy in a cyclical process. It optimizes its decision-making strategy using algorithms such as policy gradient, gradually approaching the optimal solution. The optimal decision scheme is the decision sequence that, after reinforcement learning convergence, yields the highest average reward, optimal market feedback, and optimal firm performance in the virtual environment (e.g., "small price adjustment + precise advertising + dynamic inventory optimization"). The agent's policy is initialized (e.g., randomly selecting actions) and executed cyclically in the virtual scenario: the agent observes the current state (the fused market state vector) and selects actions from the action space (e.g., "price increase by 5% + online advertising"). A Siamese network simulates the impact of the decision, outputting virtual market feedback and performance, and calculating the reward signal R. The agent updates its policy parameters using reinforcement learning algorithms (e.g., Actor-Critic), increasing the probability of selecting high-reward actions. Iteration continues until policy convergence (e.g., reward fluctuation <1% for 100 consecutive iterations), extracting the converged optimal decision sequence, and generating the optimal decision scheme verified in the virtual environment. Through continuous iteration, decision-making strategies can be fully "tried and optimized" in a virtual environment to discover the optimal solution for multiple decision combinations (such as synergistic strategies for pricing, advertising, and inventory). Virtual verification ensures that decision-making solutions are superior in both market feedback and corporate performance, reduces decision-making risks in the real market, and improves the quality of corporate decisions in response to complex market changes.
[0114] This embodiment upgrades enterprise decision-making from "experience-based trial and error" to "virtual verification-data optimization" through reinforcement learning construction, reward mechanism design, and iterative optimization. The "low-cost trial and error" in the virtual environment allows enterprises to explore optimal solutions for multiple decision combinations, avoiding real market risks; multi-dimensional reward integration balances short-term market feedback and long-term enterprise performance, ensuring decision sustainability; and in synergy with the overall methodology, it forms a complete closed loop of "market perception-virtual simulation-reinforcement learning-decision optimization," significantly improving the scientific rigor, accuracy, and competitiveness of enterprise market decisions, helping enterprises respond quickly and strategically in complex markets, and achieve sustainable growth.
[0115] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference, characterized in that, Comprise the following steps: A unified data collection framework is constructed to fuse and collect multi-modal data including traditional structured data and images, audio, and video; A data quality evaluation model is constructed in advance through a machine learning algorithm, and the multi-modal data is processed by combining a data cleaning algorithm based on deep learning to obtain reliable multi-modal data; Different reliable multi-modal data is analyzed at multiple levels to mine data attribute similarity and semantic association and potential relationship between different modal data, and a cross-department associated data network containing clustering results and data association relationships is constructed; Based on the clustering results and data association relationships in the cross-department associated data network, different types of data obtained by clustering are mapped into a spatio-temporal graph structure, and a situational awareness model based on a spatio-temporal graph convolution network is constructed through graph convolution operation and time convolution operation to capture market dynamic changes and generate an enterprise operation state portrait; wherein the data nodes in the spatio-temporal graph structure correspond to each type of data after clustering, and the edges correspond to the association relationships between the data; The market trend prediction result is obtained by analyzing the enterprise operation state portrait; If the prediction result triggers the adjustment of market decision, the market change causal relationship is mined through a preset causal inference method; Based on historical market data, enterprise operation state portraits, and mined market change causal relationships, a virtual decision-making scenario consistent with the real environment is constructed through a twin network to simulate the virtual market feedback and virtual enterprise performance corresponding to different decision-making behaviors; An reinforcement learning environment is constructed, the enterprise decision-making behavior in the virtual scenario is taken as an action, the virtual market feedback and virtual enterprise performance output by the twin network are taken as reward signals, and reinforcement learning iteration is completed in the virtual environment to generate an optimal decision-making scheme verified by virtual verification.
2. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, After the optimal decision-making scheme verified by virtual verification is generated, the following steps are included: The situational awareness model is continuously optimized based on a preset online learning module and a meta-learning module, wherein the online learning module updates the situational awareness model in an incremental manner in real time by capturing new data features and market change information, the meta-learning module automatically adjusts learning algorithm hyperparameters and model structure by analyzing historical learning processes and the performance of the situational awareness model, and the decision-making scheme is dynamically adjusted according to model optimization results and market real-time feedback.
3. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, The construction of the unified data collection framework to fuse and collect multi-modal data including traditional structured data and images, audio, and video comprises: A distributed data collection hub is built to connect enterprise resource planning systems, customer relationship management systems, and transaction databases through a structured query language interface to collect traditional structured data, and edge collection nodes are deployed to collect image data and audio data; Wherein, based on a preset multi-modal data weight distribution model, the collection priority is dynamically adjusted according to the contribution of the data to market perception, and the contribution calculation formula is: wherein, is the weight of the mth modality data, is the integrity coefficient of the modality data at time t, is the market correlation degree of the modality data at time t, M is the total number of modalities, T is the acquisition time window, and the scheduling module allocates network bandwidth and computing resources according to the weight of the modality data.
4. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, The construction of the data quality evaluation model through a machine learning algorithm, and the processing of the multi-modal data by combining a data cleaning algorithm based on deep learning to obtain reliable multi-modal data comprises: A data quality evaluation model is constructed from four dimensions of accuracy, integrity, consistency and timeliness to calculate the quality score of multi-modal data; When the quality score is lower than the quality threshold, the multi-modal data is determined as low-quality data, and a generative adversarial network is used to repair the low-quality data to obtain credible multi-modal data.
5. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, The different credible multi-modal data are clustered at multiple levels to mine data attribute similarity and semantic association and potential relationship between different modal data, and a cross-department associated data network is constructed, including: Density clustering is used for the same modal data, wherein the neighborhood radius is dynamically adjusted according to the data standard deviation; The semantic similarity of different modal clustering results is calculated through a cross-modal attention mechanism, and when the semantic similarity is greater than a similarity threshold, a weighted edge between two clusters is established in the cross-department associated data network to realize semantic association modeling of multi-modal data, and the cross-department associated data network is obtained.
6. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, Based on the clustering results and data association relationships in the cross-department associated data network, different class data obtained by clustering are mapped into a spatio-temporal graph structure, and a situation awareness model based on a spatio-temporal graph convolution network is constructed through graph convolution operation and time convolution operation to capture market dynamic changes and generate an enterprise operation state portrait, including: In the spatio-temporal graph structure, the node feature vector fuses the statistical characteristics of the clustering cluster and the timestamp information, and the edge weight is dynamically updated based on the association strength in the cross-department associated data network; The graph convolution operation adopts an improved Chebyshev polynomial approximation, and the calculation formula is: wherein, is a normalized Laplacian matrix, is a k-th order Chebyshev polynomial, is a learnable parameter, and K is a polynomial order; The time convolution operation adopts a 1D convolution layer, and the convolution kernel size is dynamically adjusted according to the time granularity. The enterprise operation state portrait containing 128-dimensional features is output through a fully connected layer.
7. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, The enterprise operation state portrait is analyzed to obtain market trend prediction results, including: From the enterprise operation state portrait, the core index time series is extracted, and the STL algorithm is used to decompose into trend item , seasonal item and residual item , that is ; For the user behavior clustering features mined from the image modal data, an attention mechanism is used to map them into a passenger flow trend auxiliary index, and the user feedback sentiment features of the audio modal are mapped into a sales conversion rate influencing factor through a sentiment analysis model and integrated into the time series decomposition logic; For the trend term, a quadratic exponential smoothing prediction is used, combined with a long short-term memory network to capture nonlinear trends, and the formula is: wherein, is the trend prediction value at time t+1, is the trend change rate, and a and β are smoothing coefficients, is the prediction supplement of the LSTM model for the recent sequence; For the seasonal term, a period replication method is used in combination with a Transformer model for prediction, and the residual term is predicted through a long short-term memory network. Finally, the market trend prediction value is obtained by fusing the three parts of the results, and the prediction error is evaluated by the mean absolute percentage error. When the mean absolute percentage error is less than a preset error threshold, it is determined as an effective prediction. At the same time, the prediction results of multiple models are integrated by using ensemble learning to improve the prediction robustness.
8. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, The market change causal relationship is mined through a preset causal inference method, including: A propensity score matching is used in combination with a causal graph model to eliminate the influence of confounding variables, and the propensity score of each sample is calculated: where Act is the treatment variable, X is the set of confounding variables, and p(x) is estimated by a logistic regression combined with a neural network model; A nearest neighbor matching method is used to match 3 control group samples for each treatment group sample, and a covariate balance test is used to ensure the matching effect, and the average treatment effect is calculated: where Y(l) is the treatment group result and Y(0) is the control group result. The robustness of the causal relationship is verified by combining the instrumental variable method and the double difference method. The causal discovery algorithm is used to automatically screen out variables with ATE>0 and p-value<0.05 as the core driving factors of market changes, and a causal relationship network diagram is constructed.
9. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, Based on historical market data, enterprise operation state portraits, and the causal relationship of market changes discovered, a virtual decision-making scenario that matches the real environment is constructed through a twin network, simulating the virtual market feedback and virtual enterprise performance corresponding to different decision-making behaviors, including: The twin network includes real scenario branches and virtual scenario branches that share weights, and the input is historical market data and enterprise operation state portrait features. The twin network adopts a Transformer architecture; The virtual scenario branch simulates the impact of decision-making through a causal relationship-guided generation module. For price adjustment-type decisions, the virtual sales change calculation formula is: wherein, is the virtual sales volume, is the benchmark sales volume, is the price change rate, and η is the price elasticity coefficient, is the kth correlation factor, is the factor weight; The similarity of the output features of the two branches of the twin network is compared to verify the degree of matching between the virtual scenario and the real environment. The generation module parameters are continuously optimized using adversarial training until the matching requirement is met. Meanwhile, the field adaptation technology is introduced to improve the generalization ability of the virtual scenario.
10. The market trend perception and decision-making method based on spatio-temporal graph convolution and causal inference according to claim 1, characterized in that, The state space of the reinforcement learning environment is the fusion vector of enterprise operation state portrait features and market trend prediction results, and the state representation is optimized through a feature selection algorithm. The action space of the reinforcement learning learning environment includes discrete decision options such as price adjustment, advertising, and inventory optimization, and a hierarchical reinforcement learning is used to decompose the actions; The reward signal adopts a weighted fusion mechanism combined with a dynamic reward strategy, and the calculation formula is: wherein, is the virtual market feedback score, is the virtual firm performance score, is the dynamic reward term, are the weight coefficients for the virtual market feedback and the virtual firm performance, respectively, and γ is the weight coefficient for the reward shaping term.