Multi-category sales volume joint prediction system and method based on large model environment semantic perception

Through the multi-category sales joint prediction system with semantic perception of large-scale environmental environment, the generalization ability and accuracy of multi-category sales prediction methods in a dynamic environment is solved, and the time-space-dependent learning of multi-category sales sequences is realized, which improves the prediction effect and the operational efficiency of the supply chain.

CN120471655APending Publication Date: 2025-08-12SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510575059.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

When faced with dynamic environmental changes, the existing multi-category sales forecasting methods have insufficient generalization capabilities and prediction accuracy, making it difficult to effectively capture the multi-dimensional space-time dependence of multi-category sales sequences.

Method used

A multi-category sales joint prediction system based on large-scale environmental semantic perception is adopted, including a data collection module, an environmental denoising guide mask text generation module, a LoRA end-to-end multi-category environmental semantic extraction enhancement module and a multi-view characterization learning module. Through unified modeling and feature screening of multi-source data, dynamic environmental information is captured to realize the space-time and space-dependent learning of multi-category sales sequences.

Benefits of technology

It improves the accuracy and robustness of sales forecasts in multiple categories, can effectively predict during normal and sudden abnormal periods, supports the operational optimization of the retail supply chain, and improves sales planning and inventory management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471655A_ABST
    Figure CN120471655A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-category sales volume joint prediction system and method based on large model environment semantic perception. Comprising a data collection module, a text generation module based on an environment denoising guide mask, an end-to-end multi-category environment semantic extraction enhancement module based on LoRA and a multi-category sales combined prediction module based on multi-view representation learning. According to the method, screening, learning and fusion of dynamic environment information for multi-category sales volume prediction are effectively realized, the difficulty of space-time dependence learning of a multi-category sales volume sequence is solved, and improvement of the accuracy of multi-category sales volume prediction is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer information computing, and in particular to multi-category sales forecasting in the retail industry, and mainly relates to a multi-category sales joint forecasting system and method based on large model environment semantic perception. Background Art

[0002] With the development of the retail industry and its infrastructure, people's shopping needs have been greatly met. Driven by the trends of smart retail and smart logistics, many large retail companies have begun to adopt intelligent technologies to improve operational efficiency and operating profits. Accurate sales forecasts for multiple categories play a crucial role in enabling and increasing the efficiency of the entire supply chain platform. Sales forecasts influence many aspects of the supply chain, such as sales planning, inventory management, and the optimization of distribution resources.

[0003] Existing sales forecasting methods can be divided into two categories: those based on heterogeneous numerical encoding and those based on text semantic embedding. Methods based on traditional heterogeneous numerical encoders collect multi-source external data, organize it in a structured table, and develop specialized encoders for each feature type. However, this data modeling approach requires sufficient training data to learn the embedding of external features. It cannot effectively model new environmental factors or the influence of sparse features, resulting in weak robustness and generalization capabilities. While the prior world knowledge learned through pre-training in methods based on text semantic embeddings can help alleviate the modeling problem of dynamic environmental features, its single autoregressive time series prediction method makes it difficult to capture the multidimensional spatiotemporal dependencies of multi-category sales series, limiting its prediction effectiveness. Summary of the Invention

[0004] The present invention is aimed at the problems of insufficient generalization ability and prediction accuracy of multi-category sales prediction methods in the existing technology, and proposes a multi-category sales joint prediction system and method based on large-model environmental semantic perception, including a data collection module, a text generation module based on environmental denoising guided mask, an end-to-end multi-category environmental semantic extraction enhancement module based on LoRA, and a multi-category sales joint prediction module based on multi-perspective representation learning. The data collection module is used to obtain multi-category sales time series data and multi-source external data; the text generation module based on environmental denoising guided mask is used to construct the collected data into a masked text description enhanced with sales prediction domain knowledge; the end-to-end multi-category environmental semantic extraction enhancement module based on LoRA performs pre-trained language model domain fine-tuning through multiple channel-independent single-category sales prediction tasks; the multi-category sales joint prediction module based on multi-perspective representation learning includes an environment-sales causal evolution module for capturing time series evolution patterns, a channel correlation mining module for capturing multi-channel correlations, a multivariate spatiotemporal dependency mining module for capturing the spatiotemporal dependencies of multivariate sales series, a multi-perspective representation fusion module, and a multi-category sales prediction module. The present invention effectively solves the problems of screening, learning and integrating dynamic environmental information for multi-category sales forecasting, solves the difficulty of learning spatiotemporal dependencies of multi-category sales series, and helps to improve the accuracy of multi-category sales forecasting.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is: a multi-category sales joint prediction system based on large-scale model environmental semantic perception, which at least includes a data collection module, a text generation module based on environmental denoising guidance mask, an end-to-end multi-category environmental semantic extraction enhancement module based on LoRA, and a multi-category sales joint prediction module based on multi-view representation learning.

[0006] The data collection module is used to obtain multi-category sales time series data and multi-source external data; the multi-source external data includes numerical, categorical, and textual data, which are used to assist in sales forecasting;

[0007] The text generation module based on the environment denoising guided mask is used to construct the collected data into a masked text description enhanced with sales forecast domain knowledge; using the pre-trained language model to extract the masked token embedding of each text description to obtain the denoised static environment semantic embedding, the predicted environment semantic embedding, and the observed environment semantic embedding;

[0008] The LoRA-based end-to-end multi-category environmental semantic extraction enhancement module: fine-tunes the pre-trained language model domain through multiple channel-independent single-category sales prediction tasks; the pre-trained language model uses the LoRA algorithm in its self-attention layer, adds historical sales numerical data based on the environmental semantic embedding as input, obtains a fused embedding through a weighted network, and learns the interactive relationship between the dynamic environment and specific category sales through a gated residual network, an autoregressive LSTM model, and a causal masked attention mechanism, achieving domain adaptation of the pre-trained language model through an end-to-end training method;

[0009] The multi-category sales joint prediction module based on multi-perspective representation learning includes an environment-sales causal evolution module, a channel association mining module, a multi-variable spatiotemporal dependency mining module, a multi-perspective representation fusion module, and a multi-category sales prediction module.

[0010] The environment-sales causal evolution module is used to capture the temporal evolution pattern; the channel association mining module is used to capture multi-channel associations; the multi-variable spatiotemporal dependency mining module is used to capture the spatiotemporal dependency of multi-variable sales series; the multi-perspective representation fusion module and the multi-category sales prediction module are used for multi-perspective feature fusion and multi-category sales prediction respectively;

[0011] In the multi-category sales joint prediction module based on multi-perspective representation learning, all environmental semantic embeddings and historical sales numerical data are fused through a weighted network based on the static environment semantic embedding, predicted environment semantic embedding and observed environment semantic embedding enhanced by domain fine-tuning. The temporal evolution representation, channel association representation and multivariate spatiotemporal dependency representation of each category are obtained through the multi-perspective representation learning method. The feature fusion and splicing are input into the gated residual network, and the multi-category sales in the future time step to be predicted are decoded through autoregressive prediction.

[0012] As an improvement of the present invention, the data collected by the data collection module includes at least the names of multiple categories corresponding to the sales sequence to be predicted, the sales regions, the names of retailers, daily sales data aligned in the time dimension, date data, holiday data, promotional activity data, weather data and news hot search data.

[0013] As an improvement of the present invention, the template in the text generation module based on the environment denoising guidance mask includes three parts: domain knowledge, environment description, and task configuration. Domain knowledge is a priori world knowledge summarized in the target prediction scenario, which is used to enhance the language model's prediction of the impact of the current environment on the sales of the target category. The format is "Product sales are time-periodic and affected by external correlation factors. Pay attention to the following predicted / observed / static environment information for <category> sales prediction:". This part of the domain knowledge can be summarized by scenario business personnel or provided by a larger language model. The environment description is a unified text description converted from prepared multi-source heterogeneous data according to predefined data-to-text generation rules, and the descriptions of different source data are separated by line breaks. The task configuration is a task instruction provided to the language model, used to align the model's understanding of the sales forecasting task and stimulate the model's reasoning ability about environmental impacts. The format is "Please infer the impact of this information on the sales of <category> is: [MASK]." Here, the [MASK] token is used in the form of an instruction-guided mask to interact with environmental information, and refined environmental information is extracted in the form of a latent vector in the high-dimensional embedding space. The final environmental impact is represented by the embedding of the [MASK] token in the last layer of the language model. The text generation module based on the environmental denoising-guided mask ultimately constructs multi-source heterogeneous data into a static context mask text description for the sales forecast of the target category, as well as a dynamic context mask text description of the time-varying observation features and time-varying prediction features.

[0014] As another improvement of the present invention, the specific structure of the gated residual network GRN used in the end-to-end multi-category environment semantic extraction enhancement module based on LoRA is as follows:

[0015] GRN θ (a,c)=LayerNorm(a+GLU θ (η1))

[0016] η1=W 1,θ η2+b 1,θ

[0017] η2=ELU(W 2,θ a+W 3,θ c+b 2,θ )

[0018] In the above formula, θ is used to identify the gated residual network with different parameters, a is the embedding to be fused, c is the environment semantic embedding, LayerNorm is the standard layer normalization function, η1 and η2 represent the intermediate embeddings of the processing, and GLU stands for the gated linear unit, which is used to control the degree of fusion of the environment semantics, expressed as GLU θ (γ)=σ(W 4,θ γ+b 3,θ)⊙(W 5,θ γ+b 4,θ ), where ⊙ represents the element-wise Hadamard product, σ represents the sigmoid function used to obtain the gate value; W in all the above formulas 1,θ To W 3,θ and b 1,θ to b 2,θ These are the neural network parameters to be learned. a and c are first concatenated and fed into a two-layer MLP perceptron network to fuse them to obtain η2 and η1, and then fed into the GLU module. The ELU stands for the exponential linear unit activation function, and the formula is:

[0019]

[0020] Here β is a hyperparameter greater than zero, e x As a natural exponential function, ELU ensures that gradient propagation occurs when the input x is less than or equal to zero, thus avoiding the problem of neuron death in traditional activation functions.

[0021] As another improvement of the present invention, in the end-to-end multi-category environment semantic extraction enhancement module based on LoRA, obtaining fusion embedding through a weighted network specifically includes the following steps:

[0022] Use the linear layer to embed the sales value and get the vector

[0023] The environment text description constructed by the environment denoising guided mask is embedded through the pre-trained language model to obtain vectors

[0024] All vector embeddings are concatenated, By doing the following:

[0025]

[0026] The resulting vector That is weighted fusion embedding, where v t It is an m-dimensional weight tensor. According to the corresponding weights calculated by the softmax function, it is embedded in the m input vectors for weighted fusion.

[0027] As another improvement of the present invention, in the environment-sales causal evolution module, the semantic embedding of the input environment text description is obtained based on the language model, and it is fused with the historical sales numerical data through a weighted network. The long-term and short-term temporal dependencies are further fused through the autoregressive LSTM model and the causal mask attention mechanism, and finally the temporal evolution representation is obtained through the gate unit and the residual module.

[0028]

[0029] Among them, c represents the category identifier, t represents the time step to be predicted, is the hidden state embedding output by the LSTM decoder, is the fused context semantic embedding provided by the aforementioned contribution weighted network, Represents the intermediate representation at the sequence level, and MaskedAttention represents the classic unidirectional causal mask attention mechanism, that is, the current time step can only aggregate the information of the previous time step, ensuring the aggregation of causal relationships through time order.

[0030] As another improvement of the present invention, in the channel association mining module, the hidden state vector obtained by the LSTM decoder and the cell state vector Based on this, different gated residual networks are used to fuse the semantic embedding of the current time step environment. Get fused embedding and Then, the fusion embedding corresponding to all C categories is input into the two-layer Transformer encoder to obtain the long-term and short-term channel correlation representation and

[0031]

[0032] The fused channel correlation representation is further obtained based on the long-term and short-term correlation gating mechanism:

[0033]

[0034] Among them, FC represents the fully connected layer, σ represents the sigmoid activation function, and ⊙ represents the element-by-element product. It is the adaptive multi-category association representation for the final fusion.

[0035] As a further improvement of the present invention, in the multivariate spatiotemporal dependency mining module, the multivariate sales sequence is formed as a whole into two-dimensional tensor data and input into the inverted Transformer encoder to obtain the two-dimensional spatiotemporal embedding of the time step to be predicted:

[0036]

[0037] Among them, x :,c is the sales volume sequence of category c, which is encoded into variable token embedding through the linear embedding layer Embedding(·) Aggregated into a two-dimensional tensor H in the category dimension ST, and then input into the Transformer encoder, and when outputting the projection, it will pass through the linear layer Projection of the corresponding time step t (·) Align with the current time step t to get the fused embedding Further, different gated residual GRN modules are used to fuse the current time step environment semantic embedding to obtain a multivariate spatiotemporal dependency representation with enhanced environment information.

[0038]

[0039] in, is the original context semantic embedding fused through the contribution weighted network, is the temporal hidden state vector obtained in the aforementioned environment-sales causal evolution module.

[0040] Finally, the temporal evolution representation, channel correlation representation, and multivariate spatiotemporal dependency representation of each category are concatenated together and decoded through a gated residual module and a linear output layer to produce the predicted sales volume for each category:

[0041]

[0042] To achieve the above objectives, the present invention also adopts a technical solution: a multi-category sales joint prediction method based on large-scale model environment semantic perception, comprising the following steps:

[0043] S1, multi-source heterogeneous external data collection: Based on sales forecasting domain knowledge and the characteristics of multiple categories to be predicted, multi-category sales time series values and multi-source heterogeneous external environment data are obtained. Based on whether the data is time-varying, the data is divided into static features, time-varying predictive features, and time-varying observation features.

[0044] S2, Masked Context Text Generation: A text generation module based on environment denoising-guided masking, which converts multi-source heterogeneous external data into static / predicted / observed context masked text descriptions enhanced with domain knowledge;

[0045] S3, domain fine-tuning of the pre-trained language model: This involves extracting contextual semantic embeddings from the pre-trained language model and combining them with multi-category sales numerical data as input. The multi-category sales numerical data corresponding to the time step to be predicted is used as the output label. This splits the multi-category sales prediction task into multiple single-category sales prediction tasks. This involves training and updating the low-rank adaptive matrix parameters introduced by the LoRA algorithm using an end-to-end multi-category contextual semantic extraction enhancement module based on LoRA.

[0046] S4, Multi-perspective Representation Extraction for Multi-category Sales Time Series: Based on the domain-fine-tuned language model, its parameters are completely frozen to obtain domain-fine-tuned enhanced static environment semantic embedding, predicted environment semantic embedding, and observed environment semantic embedding. Based on the multi-perspective representation learning method, the temporal evolution representation, channel correlation representation, and multivariate spatiotemporal dependency representation of multi-category sales time series data under the influence of multi-source heterogeneous external data are obtained respectively.

[0047] S5, multi-view representation fusion and prediction: The multi-view representations are concatenated and fused, and input into the gated residual network, and finally the multi-category sales volume at the future predicted time step is decoded through autoregressive prediction.

[0048] Compared with the existing technology, the present invention has the following beneficial effects: the present invention provides a multi-category sales joint prediction system and method based on large-model environmental semantic perception. In the multi-category sales prediction method, a pre-trained language model with world prior knowledge is introduced, and a unified text modeling and feature screening mechanism of multi-source heterogeneous external data is realized through a text generation module based on environmental denoising guided masking; the language model domain adaptation is realized through a channel-independent end-to-end training method; through a multi-perspective representation learning method, the multi-dimensional spatiotemporal dependency of multi-category sales time series under the influence of dynamic environment is successfully captured, and the effective fusion of environmental information, temporal information, spatial correlation and spatiotemporal dependency is realized; this method and system can improve the prediction effect of multi-category sales, and is effective both in normal periods and sudden abnormal periods. It can be applied to the daily operation of the retail supply chain platform, and provide a decision-making basis for its sales plan formulation, inventory management, commodity distribution resource optimization, etc., thereby improving the operating efficiency and operating profit of the entire supply chain platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a schematic diagram of the structure of the multi-category sales joint prediction system based on large-scale model environment semantic perception of the present invention;

[0050] Figure 2 This is a flowchart of the steps of the multi-category sales joint prediction method based on large model environment semantic perception of the present invention;

[0051] Figure 3 1 is a comparison chart of experimental results of different methods in the test examples of the present invention. DETAILED DESCRIPTION

[0052] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0053] Example 1

[0054] A multi-category sales joint prediction system based on large model environment semantic perception, such as Figure 1 As shown, it includes: data collection module, text generation module based on environmental denoising guided mask, end-to-end multi-category environmental semantic extraction enhancement module based on LoRA, multi-category sales joint prediction module based on multi-view representation learning,

[0055] The data collection module obtains time series data of sales of multiple categories to be predicted and multi-source heterogeneous external data that can be used to assist sales forecasting based on sales forecasting domain knowledge, including numerical, categorical and textual data. According to whether the data is time-varying, the data is further divided into time-varying features and static features. Among them, time-varying features are divided into predictive features and observation features according to whether they can only be obtained through retrospective observation.

[0056] The data collected by the data collection module should at least include the name of the category to be predicted, the sales region, and the name of the retailer; multi-category sales data, date data, holiday data, promotional activity data, weather data, and news data aligned in the time dimension with daily granularity, and numerical, categorical, and textual features to assist sales forecasting can be dynamically added based on domain knowledge.

[0057] The text generation module based on environment denoising guided mask: based on the collected multi-source heterogeneous external data, the environment denoising guided mask text generation template with enhanced annotations using domain knowledge is used to construct static context mask text descriptions of static features and dynamic context mask text descriptions of time-varying features respectively. Furthermore, a pre-trained language model such as DeBERTa is used to extract the mask token embeddings of each text description to obtain the denoised static environment semantic embedding, predicted environment semantic embedding and observed environment semantic embedding.

[0058] The text generation module based on the environmental denoising guided mask template first designs the conversion rules from multi-source heterogeneous data to text descriptions. For numerical and categorical features, we use metadata descriptions to explain the data source and annotate the features in the form of key-value pairs. For example, an example of a text description of a time feature is "Time information: Date: September 29, 2023, Day of the week: Friday, Type: Mid-Autumn Festival, Holiday: Yes." Descriptions between different source data are separated by line breaks; for text features, we also annotate the data in the form of metadata descriptions, but sort the text data by relevance or thermal value. For example, the environmental description text for hot search news is "Related news hot searches are as follows: 1. Why are moon cakes not selling well this year... 2. Traffic on the first day of the holiday attracted a peak in passenger flow..." Through predefined conversion rules, these multi-source heterogeneous features are first converted into a unified text description form.

[0059] Next, in order to improve the pre-trained large language model's ability to reason about environmental influences, we divided the environmental denoising guidance mask template into three parts, namely domain knowledge, environmental description, and task configuration. Domain knowledge is prior world knowledge summarized in the target prediction scenario. It is used to enhance the language model's prediction of the current environment's impact on the sales of the target category. The format is, "Product sales are time-periodic and affected by external factors. Pay attention to the following predicted / observed / static environmental information for <category> sales prediction:" This domain knowledge can be summarized by scenario business personnel or provided by a larger language model. The environmental description is a unified text description converted from the prepared multi-source heterogeneous data according to predefined data-to-text generation rules. Descriptions of different source data are separated by line breaks. The task configuration is the task instruction provided to the language model to align the model's understanding of the sales prediction task and stimulate the model's reasoning ability about environmental impacts. The format is, "Please infer the impact of this information on the sales of <category> is: [MASK]." Here, the [MASK] token is used in the form of a mask guided by the instruction to interact with the environmental information. The refined environmental information is extracted in the form of a latent vector in the high-dimensional embedding space. The final environmental impact is represented by the embedding of the [MASK] token in the last layer of the language model.

[0060] The LoRA-based end-to-end multi-category environmental semantic extraction enhancement module divides the multi-category sales forecasting task into multiple single-category sales forecasting tasks based on channel-independent training. It further introduces the LoRA method for efficient parameter fine-tuning, training only the newly introduced low-rank adaptive matrix parameters while ensuring the generalization capability of the pre-trained language model. Historical sales data is added to the environmental semantic embedding as input, and a fused embedding is obtained through a weighted network. The interaction between the dynamic environment and specific category sales is further learned through a gated residual network, an autoregressive LSTM model, and a causal masked attention mechanism. This end-to-end training approach achieves domain adaptation of the pre-trained language model, improving the multi-category environmental semantic embedding extraction capability.

[0061] In the end-to-end multi-category environment semantic extraction enhancement module based on LoRA, the LoRA algorithm is introduced into the self-attention layer of the pre-trained language model, and the original weight matrix W q 、W k 、W v 、W o , respectively introduce new low-rank adaptive parameters for subsequent fine-tuning adaptation.

[0062] In the end-to-end multi-category environmental semantic extraction enhancement module based on LoRA, assuming that a is the embedding to be fused and c is the environmental semantic embedding, the gated residual network GRN structure used is as follows:

[0063] GRNθ (a,c)=LayerNorm(a+GLU θ (η1))

[0064] η1=W 1,θ η2+b 1,θ

[0065] η2=ELU(W 2,θ a+W 3,θ c+b 2,θ )

[0066] a and c are concatenated and fed into a two-layer MLP perceptron network for fusion to obtain η1. ELU stands for exponential linear unit activation function, and the formula is:

[0067]

[0068] β is a hyperparameter greater than zero, usually set to 1. ELU avoids the problem of traditional activation function ReLU directly truncating negative values, thereby avoiding the problem of neuron death when processing negative values. The gated linear unit GLU is defined as GLU θ (γ)=σ(W 4,θ γ+b 3,θ )⊙(W 5,θ γ+b 4,θ ), where ⊙ represents the element-wise Hadamard product and σ is the classic sigmoid function used to calculate the gating value for the input embedding.

[0069] Specifically, through GLU module processing, the model can adaptively determine the degree of fusion of feature embedding. The ELU activation function can ensure the degree of processing of nonlinear transformation of feature embedding through the identity transformation of the negative part. These properties enable the gated residual network to provide adaptive flexibility for the fusion of environmental semantics.

[0070] The weighted network processing steps used in the fine-tuning module are as follows:

[0071] Use the linear layer to embed the sales value and get the vector

[0072] The environment text description constructed by the environment denoising guided mask is embedded through the pre-trained language model to obtain vectors

[0073] All vector embeddings are concatenated, By doing the following:

[0074]

[0075] The resulting vector That is weighted fusion embedding, where v t The m-dimensional weight tensor determines the importance of each type of input.

[0076] The multi-category sales joint prediction module based on multi-perspective representation learning is as follows: based on the domain fine-tuned language model, its parameters are completely frozen, and the static context mask text description and the dynamic context mask text description are input to obtain the static environment semantic embedding encoded by the domain fine-tuned language model, the predicted environment semantic embedding and the observed environment semantic embedding in the dynamic context mask text description. All environment semantic embeddings and historical sales numerical data are fused through a weighted network, and the temporal evolution representation, channel association representation and multivariate spatiotemporal dependency representation of each category are obtained through a multi-perspective representation learning method. These features are then fused and spliced and input into a gated residual network. Finally, the multi-category sales at the future predicted time step are decoded through an autoregressive prediction method.

[0077] The steps for processing temporal evolution representation in multi-perspective representation learning are as follows: reuse the environment-sales causal evolution module, obtain all environmental semantic embeddings based on the domain fine-tuning encoding language model with frozen parameters, fuse it with historical sales numerical data through a weighted network, further fuse long-term and short-term temporal dependencies through an autoregressive LSTM model and a causal mask attention mechanism, and finally obtain the temporal evolution representation before decoding through a gating mechanism and a residual module.

[0078] The channel association representation processing steps in multi-view representation learning are as follows: Based on the LSTM decoder embedding, the current time step environment semantic embedding is fused through different gated residual networks, and the fused embedding of all categories is further input into the two-layer Transformer encoder. Then, the long-term and short-term association gating mechanism is used to obtain the final fused adaptive multi-category association representation.

[0079] The steps for multivariate spatiotemporal dependency representation in multi-view representation learning are as follows: the multivariate sales series is input as two-dimensional data into the inverted Transformer encoder to obtain the two-dimensional spatiotemporal embedding of the time step to be predicted, and then the current time step environment semantic embedding is further fused through different gated residual modules to obtain a multivariate spatiotemporal dependency representation with enhanced environment information.

[0080] Finally, the temporal evolution representation, channel association representation, and multivariate spatiotemporal dependency representation of each category are spliced together, and then the future multi-category sales data is decoded through the gated residual module and linear output layer.

[0081] Example 2

[0082] A multi-category sales joint prediction method based on large model environment semantic perception, using the system as described in Example 1, such as Figure 2 As shown, the specific steps include:

[0083] Step S1, multi-source heterogeneous external data collection: Based on sales forecasting domain knowledge and the characteristics of multiple categories to be predicted, multi-category sales time series values and multi-source heterogeneous external environment data are obtained, including numerical, categorical, and textual data. Furthermore, based on whether the data is time-varying, the data is divided into static features, time-varying predictive features, and time-varying observation features.

[0084] Step S2, mask context text generation: A text generation module based on the environment denoising guided mask converts multi-source heterogeneous external data into static / predicted / observed context mask text descriptions enhanced with domain knowledge;

[0085] Step S3, domain fine-tuning of the pre-trained language model: The pre-trained language model is used to extract context semantic embeddings, which are combined with multi-category sales numerical data as input. The multi-category sales numerical data corresponding to the time step to be predicted is used as the output label. The multi-category sales prediction task is further divided into multiple single-category sales prediction tasks. Based on the LoRA-based end-to-end multi-category context semantic extraction enhancement module, the low-rank adaptive matrix parameters introduced by the LoRA algorithm are trained and updated.

[0086] Step S4, extracting multi-perspective representations of multi-category sales time series: Based on the domain-fine-tuned language model, completely freeze its parameters, input the static context mask text description and the dynamic context mask text description, and obtain the static environment semantic embedding encoded by the domain-fine-tuned language model, the predicted environment semantic embedding in the dynamic context mask text description, and the observed environment semantic embedding. Based on the designed multi-perspective representation learning method, the temporal evolution representation, channel correlation representation, and multivariate spatiotemporal dependency representation of the multi-category sales time series data under the influence of multi-source heterogeneous external data are obtained respectively;

[0087] Step S5, multi-view representation fusion and prediction: further concatenate and fuse the multi-view representations and input them into the gated residual network, and finally decode the multi-category sales volume at the future predicted time step through autoregressive prediction.

[0088] Test Case

[0089] The proposed method (DT-CAMP) is compared with other advanced sales prediction algorithms, TFT and NSTransformer, as well as the recent sales prediction algorithm GPT4TS based on a large language model. The results are tested on two Chinese and foreign retail datasets to verify the effectiveness of the proposed method.

[0090] Dataset 1 is a dataset from a retail website. We use order data from the Beijing area collected from public websites to calculate daily sales of multiple product categories. The categories cover 65 leaf-node categories under five primary categories (food and beverages, healthcare, sports and outdoor, home cleaning, and household daily necessities), reflecting the diverse needs of people in their daily lives. We collected multi-source external data through web crawlers and API programs, including holiday calendars, weather information, promotions, sporting events, public health events, product search indexes, and hot news topics. A total of 173,784 news items were collected. The dataset covers a period of 17 months, from April 1, 2022, to August 31, 2023.

[0091] Dataset 2 is another large supermarket dataset. This dataset comes from the M5 forecasting competition on the Kaggle platform. The data for this competition was provided by a major retail company and includes daily sales records generated by the company's actual operations, as well as auxiliary data such as holiday calendars. The M5 competition collected demand data from 21 product categories in three US states (California, Texas, and Wisconsin) between January 29, 2011, and May 22, 2016. The M5 competition also supplemented the environmental data with a US natural disaster dataset, a comprehensive event dataset, the Global News Media GDELT dataset, and weather forecast data for these states. This dataset contains a total of 629,786 news items.

[0092] The proposed method (DT-CAMP) was tested on dataset 1 and dataset 2 with other sales forecasting algorithms. The weighted mean absolute percentage error (wMAPE) was used as the evaluation metric. The experimental results are shown in Figure 2. Figure 3As shown in the figure, the results of each experiment from left to right correspond to DT-CAMP, NSTransformer, TFT and GPT4TS respectively. The results show that the sales forecasting method proposed in this invention has achieved the best prediction effect in the two data sets with different data statistical characteristics. In data set 1, the wMAPE of DT-CAMP with prediction step size of 1, 7 and 14 configurations are 0.1442, 0.1789 and 0.1901 respectively, and the performance of the best baseline is 0.1556, 0.1983 and 0.2113 respectively, which are optimized by 7.33%, 9.78% and 10.02% respectively; in data set 2, the wMAPE of DT-CAMP with prediction step size of 1, 7 and 14 configurations are 0.0681, 0.0756 and 0.0796 respectively, and the performance of the best baseline is 0.0757, 0.0896 and 0.0944 respectively, which are optimized by 10.04%, 15.63% and 15.68% respectively. From the experimental results, it can be seen that DT-CAMP not only achieves the best performance under each step size setting, but also the optimization effect gets better with the increase of prediction step size, which shows that DT-CAMP has better robustness and generalization than other baseline methods in uncertain dynamic environments.

[0093] It should be noted that the above content merely illustrates the technical idea of the present invention and cannot be used to limit the scope of protection of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications all fall within the scope of protection of the claims of the present invention.

Claims

1. A multi-category sales joint prediction system based on large-scale model environment semantic perception, characterized by: It includes at least a data collection module, a text generation module based on environmental denoising guided mask, an end-to-end multi-category environmental semantic extraction enhancement module based on LoRA, and a multi-category sales joint prediction module based on multi-view representation learning. The data collection module is used to obtain multi-category sales time series data and multi-source external data; the multi-source external data includes numerical, categorical, and textual data, which are used to assist in sales forecasting; The text generation module based on the environment denoising guided mask is used to construct the collected data into a masked text description enhanced with sales forecast domain knowledge; using the pre-trained language model to extract the masked token embedding of each text description to obtain the denoised static environment semantic embedding, the predicted environment semantic embedding, and the observed environment semantic embedding; The LoRA-based end-to-end multi-category environmental semantic extraction enhancement module: fine-tunes the pre-trained language model domain through multiple channel-independent single-category sales prediction tasks; the pre-trained language model uses the LoRA algorithm in its self-attention layer, adds historical sales numerical data based on the environmental semantic embedding as input, obtains a fused embedding through a weighted network, and learns the interactive relationship between the dynamic environment and specific category sales through a gated residual network, an autoregressive LSTM model, and a causal masked attention mechanism, achieving domain adaptation of the pre-trained language model through an end-to-end training method; The multi-category sales joint prediction module based on multi-perspective representation learning includes an environment-sales causal evolution module, a channel association mining module, a multi-variable spatiotemporal dependency mining module, a multi-perspective representation fusion module, and a multi-category sales prediction module. The environment-sales causal evolution module is used to capture the temporal evolution pattern; the channel association mining module is used to capture multi-channel associations; the multivariate spatiotemporal dependency mining module is used to capture the spatiotemporal dependency of multivariate sales series; The multi-view representation fusion module and the multi-category sales prediction module are used for multi-view feature fusion and multi-category sales prediction respectively; In the multi-category sales joint prediction module based on multi-perspective representation learning, all environmental semantic embeddings and historical sales numerical data are fused through a weighted network based on the static environment semantic embedding, predicted environment semantic embedding and observed environment semantic embedding enhanced by domain fine-tuning. The temporal evolution representation, channel association representation and multivariate spatiotemporal dependency representation of each category are obtained through the multi-perspective representation learning method. The feature fusion and splicing are input into the gated residual network, and the multi-category sales in the future time step to be predicted are decoded through autoregressive prediction.

2. The multi-category sales joint prediction system based on large-scale model environment semantic perception according to claim 1 is characterized by: The data collected by the data collection module includes at least the names of multiple categories corresponding to the sales series to be predicted, the sales regions, the names of retailers, daily sales data aligned in the time dimension, date data, holiday data, promotional activity data, weather data and news hot search data.

3. The multi-category sales joint prediction system based on large-scale model environment semantic perception according to claim 1 is characterized by: The template in the text generation module based on environmental denoising guided mask includes domain knowledge, environmental description and task configuration; the domain knowledge is the prior world knowledge summarized in the target prediction scenario, which is used to enhance the language model's prediction of the impact of the current environment on the sales of the target category; the environmental description is a unified text description converted from the prepared multi-source heterogeneous data according to predefined data-to-text generation rules; the task configuration is the task instruction provided to the language model, which is used to align the model's understanding of the sales prediction task and stimulate the model's reasoning ability about the impact of the environment.

4. The multi-category sales joint prediction system based on large-scale model environment semantic perception according to claim 3 is characterized by: The specific structure of the gated residual network (GRN) used in the end-to-end multi-category environment semantic extraction enhancement module based on LoRA is as follows: GRN θ (a,c)=LayerNorm(a+GLU θ (η1)) η1=W 1,θ η2+b 1,θ η2=ELU(W 2,θ a+W 3,θ c+b 2,θ ) In the above formula, θ is the index of the gated residual network with different parameters, a is the embedding to be fused, c is the environment semantic embedding, LayerNorm is the standard layer normalization function, η1 and η2 represent the intermediate embeddings of the processing, GLU represents the gated linear unit, and W 1,θ To W 3,θ and b 1,θ to b 2,θ are the neural network parameters to be learned, ELU represents the exponential linear unit activation function, and the formula is: Among them, β is a hyperparameter greater than zero, e x is the natural exponential function.

5. The multi-category sales joint prediction system based on large-scale model environment semantic perception according to claim 4 is characterized by: In the end-to-end multi-category environment semantic extraction enhancement module based on LoRA, the fusion embedding is obtained through the weighted network. The specific method is: the sales value is embedded using the linear layer to obtain the vector The environment text description constructed by the environment denoising guided mask is embedded through the pre-trained language model to obtain vectors All vector embeddings are concatenated, By doing the following: The resulting vector is the weighted fusion embedding, where v t It is an m-dimensional weight tensor. The corresponding weight is calculated according to the softmax function and is weightedly fused with the m input vector embeddings.

6. The multi-category sales joint prediction system based on large-scale model environment semantic perception according to claim 1 is characterized by: In the environment-sales causal evolution module, the semantic embedding of the input environment text description is obtained based on the language model, and it is fused with the historical sales numerical data through a weighted network. The long-term and short-term temporal dependencies are fused through the LSTM model and the causal mask attention mechanism, and the temporal evolution representation is obtained through the gated GLU unit and residual connection. Among them, c represents the category identifier, t represents the time step to be predicted, is the hidden state embedding output by the LSTM decoder, It is the fusion environment semantic embedding, Represents the intermediate representation at the sequence level, and MaskedAttention represents the unidirectional causal mask attention mechanism.

7. The multi-category sales joint prediction system based on large-scale model environment semantic perception according to claim 1 is characterized by: In the channel association mining module, the hidden state vector obtained by the LSTM decoder and the cell state vector Based on this, different gated residual networks are used to fuse the semantic embedding of the current time step environment. Get fused embedding and The fusion embedding corresponding to all C categories is input into the two-layer Transformer encoder to obtain the long-term and short-term channel correlation representation and The fused channel correlation representation is obtained based on the long-term and short-term correlation gating mechanism: Among them, FC represents the fully connected layer, σ represents the sigmoid activation function, and ⊙ represents the element-by-element product. It is an adaptive multi-category association representation.

8. The multi-category sales joint prediction system based on large-scale model environment semantic perception according to claim 1 is characterized by: In the multivariate spatiotemporal dependency mining module, the multivariate sales series is taken as a whole to form a two-dimensional tensor data and input into the inverted Transformer encoder to obtain the two-dimensional spatiotemporal embedding of the time step to be predicted: Among them, x :,c is the sales volume sequence of category c, which is encoded into variable token embedding through the linear embedding layer Embedding(·) Aggregated into a two-dimensional tensor H in the category dimension ST , input into the Transformer encoder, and when outputting the projection, it will pass through the linear layer Projection of the corresponding time step t (·) Align with the current time step t to obtain the fused embedding By fusing the current time step environment semantic embedding through different gated residual GRN modules, we can obtain a multivariate spatiotemporal dependency representation with enhanced environment information. in, is the environment semantic embedding, is the temporal hidden state vector obtained in the environment-sales causal evolution module; The temporal evolution representation, channel correlation representation, and multivariate spatiotemporal dependency representation of each category are spliced together, and the predicted sales volume of each category is decoded through a gated residual module and a linear output layer: in, Represents the predicted sales of category c at time step t.

9. A multi-category sales joint prediction method based on large model environment semantic perception using the system as claimed in claim 1, characterized in that: The steps include: S1, multi-source heterogeneous external data collection: Based on sales forecasting domain knowledge and the characteristics of multiple categories to be predicted, obtain multi-category sales time series values and multi-source heterogeneous external environment data, and divide the data into static features, time-varying predicted features, and time-varying observed features; S2, Masked Context Text Generation: A text generation module based on environment denoising-guided masking, which converts multi-source heterogeneous external data into static / predicted / observed context masked text descriptions enhanced with domain knowledge; S3, domain fine-tuning of the pre-trained language model: This involves extracting contextual semantic embeddings from the pre-trained language model and combining them with multi-category sales numerical data as input. The multi-category sales numerical data corresponding to the time step to be predicted is used as the output label. This splits the multi-category sales prediction task into multiple single-category sales prediction tasks. This involves training and updating the low-rank adaptive matrix parameters introduced by the LoRA algorithm using an end-to-end multi-category contextual semantic extraction enhancement module based on LoRA. S4, Multi-perspective Representation Extraction for Multi-category Sales Time Series: Based on the domain-fine-tuned language model, its parameters are completely frozen to obtain domain-fine-tuned enhanced static environment semantic embedding, predicted environment semantic embedding, and observed environment semantic embedding. Based on the multi-perspective representation learning method, the temporal evolution representation, channel correlation representation, and multivariate spatiotemporal dependency representation of multi-category sales time series data under the influence of multi-source heterogeneous external data are obtained respectively. S5, multi-view representation fusion and prediction: The multi-view representations are concatenated and fused, and input into the gated residual network, and finally the multi-category sales volume at the future predicted time step is decoded through autoregressive prediction.