Catering take-out consumption information prediction method based on large language model and text mining

By combining large language models and text mining technology with data on food delivery, socio-economic factors, and the environment, the problem of limited data sources and insufficient deep semantic capture in food delivery consumption forecasting has been solved, resulting in more accurate and refined consumption information forecasting.

CN121146820APending Publication Date: 2025-12-16GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511083446.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing methods for predicting food delivery consumption rely on a single data source, making it difficult to capture deep semantic information, resulting in insufficient accuracy and precision in prediction.

Method used

This study employs a method based on large language models and text mining. By acquiring food delivery consumption data, socioeconomic data, and environmental data of the target region, and using text mining models for vectorization, the study determines food delivery consumption preference vectors. Combined with large language models, the study predicts consumption volume, thereby improving the accuracy and precision of predictions.

Benefits of technology

It enables refined analysis of food delivery consumption data, improves the accuracy and precision of predictions, and provides reliable data references for the analysis of food delivery consumption information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121146820A_ABST
    Figure CN121146820A_ABST
Patent Text Reader

Abstract

The invention provides a catering take-out consumption information prediction method based on a large language model and text mining. The method comprises the steps of obtaining catering take-out consumption data, social economic data and environmental data of a target area in a target historical time period; based on a text mining model, performing vectorization processing on the catering take-out consumption data to determine a catering take-out consumption preference vector corresponding to the catering take-out consumption data; on the basis of a large language model, according to the social economic data and the environmental data, predicting the catering take-out consumption amount of the target area; and based on a catering take-out consumption information prediction model, according to the catering take-out consumption data, the catering take-out consumption preference vector and the catering take-out consumption amount, determining catering take-out consumption information of the to-be-predicted area. According to the invention, the prediction accuracy and fineness of the catering take-out consumption information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of information prediction, and particularly relates to a takeout consumption information prediction method based on a large language model and text mining. BACKGROUND

[0002] With the continuous expansion of the takeout market, the size of the takeout user market shows a rapid upward trend, and the takeout consumption demand shows a diversification and health trend. Based on this, the importance of takeout consumption green transformation and takeout healthy diet guidance also increases; by predicting takeout consumption information, the takeout consumption preferences and consumption structure can be analyzed in a timely manner, thereby providing data reference for takeout consumption green transformation, packaging waste management and takeout healthy diet guidance. However, in the related art, the data source dimension of the takeout consumption prediction is single, resulting in that the result obtained by the prediction is mostly the total amount of takeout consumption, and the information contained is less; and the takeout consumption prediction method in the related art is difficult to capture the deep semantic information in the takeout consumption record, resulting in that in the prediction process, the takeout consumption prediction method in the related art less utilizes data other than the takeout consumption amount for prediction, thereby causing the accuracy of the prediction result to have a deviation. It can be seen that the related art has problems such as insufficient prediction accuracy and fineness of takeout consumption information. SUMMARY

[0003] The application provides a takeout consumption information prediction method based on a large language model and text mining, which aims to improve the prediction accuracy and fineness of takeout consumption information.

[0004] The takeout consumption information prediction method based on a large language model and text mining provided by the application comprises the following steps:

[0005] Obtain takeout consumption data, social and economic data and environmental data of a target region in a target historical time period;

[0006] Based on a text mining model, the takeout consumption data is subjected to vectorization processing to determine a takeout consumption preference vector corresponding to the takeout consumption data;

[0007] Based on a large language model, the takeout consumption amount of the target region is predicted according to the social and economic data and the environmental data;

[0008] Based on a takeout consumption information prediction model, the takeout consumption information of a region to be predicted is determined according to the takeout consumption data, the takeout consumption preference vector and the takeout consumption amount.

[0009] The application carries out data mining and vectorization processing on the catering take-out consumption data in the target historical time period through a text mining model, determines the catering take-out consumption preference vector corresponding to the target area in the catering take-out consumption data, captures the semantics of unstructured data in the catering take-out consumption data, and can finely analyze and process the catering take-out consumption data, and can analyze the influence of unstructured data in the catering take-out consumption data on catering take-out consumption, to provide a basis for fine catering take-out consumption information prediction and improve the prediction accuracy of catering take-out consumption information; through the data analysis ability of large language model on social and economic data and environmental data, the accuracy of catering take-out consumption information prediction in dynamic scene is improved; and the prediction accuracy and fineness of catering take-out consumption information are improved to provide reliable data reference for analysis of catering take-out consumption information. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 A flowchart of a catering take-out consumption information prediction method based on a large language model and text mining provided by an embodiment of the application. DETAILED DESCRIPTION

[0012] The technical solutions in the embodiments of the application will be described clearly and completely in the following with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are some embodiments of the application, not all embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.

[0013] The flowchart shown in the drawings is only an example, not necessarily including all contents and operations / steps, and not necessarily executed in the described order. For example, some operations / steps can be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.

[0014] The embodiment of the application provides a catering take-out consumption information prediction method based on a large language model and text mining. The catering take-out consumption information prediction method based on a large language model and text mining can be applied to a terminal device, which can be a notebook computer, a desktop computer, a cloud server and other electronic devices.

[0015] Some embodiments of the present application will be described in detail with reference to the drawings. The following examples and features in the examples can be combined with each other without conflict.

[0016] Please refer to Figure 1 , Figure 1 A flowchart of a large language model and text mining based catering take-out consumption information prediction method provided by an embodiment of the present application.

[0017] As Figure 1 shown, the large language model and text mining based catering take-out consumption information prediction method includes steps S101 to S104.

[0018] Step S101, obtain catering take-out consumption data, social economic data and environmental data of a target region in a target historical time period.

[0019] For example, obtain catering take-out consumption data of a target region in a target historical time period, social economic data of the target historical time period and environmental data of the target historical time period from multiple data sources. In some embodiments, the data sources include but are not limited to catering take-out web pages and / or applications, social economic information public web pages or applications, and environmental information public web pages or applications. Obtain relevant public data in the above-mentioned web pages and / or applications to obtain corresponding catering take-out consumption data, social economic data and environmental data. In some embodiments, the catering take-out consumption data includes but is not limited to catering take-out dish names, catering take-out dish types, catering take-out dish descriptions, etc.; in other embodiments, the catering take-out consumption data also includes catering take-out dish carbon footprint and catering take-out dish nutritional value, etc.; and the social economic data includes but is not limited to the number of permanent residents in the target region, population migration index (inflow, outflow, net inflow), employment population in each industry, population density, catering take-out Engel coefficient, road network density, per capita GDP, mobile internet penetration rate, per capita built-up area, industrial structure, etc.; and the environmental data includes but is not limited to monthly average temperature, monthly average PM 2.5 concentration and monthly average precipitation, etc.

[0020] For example, the target region includes a region to be predicted and at least one non-predicted region, wherein the region to be predicted and the non-predicted region do not overlap; the division of the region can be divided according to administrative regions or prefecture-level cities, or can be divided according to actual catering take-out consumption information prediction needs, which is not limited by the present application.

[0021] In some embodiments, the target area's to-be-cleaned catering delivery consumption data, to-be-cleaned social and economic data, and to-be-cleaned environmental data in a target historical time period are acquired from multiple data sources; and based on a preset data cleaning rule, the to-be-cleaned catering delivery consumption data, to-be-cleaned social and economic data, and to-be-cleaned environmental data are subjected to data cleaning to obtain catering delivery consumption data, social and economic data, and environmental data. For example, special characters in various to-be-cleaned data and non-catering delivery dish texts in to-be-cleaned catering delivery consumption data are deleted to achieve the effect of data cleaning and reduce the impact of noise data on information prediction.

[0022] In step S102, the catering delivery consumption data is subjected to vectorization processing based on a text mining model to determine a catering delivery consumption preference vector corresponding to the catering delivery consumption data.

[0023] For example, the acquired catering delivery consumption data is input into the text mining model to perform text mining processing on the catering delivery consumption data in the text mining model to determine a catering delivery consumption preference vector. Specifically, the catering delivery consumption data is subjected to vectorization processing to obtain a catering delivery consumption preference vector corresponding to the catering delivery consumption data.

[0024] In some embodiments, the catering delivery consumption data includes multiple catering delivery dishes, and each catering delivery dish has corresponding catering delivery dish name and catering delivery dish type data. After the catering delivery consumption data is input into the text mining model, the catering delivery dish corresponding to the catering delivery dish name and the catering delivery dish type data can be determined to determine the catering delivery dish attribute feature vector, and thus the catering delivery consumption preference vector corresponding to each catering delivery dish can be determined.

[0025] For example, the catering delivery dish name is steamed dumplings with soup, fried rice, etc., and the catering delivery dish type is used to indicate the cuisine of the dish. The data displayed in the catering delivery dish type includes but is not limited to Chinese snacks, Chinese fast food, Western fast food, drinks, baked cakes, seafood barbecue, porridge, rice, noodles, etc. For example, the catering delivery dish type corresponding to the catering delivery dish name steamed dumplings with soup is Chinese snacks, the catering delivery dish type corresponding to the catering delivery dish name fried rice is porridge, rice, noodles, etc. The catering delivery dish name and the catering delivery dish type can be obtained when the catering delivery consumption data is acquired, so that the catering delivery dish attribute feature vector corresponding to the catering delivery dish name and the catering delivery dish type data can be determined.

[0026] In some embodiments, the vectorization processing of the catering take-out consumption data based on the text mining model to determine the catering take-out consumption preference vector corresponding to the catering take-out consumption data comprises: based on the word segmentation network in the text mining model, performing word segmentation screening processing on the unstructured data in the catering take-out consumption data to obtain a plurality of catering take-out dishes each corresponding to catering take-out dish information text; based on the vectorization network in the text mining model, performing vectorization processing on the catering take-out dish information text to obtain a first dish attribute feature vector corresponding to each of the catering take-out dishes; based on the normalization network in the text mining model, performing normalization processing on the structured data in the catering take-out consumption data to obtain a second dish attribute feature vector corresponding to each of the catering take-out dishes; and based on the vector processing network in the take-out consumption preference prediction model, determining the catering take-out consumption preference vector according to the first dish attribute feature vector and the second dish attribute feature vector.

[0027] For example, the text mining model includes a word segmentation network, a vectorization network, a normalization network, and a vector processing network. The word segmentation network is used to perform word segmentation processing on the unstructured text in the catering take-out consumption data and transmit the word segmentation processing result to the vectorization network for vectorization processing, thereby obtaining a first dish attribute feature vector corresponding to the catering take-out dish. The normalization network is used to perform normalization processing on the structured data in the catering take-out consumption data, thereby obtaining a second dish attribute feature vector corresponding to each of the catering take-out dishes, and then in the vector processing network, the catering take-out consumption preference vector is determined according to the first dish attribute feature vector corresponding to each of the catering take-out dishes and the second dish attribute vector corresponding to each of the catering take-out dishes.

[0028] In some embodiments, in the catering take-out consumption data, the catering take-out dish name, the catering take-out dish type, and the catering take-out dish description information corresponding to each of the catering take-out dishes can be determined, and the three unstructured texts are connected to obtain a catering take-out dish document. Based on a preset word segmentation algorithm, the catering take-out dish document is segmented to obtain the catering take-out dish information text corresponding to the catering take-out dish; and the TF-IDF algorithm is used to determine the weight value of each catering take-out dish information text in the catering take-out dish document corresponding to each catering take-out dish.

[0029] In the vectorization network, the take-out food information text in each take-out food document is vectorized and encoded to obtain a vector matrix of [1 x v], where v is used to indicate the number of take-out food information texts in each take-out food document. For example, taking a take-out food of spicy chicken burger as an example, the corresponding food description information is: "tender and juicy chicken leg meat, wrapped with secret spicy sauce and fried to golden crisp, with fresh lettuce and special sauce, sandwiched in soft and sweet burger buns, one bite, crispy outside and tender inside, spicy and satisfying, and fragrant!", and the corresponding take-out food type is "foreign cuisine". After segmentation, the result is obtained as: [spicy chicken burger, chicken leg meat, spicy sauce, lettuce, bread, …, crispy outside and tender inside, foreign cuisine]. The vector matrix corresponding to the spicy chicken burger is [1, 0, 0, 0, …] (dimension v).

[0030] After determining the vector matrix corresponding to the take-out food, the vector matrix is multiplied by the input weight matrix W [V x N] to obtain the embedding corresponding to the take-out food, where embedding is a matrix of [1 x N] size, and N is used to indicate the vector dimension of the output word. The average of the embedding corresponding to each take-out food is taken as the hidden layer vector h, and then multiplied by the output weight matrix W' [N x V] to obtain a vector y of size [1 x V]. The vector y is processed by using a softmax activation function to obtain a V-dim probability distribution, thereby obtaining a first food attribute feature vector, and the first food attribute feature vector is transmitted to the vector processing network.

[0031] For example, the vectorization network can be constructed by the Skip-gram model in the Word2Vec model.

[0032] For example, during the training of the text mining model, the input weight W and / or the embedding corresponding to each take-out food can be updated according to the model error, thereby improving the prediction accuracy of the text mining model.

[0033] For example, during the determination of the unstructured text in the take-out food consumption data, the structured data in the take-out food consumption data can also be determined, including but not limited to the carbon footprint of the take-out food corresponding to the take-out food and the nutritional value of the take-out food. The structured data in the take-out food consumption data is normalized to obtain a corresponding second food attribute feature vector, and the second food attribute feature vector is input into the vector processing network.

[0034] For example, in the normalization network, the structured data in the catering take-out consumption data is normalized based on the Z-score normalization method to obtain a second dish attribute feature vector.

[0035] For example, the vector processing network receives the first dish attribute feature vector output by the vectorization network and the second dish attribute feature vector output by the normalization network to determine a catering take-out consumption preference vector according to the first dish attribute feature vector and the second dish attribute feature vector.

[0036] In some embodiments, the vector processing network in the text mining model determines the catering take-out consumption preference vector according to the first dish attribute feature vector and the second dish attribute feature vector, including: splicing the first dish attribute feature vector of each catering take-out dish with the corresponding second dish attribute feature vector to obtain a target feature vector corresponding to each catering take-out dish; determining a mean vector according to the target feature vectors of multiple catering take-out dishes, and determining the mean vector as the catering take-out consumption preference vector.

[0037] For example, according to the catering take-out dishes, the corresponding first dish attribute feature vector and the second dish attribute feature vector are determined, and the corresponding first dish attribute feature vector and the second dish attribute feature vector are subjected to vector splicing processing to obtain a target feature vector corresponding to the catering take-out dish. For example, the first dish attribute feature vector corresponding to the catering take-out dish str1 is (w11, w12, …, w1m), and the second dish attribute feature vector corresponding to the catering take-out dish str1 is (b1, c1, …). After vector splicing, the target feature vector corresponding to the catering take-out dish str1 is (w11, w12, …, b1, c1, …), and so on. Thus, the target feature vectors corresponding to the catering take-out dishes str2 and str3 can be obtained.

[0038] For ease of understanding, the catering take-out dishes str1, str2 and str3 and their corresponding target feature vectors can be shown in the following catering take-out consumption data table.

[0039]

[0040] 5For example, after obtaining the target feature vectors corresponding to each catering take-out dish, the target feature vectors of multiple catering take-out dishes are used to determine a catering take-out consumption preference vector. For example, taking the target feature vectors (w11, w12, w13, …, b1, c1, …) and (w21, w22, w23, …, b2, c2, …) as examples, the catering take-out consumption preference vector is

[0041] ​ It should be understood that the mean vector obtained by averaging the vectors corresponding to each column of the above table is the catering take-out consumption preference vector.

[0042] By text mining the catering take-out consumption data through the segmentation network and vectorization network in the text mining model, the text semantics in the catering take-out consumption data can be identified, so that the influence of the text semantics on the catering take-out consumption information prediction process can be paid attention to, thereby improving the accuracy of the catering take-out consumption information prediction; and the target feature vector obtained by text mining includes multi-dimensional data, thereby improving the refinement degree of predicting the catering take-out consumption information.

[0043] In step S103, based on the large language model, the catering take-out consumption of the target region is predicted according to the social and economic data and the environmental data.

[0044] For example, by inputting the social and economic data and the environmental data into the large language model to predict the catering take-out consumption, the influence of the social and economic factors and the environmental factors on the catering take-out consumption is predicted, thereby improving the accuracy of the catering take-out consumption information prediction.

[0045] In some embodiments, the social and economic data of a region includes the number of permanent residents of the region, the population migration index (inflow, outflow, net inflow), the number of employed persons in each industry, the population density, the catering take-out Engel coefficient, the per capita GDP, the age structure, the industrial structure, etc., and the environmental data of the region includes the monthly average temperature, the monthly average PM 2.5 emission, the monthly average precipitation, etc., so that by the large language model, the catering take-out consumption of the region is predicted according to the number of permanent residents, the population density, the catering take-out Engel coefficient, the mobile internet penetration rate, the road network density, the per capita GDP, the age structure, the industrial structure, the number of employed persons in each industry, the monthly average temperature, the monthly average PM 2.5 emission, the monthly average precipitation, etc., so that the prediction of the catering take-out consumption can adapt to dynamic scenarios, thereby improving the accuracy of the catering take-out consumption information.

[0046] In some embodiments, the prediction of the catering take-out consumption of the target region based on the large language model and the social and economic data and the environmental data includes: based on the data extraction network of the large language model, extracting first target data from the social and economic data and second target data from the environmental data, and filling the first target data and the second target data into a preset text template to obtain a prompt text; based on the prediction network of the large language model, predicting the catering take-out consumption of the target region according to the prompt text and historical catering take-out consumption information.

[0047] Exemplarily, after the socio-economic data and the environmental data are input into the large language model, the first target data is extracted from the socio-economic data and the second target data is extracted from the environmental data by a data extraction network in the catering take-out consumption model. For example, the socio-economic data and the environmental data are mostly obtained in the form of reports, and the data extraction network is used to extract the corresponding data in the corresponding report.

[0048] Exemplarily, after the first target data and the second target data are obtained, the first target data and the second target data are filled into a preset text template to obtain a prompt text. The prompt text is used to prompt the prediction network to make a prediction according to the corresponding data, so as to realize the attention to the influence of the socio-economic factors and the environmental factors on the catering take-out consumption.

[0049] Exemplarily, in the prediction network of the large language model, the catering take-out consumption in the target region is predicted according to the input prompt text and obtained historical catering take-out consumption information. In some embodiments, the historical catering take-out consumption information at least includes historical catering take-out consumption, socio-economic information corresponding to the historical catering take-out consumption, and environmental information corresponding to the historical catering take-out consumption, so that the prediction network can predict the catering take-out consumption in the target region according to the historical catering take-out consumption information and the actual socio-economic data and environmental data of the target region.

[0050] For example, the prediction network includes a large language model, and the large language model can reconstruct the prediction task into a natural language processing problem, so that the large language model can fully consider the dynamic response of multiple heterogeneous factors to a variable over time in the natural language processing process, and realize dynamic scene prediction of the catering take-out consumption.

[0051] In some embodiments, the preset text template is as follows:

[0052] "You are a catering take-out market researcher, and you are good at predicting the catering take-out consumption in a region through historical catering take-out consumption information and regional related information.

[0053] We are currently located in the XX region {regional feature description}.

[0054] The monthly sales of catering take-out in this region in the past 12 months are:

[0055] The population migration index in this region in the past 12 months is:

[0056] The average monthly temperature in this region in the past 12 months is:

[0057] The catering take-out Engel coefficient in this region is:

[0058] The average monthly PM 2.5Discharge:

[0059] Internet penetration rate of the region: %

[0060] Road network density of the region: km / km 2 ;

[0061] Population density of the region: person / square kilometer

[0062] GDP per capita of the region: ten thousand yuan

[0063] Population age structure of the region: %

[0064] Industrial structure of the region: %

[0065] Employment population of the region by industry: people

[0066] It should be understood that the colon after each sentence of the prompt template is used to fill in the corresponding first target data and second target data to obtain the prompt text after completing the data filling. It should be noted that those skilled in the art can add, delete or adjust the expressions of the sentences in the preset text template according to the data used for prediction and the actual situation of the region, and the present application does not limit the specific sentences in the preset text template.

[0067] For example, using the prediction and derivation process of the large language model based on the prompt text, the influence mechanism of multiple heterogeneous factors on the consumption of catering take-out can be visualized, and the prediction result of the consumption of catering take-out can be obtained.

[0068] For example, the input expression of the large language model is X=(P, T, Φ), where P is used to indicate the preset text template, T is used to represent the region feature description, the first target data and / or the second target data, and Φ is used to indicate the historical catering take-out consumption information corresponding to the preset historical time period. The large language model can convert the input X into discrete tokens of a vocabulary, so as to predict the corresponding catering take-out consumption according to the discrete tokens.

[0069] For ease of understanding, the following will input the provided region A and its related information, and the region B and its related information into the large language model for prediction. It has been verified that the large language model can realize the prediction of the catering take-out consumption of the region according to the input of the corresponding related information of the region. The specific process can be referred to in the following.

[0070] Taking region A as an example, the restaurant take-out information and the region related information corresponding to region A are input into a preset text template to obtain a prompt text, and the prompt text is input into a large language model to obtain a prediction result of the large language model, wherein the historical restaurant take-out information corresponding to region A provided in this embodiment includes historical restaurant take-out consumption in the last 10 months, and the region related information provided includes social and economic data and environmental data, the social and economic data specifically includes population inflow, population outflow, population net inflow and night light data in the last 10 months; the environmental data includes air temperature and precipitation in the last 10 months. In addition to the prediction result of the restaurant take-out consumption of region A, it should be noted that the prediction result of other information is used to indicate other information of region A in the to-be-predicted time period, since the to-be-predicted time period is a future time, other information also needs to be predicted to determine the restaurant take-out information of region A in the to-be-predicted time period except for the restaurant take-out consumption, and the region related information in the to-be-predicted time period, so as to be able to predict the restaurant take-out consumption of region A according to these information.

[0071] For example, the prompt text corresponding to region A is as follows:

[0072] "You are a restaurant take-out market research personnel, good at predicting the restaurant take-out consumption of a region through historical restaurant take-out consumption information and region related information.

[0073] We are currently located in region A.

[0074] The following historical restaurant take-out consumption information and region related information corresponding to the historical time are provided:

[0075] Historical restaurant take-out consumption: {49837918, 86509888, 73604312, 75064669, 73648274, 75752875, 83432396, 42587881, 87656780, 42007811}.

[0076] Population inflow in the last 10 months: {252.7243, 358.5482, 232.8703, 192.6351, 269.6569, 332.8365, 357.5766, 388.1798, 322.5205, 333.6227}.

[0077] Population outflow in the last 10 months: {416.4049, 198.0849, 227.7947, 216.3218, 268.3228, 336.0689, 372.24, 379.0483, 330.1338, 328.2513}.

[0078] Population net inflow for the last 10 months: {-163.6806, 160.4633, 5.0756, -23.6867, 1.3341, -3.2324, -14.6634, 9.1315, -7.6136, 5.3714}.

[0079] Night light data for the last 10 months: 27.2644, 25.4147, 30.0196, 28.2005, 25.0444, 27.8743, 0.2357, 26.9861, 27.7245, 29.7993.

[0080] Environmental data includes air temperature and precipitation, specifically, the average monthly air temperature for the last 10 months: 14.9212, 12.1623, 20.2402, 21.9913, 23.5746, 27.6743, 29.6143, 28.9371, 28.2362, 24.387; the average monthly precipitation for the last 10 months: 19.45, 230.0158, 154.1247, 86.2382, 400.1657, 295.1908, 154.1818, 243.9956, 66.0453, 25.9762.

[0081] In addition, the time to be predicted is next month, and other relevant prediction information for next month is as follows:

[0082] Air temperature: {26.7}.

[0083] Precipitation: {340.5}.

[0084] Population inflow: {332.98}.

[0085] Population outflow: {300.95}.

[0086] Night light data: {30.134}.

[0087] Population net inflow: {22.03}.

[0088] Now please note that your task is to predict the takeout consumption for next month based on the information provided above, please provide the prediction value directly in the angle brackets <value>."

[0089] Thus, the large language model can output the prediction result corresponding to region A based on the above prompt text, as follows:

[0090] "Okay, I need to predict the takeout consumption in region A next month. First, I need to carefully analyze the historical data provided by the user and the predicted data for the next month. The user has provided 10 months of takeout consumption data, temperature, precipitation, population migration, and nighttime light data, as well as specific values for these variables for the next month.

[0091] First, let's look at the historical takeout consumption: 49837918, 86509888, 73604312, 75064669, 73648274, 75752875, 83432396, 42587881, 87656780, 42007811. These numbers fluctuate quite a bit, such as suddenly rising to 86.5 million in the second month, and then falling. There may be seasonal factors or special events that affect this. We need to see if there are any obvious trends or periodicities in these data.

[0092] Next is the temperature data, ranging from 14.92 to 29.61, with a prediction of 26.7 for next month. Region A is in a subtropical climate with relatively high temperatures, but the temperature for next month seems to be within the normal range. We need to consider whether high temperatures will affect takeout consumption. For example, high temperatures may cause people to prefer convenient food, or conversely, because of the hot weather, they may not want to go out to eat and choose takeout more. However, in the historical data, the takeout consumption corresponding to the high temperature month (such as 29.61) is not particularly high, so the impact of temperature may be limited.

[0093] In terms of precipitation data, there is a lot of fluctuation in the historical data, such as 230.02 mm of precipitation in the second month, and a prediction of 340.5 mm for next month. Region A has more precipitation during the rainy season, and heavy rain may affect travel, leading more people to choose takeout. However, we also need to see if the amount of precipitation is too large, causing delivery difficulties and thus affecting order volume. For example, when the monthly precipitation is 400.17 mm, the takeout consumption is 73.64 million, and the predicted precipitation for next month is higher, which may have a similar impact.

[0094] In terms of population inflow and outflow data, the population inflow and outflow fluctuate greatly. The net population inflow for next month is 22.03, which means net inflow. Population growth may bring more consumers, thereby increasing takeout demand. We need to consider the changes in population inflow and outflow to determine the population trend. For example, in the past few months, net inflow and outflow have alternated, but next month's net inflow is 22.03, which may mean population growth, thereby promoting takeout consumption.

[0095] Nightlight data reflects the activity of the city, higher light value may mean more business activities and consumption. In historical data, the light value is as high as 30.02, and the next month is predicted to be 30.134, close to the historical highest, which means that the city activity may be more active, which helps the growth of take-out consumption.

[0096] In addition, the knowledge base mentions that region A is a tourist city, and July is the tourist peak season, with more than 10% growth in off-site orders. Although it is currently July, the user's question time is July 24, 2025, and the prediction is for the next month (August). August may still be the tourist peak season, especially during the summer vacation, and the increase in the number of tourists will drive the demand for take-out. In particular, snacks such as rice noodles and lemon tea are popular and may increase the number of orders.

[0097] In summary, the precipitation is higher next month, which may lead to more people staying indoors ordering take-out; the temperature is moderate and will not have a significant impact on consumption; the net population migration may increase potential consumers; the night light activity is high, and the city's economic activity is thriving.

[0098] In addition, the tourists brought by the tourist peak season will also increase the number of take-out orders. Therefore, it is expected that the take-out consumption will increase next month, but considering the fluctuations in historical data, it may not reach the level of peak months, but should be higher than the average level;

[0099] Respondse: Predict the consumption for next month:<4684593>4>.”

[0100] In other embodiments, taking region B as an example, the take-out information corresponding to region B and the region-related information are input into the preset text template to obtain the prompt text, and the prompt text is input into the large language model to obtain the prediction result of the large language model. For ease of understanding and comparison, region B also provides historical take-out consumption, population inflow, population outflow, net population inflow, night light data, temperature, and precipitation in the last 10 months.

[0101] For example, the prompt text corresponding to region B is as follows:

[0102] "You are a take-out market researcher who is good at predicting the take-out consumption of a region by historical take-out consumption information and region-related information.

[0103] We are currently in region B.

[0104] The following historical take-out consumption information and region-related information corresponding to the historical time are provided:

[0105] Recent 10 months of historical food delivery consumption: {29320246, 39041471, 34398937, 36274955, 33149689, 34819844, 39267532, 39939732, 39258859, 33335718}.

[0106] Recent 10 months of population inflow: {150.1607, 270.7326, 107.8995, 194.1586, 249.841, 252.1406, 252.0153, 284.0833, 189.2216, 252.6352}.

[0107] Recent 10 months of population outflow: {355.8042, 137.185, 124.3513, 224.4826, 255.7715, 277.5458, 305.8628, 312.3825, 235.162, 270.4094}.

[0108] Recent 10 months of net population inflow: {-205.6435, 133.5476, -16.4518, -30.324, -5.9305, -25.4052, -53.8475, -28.3002, -45.9404, -17.7742}.

[0109] Recent 10 months of night light data: {32.804, 29.1672, 33.5449, 38.2444, 35.8513, 36.4955, 0.0491, 32.9987, 34.715, 37.4537}.

[0110] Recent 10 months of average monthly temperature: {14.9212, 12.1623, 20.2402, 21.9913, 23.5746, 27.6743, 29.6143, 28.9371, 28.2362, 24.387}.

[0111] Recent 10 months of average monthly precipitation: {14.0154, 164.2093, 115.86, 68.423, 393.7216, 354.5433, 247.2737, 358.8201, 111.0284, 38.8588}.

[0112] In addition, the time to be predicted is next month, and other related prediction information for next month is as follows:

[0113] Temperature: {28}.

[0114] Precipitation: {274.98}.

[0115] Population inflow: {340.13}.

[0116] Population outflow: {280.76}.

[0117] Nightlight data: {36.594}.

[0118] Net population inflow: {59.37}.

[0119] Now please note that your task is to predict the next month's food delivery consumption using the information provided above. Please provide the predicted value directly in the angle brackets. <value>."

[0120] Thus, the large language model can output the prediction result corresponding to region B based on the above prompt text, as follows:

[0121] "Okay, I need to predict the restaurant delivery consumption in region B next month. The user has provided historical data and prediction data for next month. First, I need to understand the relationship between these data and how various factors affect consumption.

[0122] First, looking at the historical data, the restaurant delivery consumption data fluctuates, such as 29.32 million in the first month, then jumps to 39.04 million, and then drops to 34.39 million, and then fluctuates. This may be related to temperature, precipitation, population migration, and nighttime light. I need to analyze the impact of each of these factors.

[0123] In terms of temperature data, historical temperature ranges from 14.9 to 29.6 degrees. Generally, when the temperature is low, people may prefer to order takeout instead of going out; however, if it is warm weather, it may have little effect. I need to look at the trends in the data. For example, when the temperature rises from 14.9 to 20.24, consumption increases from 29.32 million to 34.39 million, which may indicate that moderate temperatures increase consumption, but high temperatures may affect consumption? For example, in the 7th month, the temperature was 29.61, and the consumption was 39.26 million, but in the 8th month, the temperature was 28.94, and the consumption was 39.93 million, which actually increased slightly, so the temperature may not have a significant impact, or there may be other factors.

[0124] In terms of precipitation, historical precipitation data fluctuates greatly, such as 400.17 mm in the 5th month, when consumption was 33.14 million, down from 36.27 million in the previous month. It may be that heavy rain causes more people to order takeout, but the data shows that consumption decreases in months with high precipitation, which may be contradictory. I need to look at the situation in other months more carefully. For example, in the 6th month, the precipitation was 295.19, and the consumption was 34.81 million, which increased slightly from the previous month, which may indicate that the impact of precipitation is not significant, or there may be other factors at work. It is possible that increased precipitation will lead to increased consumption, but this needs to be combined with other factors.

[0125] In terms of population inflow and outflow, over the past 10 months, the net population inflow has been positive and negative. For example, in the first month, the net inflow was -205.64 (i.e., there was a population outflow), and the consumption was 29.32 million; in the second month, the net inflow was 133.55, and the consumption increased to 39.04 million. This may indicate that an increase in population migration promotes consumption, but I need to look at the situation in other months. For example, in the third month, the net outflow was 16.45, and the consumption was 34.39 million, which decreased from the previous month, which may be related to the increase in the number of people migrating out. The net inflow for the next month is 59.37, which means that the net population migration is high, which may promote consumption growth.

[0126] Night light data, usually related to urban activity, light intensity may represent economic activity, high consumption. For example, historical light data from 0.0491 to 38.24, the 7th month light is very low (0.2357), consumption 3926 million, which may be an outlier, or data error? Need to confirm, but there may be special circumstances such as power outages or other factors. But the data provided by the user may have been processed, so we need to accept this data. The next month's light is 36.594, higher than most historical months, which may mean an increase in activity and promote consumption.

[0127] Now integrate the next month's data: temperature 28 degrees, which is relatively high, but may have little impact on consumption; precipitation 274.98 mm, which is moderately high, may promote takeout consumption; 340.13 people move in, 280.76 people move out, net migration 59.37, which may increase the population base and drive consumption; night light 36.594, at a high level, may reflect economic activity and support high consumption.

[0128] Next, I need to analyze these factors together. Historical data shows that when net migration increases, consumption usually rises, such as the second month with net migration of 133.55, consumption increased from 2932 million to 3904 million. Similarly, the next month's net migration of 59.37 may drive consumption growth. At the same time, the high precipitation may increase takeout orders, but we need to refer to the historical relationship between precipitation and consumption. For example, consumption fluctuates in months with high precipitation, but we can look at temperature as well. If the temperature is moderate, the precipitation will promote consumption. The current temperature is 28 degrees, which may be warm but not extreme, combined with precipitation, people may prefer to order takeout at home.

[0129] In addition, as a first-tier city, the consumption level of Region B is high, and the high night light data may reflect commercial activity, so consumption may remain at a high level. Looking at historical data, consumption has fluctuated between 3300 million and 3993 million in the last few months, and the last month (10th month) consumption was 3333 million, which may be due to the precipitation of 38.8588 mm, temperature 24.38 degrees, and net migration of 17.77 million, resulting in lower consumption. The next month's net migration increases, and the precipitation increases, which may make consumption rebound.

[0130] Taking into account the higher net migration next month, higher precipitation, and higher light data, it is likely that consumption will increase significantly from last month, approaching historical highs. For example, the previous high was 39.93 million in the 8th month, and it is likely to reach around 40 million next month. The specific number needs to be adjusted, taking into account the weight of each factor, such as the impact of population migration, followed by precipitation, temperature, and light assistance. Therefore, the final prediction value is likely to be around 40 million.

[0131] Response: Predict the consumption next month:<4050000>0>.”.

[0132] Through the examples of Area A and Area B, it should be understood that historical food delivery consumption, temperature, precipitation, population inflow and outflow, and nighttime light data all affect the prediction of food delivery consumption. Therefore, by collecting information corresponding to historical time periods as data support for large language model prediction of food delivery consumption, the prediction accuracy of food delivery consumption can be improved. The food delivery consumption derivation process output by the large language model makes the impact of multiple heterogeneous factors, such as food delivery consumption data, social and economic data, and environmental data, on food delivery consumption visualized, thereby facilitating relevant technical personnel to adjust the input data and relevant personnel to understand how food delivery consumption is predicted, thereby improving the credibility of food delivery consumption.

[0133] It should be noted that the social and economic data and environmental data corresponding to Area A and Area B in the above embodiments are for illustration only. Those skilled in the art can also add actual social and economic data used, such as population density data, and environmental data used, such as monthly PM 2.5 emissions data, which are not limited by the present application.

[0134] In some embodiments, the method further comprises: obtaining training data, the training data comprising food delivery consumption data of a plurality of regions; performing data partitioning processing on the food delivery consumption data corresponding to each region to obtain a first food delivery consumption data set corresponding to each region and a second food delivery consumption data set corresponding to each region; training a large language model to be trained according to the first food delivery consumption data set and the second food delivery consumption data set corresponding to each region, to obtain the large language model.

[0135] For example, before using the large language model for prediction, the large language model needs to be trained to improve the prediction accuracy of food delivery consumption.

[0136] Exemplarily, the meal take-out consumption data corresponding to a plurality of regions is acquired as training data of the to-be-trained large language model, and the meal take-out consumption data corresponding to each region is subjected to data division processing to obtain a first meal take-out consumption data set corresponding to each region and a second meal take-out consumption data set corresponding to each region. Specifically, the first meal take-out consumption data set is taken as a support set, and the second meal take-out consumption data set is taken as a query set.

[0137] The to-be-trained large language model is trained by using the support set and the query set corresponding to at least one region to update the parameters of the to-be-trained large language model, so as to obtain the large language model.

[0138] In some embodiments, the training of the to-be-trained large language model according to the first meal take-out consumption data set and the second meal take-out consumption data set corresponding to each region to obtain the large language model comprises: in the execution of the nth training, at least one training region is randomly sampled from the plurality of regions; based on a preset optimizer, a random gradient descent process is performed on the first meal take-out consumption data set corresponding to the training region and the second meal take-out consumption data set corresponding to the training region to update the model parameters of the to-be-trained large language model; in the case that the model parameters meet the model training condition, it is determined that the training of the to-be-trained large language model is completed, and the large language model is obtained.

[0139] For example, for the nth model training, one region is randomly sampled from the plurality of regions as a training region i, and the support set and the query set corresponding to the training region i are obtained; the random gradient descent (SGD) process is performed on the support set and the query set corresponding to the training region i, and the model parameters of the to-be-trained large language model are updated according to the processing result, so as to complete the model training.

[0140] It should be understood that if it is determined that the model parameters meet the model training condition after the nth model training is completed, the model training is ended, and the large language model is obtained; if it is determined that the model parameters do not meet the model training condition after the nth model training is completed, the (n+1) th model training is performed, until the model parameters meet the model training condition, the model training is ended, and the corresponding large language model is output.

[0141] In some embodiments, if the model parameters converge, it is determined that the model parameters meet the model training condition; otherwise, it is determined that the model parameters do not meet the model training condition.

[0142] In some embodiments, the preset optimizer is used to perform random gradient descent processing on the first take-out consumption data set corresponding to the training area and the second take-out consumption data set corresponding to the training area to update the model parameters of the large language model to be trained, including: based on the model parameters obtained by the n-1 training, using the preset optimizer to perform S-step random gradient descent processing on the first take-out consumption data set corresponding to the training area, to obtain temporary parameters; based on the temporary parameters, using the preset optimizer to perform 1-step random gradient descent processing on the second take-out consumption data set corresponding to the training area, to obtain optimized parameters; updating the model parameters obtained by the n-1 training according to the optimized parameters and the model parameters obtained by the n-1 training, to obtain the model parameters corresponding to the n training.

[0143] For example, in the n training process, starting from the model parameters obtained by the n-1 training, the preset optimizer is used to perform S-step random gradient descent processing on the support set corresponding to the selected training area i, to obtain temporary parameters, wherein S is a positive integer greater than 1; then starting from the temporary parameters, the preset optimizer is used to perform 1-step random gradient descent processing on the query set corresponding to the training area i, to obtain optimized parameters; and then updating the model parameters obtained by the n-1 training according to the optimized parameters and the model parameters obtained by the n-1 training, to obtain the model parameters corresponding to the n training.

[0144] For example, the updating process of the model parameters can be as follows:

[0145]

[0146] wherein, θ im is used to indicate the model parameters updated by the n training, θ i(m-1) is used to indicate the model parameters obtained by the n-1 training, is used to indicate the optimized parameters in the n training process.

[0147] It should be noted that the above training process can improve the generalization ability of the large language model, thereby improving the prediction accuracy of the take-out consumption in different areas.

[0148] In step S104, based on the take-out consumption information prediction model, the take-out consumption information of the area to be predicted is determined according to the take-out consumption data, the take-out consumption preference vector and the take-out consumption.

[0149] For example, after obtaining the take-out consumption preference vector and the predicted take-out consumption, the take-out consumption of the area to be predicted is predicted by the take-out consumption information prediction model, to obtain the take-out consumption information of the area to be predicted.

[0150] For example, the target region includes a to-be-predicted region and at least one non-predicted region; when determining the catering take-out consumption preference vector, the catering take-out consumption data corresponding to the to-be-predicted region and / or the non-predicted region can be determined; and when determining the catering take-out consumption quantity, the data corresponding to the to-be-predicted region and / or the non-predicted region can also be predicted, which is not limited in the present application. The to-be-predicted region is adjacent to at least one non-predicted region, and / or the social and economic data and environmental data of the to-be-predicted region are similar to the social and economic data and environmental data of at least one non-predicted region.

[0151] In some embodiments, the catering take-out consumption information prediction model is used to determine the catering take-out consumption information of the to-be-predicted region according to the catering take-out consumption data, the catering take-out consumption preference vector and the catering take-out consumption quantity, including: determining the similarity between the catering take-out dish information of each catering take-out dish in the catering take-out consumption data and the catering take-out consumption preference vector based on the similarity calculation network of the catering take-out consumption information prediction model; and determining the catering take-out consumption information of the to-be-predicted region according to the similarity and the catering take-out consumption quantity based on the information prediction network of the catering take-out consumption information prediction model.

[0152] For example, the catering take-out consumption information prediction model includes a similarity calculation network and an information prediction network. After receiving the catering take-out consumption data, the similarity calculation network determines the similarity between the catering take-out consumption data corresponding to the non-predicted region and the catering take-out consumption preference vector, so that the information prediction network determines the catering take-out consumption information of the to-be-predicted region according to the similarity and the catering take-out consumption quantity.

[0153] For example, the catering take-out consumption data corresponding to the non-predicted region includes a plurality of catering take-out dishes and their corresponding catering take-out dish information, and the similarity calculation network is used to calculate the similarity between the catering take-out dish information corresponding to each catering take-out dish and the catering take-out consumption preference vector, and output the similarity corresponding to each catering take-out dish. For example, the catering take-out dish information in the catering take-out consumption data can exist in the form of a catering take-out consumption record sheet, each catering take-out consumption record sheet including at least one catering take-out dish and its corresponding catering take-out dish information, so that in the process of determining the similarity with the catering take-out consumption preference vector, the similarity between each catering take-out consumption record sheet and the catering take-out consumption preference vector can be determined according to the catering take-out dish information included in each catering take-out consumption record sheet; or the similarity between each catering take-out dish and the catering take-out consumption preference vector can be determined according to the catering take-out dish information corresponding to each catering take-out dish, which is not limited in the present application.

[0154] For example, the similarity between the food delivery dish information corresponding to the food delivery dish and the food delivery consumption preference vector is calculated based on a cosine similarity calculation rule.

[0155] In some embodiments, the determining of the food delivery consumption information of the to-be-predicted region according to the similarity and the food delivery consumption quantity comprises: determining a food delivery dish with a similarity greater than or equal to a preset similarity threshold as a to-be-selected food delivery dish; if the number of the to-be-selected food delivery dishes is greater than or equal to the food delivery consumption quantity, determining at least one food delivery dish from the to-be-selected food delivery dishes as a target food delivery dish; and determining the food delivery consumption information of the to-be-predicted region according to the target food delivery dish and the food delivery dish information corresponding to the target food delivery dish.

[0156] For example, a food delivery dish with a similarity greater than or equal to a preset similarity threshold is determined, and the food delivery dish is determined as a to-be-selected food delivery dish. The number of the to-be-selected food delivery dishes is compared with the food delivery consumption quantity. If the number of the to-be-selected food delivery dishes is greater than or equal to the food delivery consumption quantity, at least one food delivery dish from the to-be-selected food delivery dishes is determined as a target food delivery dish. The food delivery consumption information of the to-be-predicted region is determined according to the target food delivery dish and the food delivery dish information corresponding to the target food delivery dish.

[0157] If the number of the to-be-selected food delivery dishes is less than the food delivery consumption quantity, all the to-be-selected food delivery dishes are determined as target food delivery dishes.

[0158] For example, after the target food delivery dish is determined, the target food delivery dish and the food delivery dish information corresponding to the target food delivery dish are determined as the food delivery consumption information. It can be understood that the food delivery dish information of the target food delivery dish at least includes a food delivery dish name, a food delivery dish type and food delivery dish description information. Therefore, the obtained food delivery consumption information at least includes the food delivery dish name, the food delivery dish type and the food delivery dish description information of one food delivery dish.

[0159] In some embodiments, if the number of the to-be-selected food delivery dishes is greater than or equal to the food delivery consumption quantity, at least one food delivery dish from the to-be-selected food delivery dishes is determined as a target food delivery dish, which comprises: sorting a plurality of the to-be-selected food delivery dishes according to their respective similarities from large to small to obtain a to-be-selected food delivery dish sequence; extracting a corresponding number of to-be-selected food delivery dishes from the first to-be-selected food delivery dish in the to-be-selected food delivery dish sequence according to the food delivery consumption quantity; and determining the extracted to-be-selected food delivery dishes as target food delivery dishes.

[0160] For example, when the number of the candidate catering take-out dishes is greater than or equal to the number of the catering take-out dishes corresponding to the catering take-out consumption, the catering take-out dishes are sorted, and the corresponding candidate catering take-out dishes are extracted as the target catering take-out dishes according to the sorting result.

[0161] For example, the candidate catering take-out dishes include A, B, C, D, and E, and the number is 5, and the catering take-out consumption is 4, then the multiple catering take-out dishes are sorted according to the similarity corresponding to each catering take-out dish, and the candidate catering take-out dish sequence CADEB is obtained, and the candidate catering take-out dishes corresponding to the number of the catering take-out consumption are extracted from the first candidate catering take-out dish in the candidate catering take-out dish sequence, that is, the extraction starts from the candidate catering take-out dish C and ends at E, and the candidate catering take-out dishes C, A, D, and E are taken as the target catering take-out dishes, the corresponding catering take-out dish information is determined, and the catering take-out consumption information output is generated; it should be understood that the output catering take-out consumption information at least includes the catering take-out dish name, the catering take-out dish type, and the catering take-out dish description information corresponding to each of the catering take-out dishes CADE.

[0162] In other embodiments, the predicted catering take-out consumption information can also include the generation amount of catering take-out packaging waste, catering take-out nutrition quantification information, catering take-out carbon footprint, water footprint, and phosphorus footprint; thereby realizing the prediction of the generation amount of catering take-out packaging waste, catering take-out quantification information, catering take-out carbon footprint, water footprint, and phosphorus footprint corresponding to the catering take-out consumption; those skilled in the art can adjust the information contained in the input target historical period catering take-out consumption data according to the actual need of the predicted object, so as to adjust the specific prediction object in the predicted catering take-out consumption information, and the application does not limit the specific prediction object in the predicted catering take-out consumption information.

[0163] For example, when the generation amount of catering take-out packaging waste needs to be predicted, in an embodiment, the catering take-out consumption data of the target historical time period obtained in step S101 includes the generation amount of catering take-out packaging waste, and then the catering take-out consumption preference vector obtained in step S102 also carries the related information such as the generation amount of catering take-out packaging waste; thereby the predicted catering take-out consumption information obtained in step S104 also includes the prediction result of the generation amount of catering take-out packaging waste; achieving the prediction purpose of the generation amount of catering take-out packaging waste. It should be noted that the specific processing process of the generation amount of catering take-out packaging waste in the above steps can refer to the processing process of the carbon footprint in the catering take-out consumption data in the above embodiments, which will not be described here.

[0164] In other embodiments, since different types of catering take-out dishes use different probabilities of catering take-out packaging, the predicted types of catering take-out dishes and their corresponding quantities can be used to predict the amount of catering take-out packaging waste. The prediction process of the amount of catering take-out packaging waste is as follows:

[0165]

[0166] wherein PackageSum i class_consumption is used to indicate the amount of catering take-out packaging waste of packaging type i. j p is used to indicate the quantity (consumption) of catering take-out dish type j. ij n is the total number of packaging types, and m is the total number of catering take-out dish types.

[0167] For example, the predicted types of catering take-out dishes can be used to predict the amount of catering take-out packaging waste, so as to realize the prediction of the amount of catering take-out packaging waste and provide data guidance for the processing scheme of catering take-out packaging waste.

[0168] It should be noted that the embodiments provided in the present application can also be used to predict at least one of the carbon footprint, water footprint, phosphorus footprint and nutritional value of catering take-out consumption; the prediction process can refer to the prediction process of the amount of catering take-out packaging waste, which will not be repeated here.

[0169] The method for predicting catering take-out consumption information based on a large language model and text mining provided by the above embodiments realizes capturing of the semantics of unstructured text by performing text mining and vectorization processing on the unstructured text in the catering take-out consumption data in a target historical time period to obtain a catering take-out consumption preference vector, thereby providing a basis for fine analysis and processing of the catering take-out consumption data and predicting more refined catering take-out consumption information; and capturing the semantics of unstructured text can also determine the influence of the semantics of unstructured text in the catering take-out consumption data on the prediction of catering take-out consumption information, thereby improving the accuracy of the prediction of catering take-out consumption information; and the data analysis capability of the large language model on social and economic data and environmental data is used to predict the catering take-out consumption, thereby realizing the prediction of catering take-out consumption information using multi-dimensional data, avoiding the problem of single dimension of the source of the prediction data, and improving the accuracy of the prediction of catering take-out consumption information in a dynamic scenario and the scenario applicability of the method for predicting catering take-out consumption information based on a large language model and text mining, thereby providing reliable data reference for green transformation of catering take-out consumption and guidance of healthy diet for catering take-out.

[0170] It should be understood that the terms used in this specification of the application are only for the purpose of describing particular embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an" and "the" are intended to include plural forms unless the context clearly dictates otherwise.

[0171] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. It should be noted that in this text, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or system. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0172] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.< / value> < / value>

Claims

1. A method for predicting food delivery consumption information based on large language models and text mining, characterized in that, include: Acquire food delivery consumption data, socioeconomic data, and environmental data for the target area within the target historical time period; Based on a text mining model, the food delivery consumption data is vectorized to determine the food delivery consumption preference vector corresponding to the food delivery consumption data. Based on a large language model, the food delivery consumption in the target area is predicted according to the socio-economic data and the environmental data. Based on the food delivery consumption information prediction model, the food delivery consumption information of the region to be predicted is determined according to the food delivery consumption data, the food delivery consumption preference vector, and the food delivery consumption volume.

2. The method for predicting food delivery consumption information based on large language models and text mining as described in claim 1, characterized in that, The text mining model is used to vectorize the food delivery consumption data to determine the corresponding food delivery consumption preference vector, including: Based on the word segmentation network in the text mining model, the unstructured data in the food delivery consumption data is segmented and filtered to obtain the food delivery information text corresponding to each of the multiple food delivery dishes. Based on the vectorization network in the text mining model, the text of the catering takeaway dishes is vectorized to obtain the first dish attribute feature vector corresponding to each of the catering takeaway dishes; Based on the normalization network in the text mining model, the structured data in the food delivery consumption data is normalized to obtain the second dish attribute feature vector corresponding to each of the food delivery dishes. Based on the vector processing network in the aforementioned consumer preference prediction model, the food delivery consumer preference vector is determined according to the first dish attribute feature vector and the second dish attribute feature vector.

3. The method for predicting food delivery consumption information based on large language models and text mining as described in claim 2, characterized in that, The vector processing network based on the text mining model determines the food delivery consumption preference vector according to the first dish attribute feature vector and the second dish attribute feature vector, including: The first feature vector of each of the above-mentioned food delivery dishes is concatenated with the corresponding second feature vector of each of the above-mentioned food delivery dishes to obtain the target feature vector corresponding to each of the above-mentioned food delivery dishes. A mean vector is determined based on the target feature vectors of multiple food delivery dishes, and the mean vector is determined as the food delivery consumption preference vector.

4. The method for predicting food delivery consumption information based on large language models and text mining as described in claim 1, characterized in that, The method of predicting food delivery consumption in the target area based on a large language model, according to the socioeconomic data and the environmental data, includes: Based on the data extraction network of the large language model, first target data is extracted from the socio-economic data and second target data is extracted from the environmental data. The first target data and the second target data are then filled into a preset text template to obtain prompt text. Based on the prediction network of the large language model, the catering and takeout consumption in the target area is predicted according to the prompt text and historical catering and takeout consumption information.

5. The method for predicting food delivery consumption information based on large language models and text mining as described in any one of claims 1-4, characterized in that, The food delivery consumption information prediction model determines the food delivery consumption information of the region to be predicted based on the food delivery consumption data, the food delivery consumption preference vector, and the food delivery consumption volume, including: Based on the similarity calculation network of the food delivery consumption information prediction model, the similarity between the food delivery dish information of each food delivery dish in the food delivery consumption data and the food delivery consumption preference vector is determined. Based on the information prediction network of the food delivery consumption information prediction model, the food delivery consumption information of the region to be predicted is determined according to the similarity and the food delivery consumption volume.

6. The method for predicting food delivery consumption information based on large language models and text mining as described in claim 5, characterized in that, The step of determining the food delivery consumption information of the region to be predicted based on the similarity and the food delivery consumption volume includes: Food delivery dishes with a similarity greater than or equal to a preset similarity threshold are identified as candidate food delivery dishes; If the number of candidate food delivery dishes is greater than or equal to the food delivery consumption, then at least one food delivery dish is selected as the target food delivery dish from the candidate food delivery dishes. The catering consumption information of the area to be predicted is determined based on the target catering takeaway dishes and the catering takeaway dish information corresponding to the target catering takeaway dishes.

7. The method for predicting food delivery consumption information based on large language models and text mining as described in claim 6, characterized in that, If the number of candidate food delivery dishes is greater than or equal to the food delivery consumption, then at least one food delivery dish is selected from the candidate food delivery dishes as the target food delivery dish, including: The multiple candidate catering takeaway dishes are sorted from largest to smallest according to their respective similarity to obtain a candidate catering takeaway dish sequence; Based on the food delivery consumption, extract a corresponding number of candidate food delivery dishes starting from the first candidate food delivery dish in the candidate food delivery dish sequence; The extracted candidate catering takeout dishes are selected as target catering takeout dishes.

8. The method for predicting food delivery consumption information based on large language models and text mining as described in any one of claims 1-4, characterized in that, The method further includes: Acquire training data, which includes food delivery consumption data from multiple regions; The food delivery consumption data for each region is divided and processed to obtain the first food delivery consumption dataset and the second food delivery consumption dataset for each region. The large language model is trained based on the first and second food delivery consumption datasets corresponding to each region, and the large language model is obtained.

9. The method for predicting food delivery consumption information based on large language models and text mining as described in claim 8, characterized in that, The process of training the large language model to be trained based on the first and second food delivery consumption datasets corresponding to each region to obtain the large language model includes: In the nth training iteration, at least one training region is randomly sampled from among the multiple regions. Based on the preset optimizer, stochastic gradient descent is performed on the first food delivery consumption dataset and the second food delivery consumption dataset corresponding to the training region to update the model parameters of the large language model to be trained. If the model parameters meet the model training conditions, the large language model to be trained is determined to have completed training, and the large language model is obtained.

10. The method for predicting food delivery consumption information based on large language models and text mining as described in claim 9, characterized in that, The step of performing stochastic gradient descent on the first and second food delivery consumption datasets corresponding to the training region, based on a preset optimizer, to update the model parameters of the large language model to be trained, includes: Based on the model parameters obtained from the (n-1)th training, an S-step stochastic gradient descent process is performed on the first food delivery consumption dataset corresponding to the training region using a preset optimizer to obtain temporary parameters. Based on the temporary parameters, the preset optimizer is used to perform a one-step stochastic gradient descent process on the second food delivery consumption dataset corresponding to the training region to obtain the optimized parameters. The model parameters obtained from the (n-1)th training are updated based on the optimized parameters and the model parameters obtained from the (n-1)th training to obtain the model parameters corresponding to the nth training.