Data analysis method, electronic device, storage medium and program product
By acquiring relevant parameters of the target input information and merchant user information, and using multi-model parallel reasoning to generate structured analysis results, the problem of large language models lacking deep understanding in spatiotemporal data analysis is solved, and fast and accurate data analysis and decision support are achieved.
Patent Information
- Application Number
- CN202512060348.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Large language models lack a deep understanding of data indicators in spatiotemporal data analysis, making it difficult to accurately process and obtain analytical results.
By acquiring relevant parameters of the target input information, including target data metrics and information on merchants and users in time and space, the contribution of the target data distribution information is determined. Multiple models are used for parallel reasoning and similarity judgment to generate structured analysis results and suggestions.
Accurately mapping unstructured problems into spatiotemporal data distribution, quantifying the degree of impact, quickly locating the causes of problems, reducing the threshold and time cost of data analysis, and improving the accuracy of analysis and the robustness of decision-making.
Smart Images

Figure CN121833803A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and more particularly, to a data analysis method, an electronic device, a storage medium, and a program product. BACKGROUND
[0002] With the rapid development of artificial intelligence and natural language processing (NLP) technologies, large language models (LLMs) based on artificial intelligence and NLP technologies have been able to analyze input information and give answers. However, in professional data fields such as spatio-temporal data analysis that rely on accurate reasoning, LLMs often lack a deep understanding of data indicators, and their answers are superficial and fail to touch the root cause of the problem, making it difficult to accurately process input information and obtain analysis results.
[0003] Therefore, how to accurately analyze complex input problems has become a problem to be solved. SUMMARY
[0004] The present application provides a data analysis method, an electronic device, a storage medium, and a program product, which can accurately analyze complex input problems.
[0005] In a first aspect, a data analysis method is provided, comprising: obtaining target input information, determining related parameters, the related parameters including target data indicators; obtaining at least one target data distribution information related to the target data indicators, the target data distribution information including information of merchants and users in time and / or space; determining a target contribution degree corresponding to the target data distribution information, the target contribution degree being used to represent the influence degree of the target data distribution information on the target data indicators; and determining a target analysis result corresponding to the target input information based on the target contribution degree.
[0006] In the above technical solution, the related parameters including the target data indicators are determined from the obtained target input information, and at least one target data distribution information related to the target data indicators is obtained. Further, based on the target contribution degree corresponding to the determined target data distribution information, the target analysis result corresponding to the target input information is determined. The contribution degree can be accurately determined by referring to the information of merchants and users in time and space, and the unstructured natural language question can be accurately mapped to specific related parameters and spatio-temporal data distribution. The influence degree of each data distribution information on the business indicators is quantified, the complex input problem is accurately analyzed, the cause of the problem can be quickly and accurately located, and the threshold and time cost of data analysis are reduced.
[0007] In conjunction with the first aspect, in some possible implementations, the relevant parameters also include a target time period, and determining the target contribution corresponding to the target data distribution information, including: selecting first-stage data and second-stage data from the target data distribution information based on the target time period, where the first-stage data is used to characterize the data distribution at the beginning of the target time period, and the second-stage data is used to characterize the data distribution at the end of the target time period; and determining the target contribution based on the first-stage data and the second-stage data.
[0008] In the above technical solution, the target time period in the relevant parameters is used, and the first stage data at the start time and the second stage data at the end time of the target time period are selected from the target data distribution information. Then, the target contribution is determined by using the first stage data and the second stage data. This method can extract the data required to determine the target contribution through two important time nodes in the target time period, reducing the actual amount of data for data analysis and enabling accurate and rapid determination of the target contribution.
[0009] Combining the first aspect and the above-mentioned implementation methods, in some possible implementation methods, the target contribution is determined based on the data from the first stage and the data from the second stage, including: determining a first ratio indicator value based on the data from the first stage; determining a second ratio indicator value based on the data from the second stage; and determining the target contribution based on the first ratio indicator value and the second ratio indicator value.
[0010] In the above technical solution, the target contribution is determined by the first ratio indicator value determined by the first stage data and the second ratio indicator value determined by the second stage data. This can quantify the overall fluctuation of the target business indicator into a clear target contribution, thereby improving the accuracy of determining the target contribution.
[0011] Combining the first aspect and the above implementation methods, in some possible implementation methods, the target data distribution information has corresponding influencing factors. Based on the first ratio index value and the second ratio index value, the target contribution is determined, including: determining the target percentage contribution based on the second ratio index value, whereby the target percentage contribution is used to characterize the impact of changes in the weight of the influencing factors on the target contribution; determining the target ratio contribution based on the first ratio index value and the second ratio index value, whereby the target ratio contribution is used to characterize the impact of changes in the performance level of the influencing factors on the target contribution; and using the sum of the target percentage contribution and the target ratio contribution as the target contribution.
[0012] In the above technical solution, the target percentage contribution is determined by using the second ratio indicator value, and the target ratio contribution is determined by using the first ratio indicator value and the second ratio indicator value. Then, the sum of the target percentage contribution and the target ratio contribution is used as the target contribution degree. This can decompose the target contribution degree into two aspects: the weight change of the influencing factor and the change of the performance level of the influencing factor itself. This can clearly distinguish the reasons for the fluctuation of the target data indicator.
[0013] Combining the first aspect and the above implementation methods, in some possible implementation methods, the target data distribution information has a corresponding influence factor. The method further includes: determining the target influence factor from the influence factors corresponding to at least one target data distribution information, wherein the target contribution degree corresponding to the target influence factor is the maximum value among the target contributions; determining the target recommendation information based on the target influence factor; and outputting the target recommendation information.
[0014] In the above technical solution, the target influence factor is determined from the influence factors corresponding to at least one target data distribution information, and then the target suggestion information is determined and output using the target influence factor. Suggestions can be generated through the target influence factor with the largest contribution, thereby completing the automated closed loop from data analysis to decision-making suggestions, improving the guiding value of the target analysis results, and providing an actionable solution for the questioner.
[0015] Combining the first aspect and the above implementation methods, in some possible implementation methods, determining target suggestion information based on the target impact factor includes: obtaining at least one retrieval enhancement content from the target knowledge base based on the target impact factor; determining first suggestion information and second suggestion information based on the retrieval enhancement content, wherein the first suggestion information is generated by a first model and the second suggestion information is generated by a second model, and the first model and the second model are different; and determining target suggestion information based on the first suggestion information and the second suggestion information.
[0016] In the above technical solution, at least one retrieval enhancement content is obtained from the target knowledge base using the target impact factor. Then, based on the retrieval enhancement content, different models are used to determine the first and second suggestion information, and finally the target suggestion information to be output is obtained. The retrieval enhancement content provides a reliable data foundation for suggestion generation, avoiding random model generation. At the same time, parallel reasoning using two different architecture models can generate suggestion information with advantages from different perspectives, and finally obtain the target suggestion information, thus improving the accuracy and reliability of the target suggestion information.
[0017] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, determining target suggestion information based on first suggestion information and second suggestion information includes: determining the target similarity between the first suggestion information and the second suggestion information; if the target similarity is greater than or equal to a calibrated similarity threshold, fusing the first suggestion information and the second suggestion information to obtain the target suggestion information; if the target similarity is less than the calibrated similarity threshold, determining the first confidence level corresponding to the first suggestion information and the second confidence level corresponding to the second suggestion information based on the enhanced retrieval content, and determining the target suggestion information based on the first confidence level and the second confidence level.
[0018] In the above technical solution, the target similarity between the first suggestion information and the second suggestion information is determined. Then, based on the relationship between the target similarity and the calibration similarity threshold, the first suggestion information and the second suggestion information are fused, or the first suggestion information or the second suggestion information is selected as the target suggestion information based on the confidence level. When the suggestions of the two models are relatively consistent, reliable target suggestion information is generated quickly by fusion, which improves the response speed. When there is a discrepancy in the suggestions, a decision is made based on factual evidence, avoiding misjudgments caused by meaningless conflicts or illusions between models. This ensures that the output target suggestion information is always based on reliable information, which greatly improves the decision robustness of the system in complex scenarios.
[0019] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the method further includes: obtaining a prompt text template; filling the prompt text template with the target data distribution information to obtain the target prompt text; and, based on the target prompt text and the target model, performing the step of determining the target contribution corresponding to the target data distribution information.
[0020] In the above technical solution, the target data distribution information is filled into the prompt text template to obtain the target prompt text. Then, the target model is used to determine the target contribution based on the target prompt text. The target data distribution information is filled into the preset template in a structured format. The target model can obtain clear and unambiguous task instructions and context, which reduces the risk of the model producing illusions or calculation deviations due to ambiguous instructions and improves the accuracy of the target contribution.
[0021] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the prompt text template includes at least one of the following: a first prompt text, a second prompt text, and a third prompt text, wherein the first prompt text is used to prompt the role corresponding to the target model, the second prompt text is used to prompt the capabilities of the target model, and the third prompt text is used to prompt the constraints of the target model.
[0022] In the above technical solution, the prompt text template is broken down into prompt texts for roles, abilities and restrictions, which realizes refined and standardized control over the behavior of large models, ensures that the target model always maintains the correct identity positioning, and improves the accuracy of data analysis.
[0023] Secondly, a data analysis device is provided, comprising: The acquisition module is used to acquire target input information, determine relevant parameters, including target data indicators, and acquire at least one target data distribution information that is correlated with the target data indicators. The target data distribution information includes information about merchants and users in time and / or space. The determination module is used to determine the target contribution degree corresponding to the target data distribution information. The target contribution degree is used to characterize the degree of influence of the target data distribution information on the target data indicators. Based on the target contribution degree, the target analysis results corresponding to the target input information are determined.
[0024] In conjunction with the second aspect, in some possible implementations, the relevant parameters also include a target time period and a determination module, which is used to select first-stage data and second-stage data from the target data distribution information based on the target time period. The first-stage data is used to characterize the data distribution at the beginning of the target time period, and the second-stage data is used to characterize the data distribution at the end of the target time period. Based on the first-stage data and the second-stage data, the target contribution is determined.
[0025] Combining the second aspect and the above implementation methods, in some possible implementation methods, a determining module is used to determine the first ratio indicator value based on the first stage data; determine the second ratio indicator value based on the second stage data; and determine the target contribution based on the first ratio indicator value and the second ratio indicator value.
[0026] Combining the second aspect and the above implementation methods, in some possible implementation methods, the target data distribution information has corresponding influencing factors. The determination module is used to determine the target percentage contribution based on the second ratio index value. The target percentage contribution is used to characterize the impact of the weight change of the influencing factor on the target contribution. Based on the first ratio index value and the second ratio index value, the target ratio contribution is determined. The target ratio contribution is used to characterize the impact of the change of the performance level of the influencing factor on the target contribution. The sum of the target percentage contribution and the target ratio contribution is taken as the target contribution.
[0027] Combining the second aspect and the above implementation methods, in some possible implementation methods, the target data distribution information has a corresponding influence factor. The determining module is used to determine the target influence factor from the influence factors corresponding to at least one target data distribution information, and the target contribution degree corresponding to the target influence factor is the maximum value among the target contribution degrees. Based on the target influence factor, target recommendation information is determined. The device also includes an output module for outputting the target recommendation information.
[0028] Combining the second aspect and the above implementation methods, in some possible implementation methods, a determining module is used to obtain at least one retrieval enhancement content from the target knowledge base based on the target impact factor; based on the retrieval enhancement content, determine first suggestion information and second suggestion information, wherein the first suggestion information is generated by a first model and the second suggestion information is generated by a second model, and the first model and the second model are different; and based on the first suggestion information and the second suggestion information, determine target suggestion information.
[0029] Combining the second aspect and the above implementation methods, in some possible implementation methods, a determining module is used to determine the target similarity between the first suggestion information and the second suggestion information; when the target similarity is greater than or equal to the calibrated similarity threshold, the first suggestion information and the second suggestion information are fused to obtain the target suggestion information; when the target similarity is less than the calibrated similarity threshold, based on the retrieval enhancement content, the first confidence level corresponding to the first suggestion information and the second confidence level corresponding to the second suggestion information are determined, and the target suggestion information is determined based on the first confidence level and the second confidence level.
[0030] Combining the second aspect and the above implementation methods, in some possible implementation methods, the acquisition module is used to acquire the prompt text template; the device also includes a processing module used to fill the prompt text template with the target data distribution information to obtain the target prompt text; and through the target model, based on the target prompt text, the step of determining the target contribution corresponding to the target data distribution information is performed.
[0031] In combination with the second aspect and the above implementation methods, in some possible implementation methods, the prompt text template includes at least one of the following: a first prompt text, a second prompt text, and a third prompt text, wherein the first prompt text is used to prompt the role corresponding to the target model, the second prompt text is used to prompt the capabilities of the target model, and the third prompt text is used to prompt the limitations of the target model.
[0032] Fifthly, an electronic device is provided, including a memory and a processor, wherein the memory is used to store executable program code; and the processor is used to call and run the executable program code from the memory, causing the electronic device to perform the store opening assistance method in the first aspect or any possible implementation thereof.
[0033] In a sixth aspect, a computer-readable storage medium is provided that stores computer program code, which, when executed on a computer, causes the computer to perform the store opening assistance method described in the first aspect or any possible implementation thereof.
[0034] In a seventh aspect, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to execute the store opening assistance method in the first aspect or any possible implementation thereof.
[0035] The method provided in this application determines relevant parameters, including target data indicators, from the acquired target input information, then obtains at least one target data distribution information that is correlated with the target data indicators, and further determines the target analysis result corresponding to the target input information based on the target contribution degree corresponding to the determined target data distribution information. It can accurately determine the contribution degree by referring to the information of merchants and users in time and space, accurately map unstructured natural language problems into specific relevant parameters and spatiotemporal data distributions, quantify the impact of each data distribution information on business indicators, accurately perform data analysis on complex input problems, and quickly and accurately locate the cause of the problem, reducing the threshold and time cost of data analysis. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of an interactive data analysis process provided in an embodiment of this application; Figure 2 This is a schematic diagram of the architecture of a data analysis system provided in an embodiment of this application; Figure 3 This is an illustrative flowchart of a data analysis method provided in an embodiment of this application. Figure 1 ; Figure 4 This is an illustrative flowchart of a data analysis method provided in an embodiment of this application. Figure 2 ; Figure 5 This is a schematic diagram of the structure of a data analysis device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0037] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0038] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0039] Before introducing the solutions of the embodiments of this application, the technical terms that may be involved in the embodiments of this application will be explained first.
[0040] Natural Language Processing (NLP) is an interdisciplinary field combining computer science, artificial intelligence (AI), and linguistics. Its aim is to enable computers to understand, interpret, and generate human language. The core goal of NLP is to use algorithms and models to allow machines to process natural language, including text and speech, just like humans, and to achieve effective human-machine communication. Specifically, NLP mainly includes Natural Language Understanding (NLU) and Natural Language Generation (NLG). NLU aims to convert human language into a form that computers can process; NLG aims to convert computer-generated information into natural language for output or interaction.
[0041] LLMs, also known as Large Language Models, are artificial intelligence models designed to understand and generate human language. They are trained on massive amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. LLMs are characterized by their massive scale, containing billions of parameters that help them learn complex patterns in language data. These models are typically based on deep learning architectures, such as transformers, which contributes to their impressive performance on various NLP tasks.
[0042] Retrieval-augmented generation (RAG) is an NLP technique that combines the capabilities of information retrieval and generative models to generate more accurate and informative answers or text. This approach is particularly effective in handling complex question-answering systems and knowledge-enhanced dialogue systems.
[0043] Before introducing the solutions of the embodiments of this application, we will first introduce the application scenarios of the embodiments of this application.
[0044] With the rapid development of artificial intelligence and natural language processing (NLP) technologies, LLM (Limited Learning Modeling), based on these technologies, is now capable of analyzing input information and providing answers. However, in specialized data fields such as spatiotemporal data analysis, which rely on precise reasoning, situations arise where there is a large amount of report and indicator data. LLM often lacks a deep understanding of these data indicators, resulting in answers that are merely superficial descriptions of phenomena, failing to address the root causes of the problems, and thus unable to accurately process the input information and obtain analytical results.
[0045] Therefore, how to accurately perform data analysis on complex input problems has become an urgent problem to be solved.
[0046] It should be noted that, in the case of user information involved in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0047] The following is combined with Figures 1-2 The data analysis method provided in the embodiments of this application is described with examples.
[0048] Figure 1 This is a schematic diagram of an interactive data analysis process provided in an embodiment of this application.
[0049] For example, the data analysis method provided in this application supports warehouse / store scenarios and experience scenarios. The warehouse / store scenario focuses on how supply-side resources are allocated to service points on the demand side. The experience scenario focuses on the quality and certainty of the entire process from user order placement to fulfillment. After obtaining the target input information, relevant parameters including target data indicators, target time periods, and target scenarios are extracted from the target input information. Then, target data distribution information that is correlated with the target data indicators is obtained from the data distribution pool. The data distribution pool stores various data distribution information, including: time distribution, distance distribution, city distribution, time-distance distribution, time-city distribution, and time-industry distribution. The data distribution information in the data distribution pool can be generated from various data sources, such as MySQL sources, Hologres sources, and Open Data Processing Service (ODPS) sources. Furthermore, the target model calls relevant computing engines to determine the target contribution corresponding to the target data distribution information. The target model includes a first model, a second model, and a third model with different functions. The relevant computing engines include data distribution calculation, indicator calculation, and contribution calculation. The target model can calculate contribution, determine analysis results, and generate recommendations.
[0050] Figure 2 This is a schematic diagram of the architecture of a data analysis system provided in an embodiment of this application.
[0051] For example, a data analysis system is divided into a configuration side, a system capability side, and an operation side. On the configuration side, the data analysis system is configured in the following order: defining scenarios, entering data metrics, defining attribution rules, defining data sources, and configuring the target model. On the system capability side, it is divided into a presentation layer, engine capabilities, and a basic model. In the presentation layer, operators can view scenario views on the system's homepage, while administrators can perform operations such as scenario management, data metric management, attribution rule configuration, model data source configuration, model prompt word configuration, and process configuration. The engine capabilities include a model execution engine, model dialogue memory, attribution, optimization, process execution, and intent recognition. The basic model includes basic capabilities such as model scenarios, data metrics, attribution rules, contribution results, data sources, prompt words, operation processes, and reports. On the operation side, data analysis is performed according to the following process: accessing the platform homepage, selecting a scenario, entering target input information, calculating contribution, generating a report, and downloading the report. The report includes the target analysis results and target suggestion information corresponding to the target input information.
[0052] The above embodiments combined Figures 1-2 This paper introduces the execution logic and specific results of data analysis methods from the perspectives of interaction and architecture. The following will combine... Figure 3This section introduces the underlying implementation process upon which this method depends.
[0053] Figure 3 This is an illustrative flowchart of a data analysis method provided in an embodiment of this application. Figure 1 It should be understood that this method 300 can be applied to electronic devices such as servers. The server can be a distributed server cluster composed of multiple servers, or it can be implemented as a single server. The server can also be a server in a distributed system, or a server combined with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. This application does not limit the type of electronic device. For example, such as... Figure 3 As shown, the data analysis method 300 includes the following steps.
[0054] 301. Obtain target input information and determine relevant parameters, including target data indicators.
[0055] The target input information is the information used to trigger data analysis. In some embodiments, the target input information can be of any suitable type, such as text, audio, etc. In response to receiving a target operation for input information at the target user interface (UI), the target input information is acquired and relevant parameters are determined. The target user interface is the user interface in the target application equipped with the data analysis system. The relevant parameters are information extracted from the target input information and can provide auxiliary functions in data analysis. In some embodiments, the relevant parameters may include, but are not limited to, at least one of the following: target data metrics, target time period, target scenario, etc. The target data metrics are the data metrics to be analyzed. In some embodiments, the target data metrics may include, but are not limited to, one of the following: fulfillment rate, on-time rate, order completion rate, positive review rate, etc. The target time period is the time span that the data analysis needs to analyze. The target scenario is the specific application scenario corresponding to the data analysis. In some embodiments, the target scenario may include, but is not limited to, one of the following: warehouse / store scenario, experience scenario, etc.
[0056] 302. Obtain at least one target data distribution information that is related to the target data indicator. The target data distribution information includes information about merchants and users in time and / or space.
[0057] Data distribution describes the frequency and distribution of each value in the dataset, providing a structured description of the data's characteristics, trends, and variability. Target data distribution information refers to data distribution information that is correlated with a target data indicator. In some embodiments, the number of target data distribution information is at least one, and at least one target data distribution information can be obtained from the data distribution information of a preset database based on the target data indicator.
[0058] It should be noted that in the food delivery industry, when conducting data analysis on target data metrics, it is necessary to refer to at least one target data distribution information. Different target data distribution information may describe different specific information. At the same time, different target data distribution information can be set for different industries. Therefore, target data distribution information includes information on merchants and users in different industries in terms of time and / or space. By referring to information on merchants and users in different dimensions of time and space, a more comprehensive analysis of target data metrics can be conducted.
[0059] 303. Determine the target contribution degree corresponding to the target data distribution information. The target contribution degree is used to characterize the degree of influence of the target data distribution information on the target data indicators.
[0060] The target contribution rate is used to characterize the degree of influence of the target data distribution information on the target data indicator. In some embodiments, the target contribution rate can be expressed as a percentage, and can be any suitable size, such as 19.85%, 30.27%, etc. There is at least one target contribution rate. For a single target data distribution information, multiple dimensions of information can be stored; therefore, a single target data distribution information can correspond to at least one target contribution rate. In some embodiments, a larger target contribution rate indicates a greater degree of influence of the target data distribution information on the target data indicator.
[0061] 304. Based on the target contribution, determine the target analysis results corresponding to the target input information.
[0062] The target analysis results are used to characterize the reasons for fluctuations in the target data indicators in the target input information. In some embodiments, different target input information corresponds to different target analysis results. The target analysis result corresponding to the target input information can be determined by selecting the target data distribution information with the highest target contribution based on the target contribution information corresponding to the target data distribution information.
[0063] The method provided in this application determines relevant parameters, including target data indicators, from the acquired target input information, then obtains at least one target data distribution information that is correlated with the target data indicators, and further determines the target analysis result corresponding to the target input information based on the target contribution degree corresponding to the determined target data distribution information. It can accurately determine the contribution degree by referring to the information of merchants and users in time and space, accurately map unstructured natural language problems into specific relevant parameters and spatiotemporal data distributions, quantify the impact of each data distribution information on business indicators, accurately perform data analysis on complex input problems, and quickly and accurately locate the cause of the problem, reducing the threshold and time cost of data analysis.
[0064] It should be noted that steps 301-304 above are a simplified explanation of the data analysis method provided in the embodiments of this application. The data analysis method provided in the embodiments of this application will be explained in more detail below with some examples. See [link to relevant documentation]. Figure 4 , Figure 4 This is an illustrative flowchart of a data analysis method provided in an embodiment of this application. Figure 2 .
[0065] It should be understood that this method 400 can be applied to electronic devices such as servers, and the embodiments of this application do not limit the type of electronic device. For example, such as... Figure 4 As shown, the data analysis method 400 includes the following steps.
[0066] 401. Obtain target input information and determine relevant parameters, including target data indicators.
[0067] The target input information is the information used to trigger data analysis. In some embodiments, the target input information can be of any suitable type, such as text, audio, etc. In response to receiving a target operation for input information at the target user interface, the target input information is acquired and relevant parameters are determined. The target user interface is the user interface of the target application equipped with the data analysis system. The relevant parameters are information extracted from the target input information and can provide auxiliary functions in data analysis. In some embodiments, the relevant parameters may include, but are not limited to, at least one of the following: target data metrics, target time period, target scenario, etc. The target data metrics are the data metrics to be analyzed. In some embodiments, the target data metrics may include, but are not limited to, one of the following: fulfillment rate, on-time rate, order completion rate, positive review rate, etc. The target time period is the time span that the data analysis needs to analyze. In some embodiments, the target time period includes a start time and an end time. The target scenario is the specific application scenario corresponding to the data analysis. In some embodiments, the target scenario may include, but is not limited to, one of the following: warehouse scenario, experience scenario, etc.
[0068] In some embodiments, NLP techniques can be used to identify and determine relevant parameters from the acquired target input information. Specifically, Named Entity Recognition (NER) can be used to extract explicit entities (e.g., order completion rate, this month, warehouse, etc.) from the target input information. Then, semantic analysis is used to further classify the extracted explicit entities into corresponding relevant parameters. At the same time, domain rules can be invoked to verify the rationality and completeness of the parameter combinations, ultimately obtaining the relevant parameters.
[0069] 402. Obtain at least one target data distribution information that is related to the target data indicator. The target data distribution information includes information about merchants and users in time and / or space.
[0070] Data distribution describes the frequency and distribution of each value in the dataset, providing a structured description of the data's characteristics, trends, and variability. Target data distribution information is data distribution information that is correlated with the target data indicator. In some embodiments, the number of target data distribution information is at least one, and at least one target data distribution information can be obtained from the data distribution information of a preset database based on the target data indicator. The database data distribution information can be generated from various data sources, such as MySQL sources, Hologres sources, ODPS sources, etc.
[0071] For example, the data distribution information included in the preset database can be as follows: 1. Distance distribution of merchants and users; 2. Time distribution of merchants and users; 3. Distance-time distribution of merchants and users; 4. Changes in the distance distribution of merchants and users; 5. Changes in the time distribution of merchants and users; 6. Changes in the time-distance distribution of merchants and users; 7. Distribution of merchants by city; 8. Which delivery channels merchants have chosen.
[0072] Different data distributions can be set for different industries, including but not limited to: food, convenience stores, and pharmacies. Different data distributions can represent different content. For example, changes in the distance distribution between merchants and users can indicate whether users are getting farther or closer; changes in the time distribution of merchants and users can indicate user consumption at different times; and the distribution of merchants by city can indicate the density of merchants in different cities, indirectly reflecting the supply density.
[0073] In some embodiments, when retrieving at least one target data distribution information that is correlated with the target data indicator from the database, the degree of correlation between the data distribution information in the database and the target data indicator can be determined. This process is repeated for each data distribution information in the database, and then at least one data distribution information whose correlation ranking is in the top N or whose correlation is greater than a preset level is selected as the target data distribution information. Here, N can be any suitable positive integer, such as 5, 4, etc.
[0074] 403. Based on the target time period, select the first stage data and the second stage data from the target data distribution information. The first stage data is used to characterize the data distribution at the beginning of the target time period, and the second stage data is used to characterize the data distribution at the end of the target time period.
[0075] The target time period is the time span that the data analysis needs to cover. The first-stage data refers to the data distribution at the beginning of the target time period. The second-stage data refers to the data distribution at the end of the target time period. In some embodiments, the first-stage data and the second-stage data are different. The data distribution information has corresponding timestamps. After determining the target time period from the target input information, the first-stage data and the second-stage data can be selected from the target data distribution information based on the timestamps of the start and end times of the target time period.
[0076] In one possible implementation, a prompt text template is obtained. The target data distribution information is then filled into the prompt text template to obtain the target prompt text. Based on the target prompt text and using the target model, the step of determining the target contribution corresponding to the target data distribution information is performed.
[0077] The prompt text template is a keyword or phrase used to guide the target model in performing a specific task or dialogue. It helps the target model understand the current settings and more accurately parse and process the input data. In some embodiments, the methods for obtaining the prompt text template may include, but are not limited to, obtaining the prompt text template from the cloud or a server, or obtaining the prompt text template through an Application Programming Interface (API). An API is a medium for communication and data exchange between software systems. Requests can be sent to external services through an API, and the returned data constitutes the prompt text template.
[0078] The target suggestion text is data obtained by filling the suggestion text template with target data distribution information. In some embodiments, the suggestion text template has reserved positions for filling target data distribution information, which can be filled into the corresponding positions to obtain the target suggestion text. The target model refers to a machine learning model with large-scale parameters and complex computational structures, capable of processing massive amounts of data and completing various complex tasks. In some embodiments, the target model is an LLM (Limited Learning Model). The target model can be integrated into an application for data analysis. The target model includes at least one different sub-model. These sub-models may include, but are not limited to, a first model, a second model, and a third model. The first and second models can be used to generate target suggestion information. The third model can, based on the target suggestion text, perform the step of determining the target contribution corresponding to the target data distribution information.
[0079] In this implementation, the target data distribution information is filled into the prompt text template to obtain the target prompt text. Then, the target model is used to determine the target contribution based on the target prompt text. The target data distribution information is filled into the preset template in a structured format. The target model can obtain clear and unambiguous task instructions and context, which reduces the risk of the model producing illusions or calculation deviations due to ambiguous instructions and improves the accuracy of the target contribution.
[0080] In one possible implementation, the prompt text template includes at least one of the following: a first prompt text, a second prompt text, and a third prompt text, wherein the first prompt text is used to prompt the role corresponding to the target model, the second prompt text is used to prompt the capabilities of the target model, and the third prompt text is used to prompt the limitations of the target model.
[0081] The prompt text template may include, but is not limited to, at least one of the following: a first prompt text, a second prompt text, and a third prompt text. In some embodiments, the first prompt text, the second prompt text, and the third prompt text have different functions. Specifically, the first prompt text is used to indicate the role corresponding to the target model, the second prompt text is used to indicate the capabilities of the target model, and the third prompt text is used to indicate the limitations of the target model.
[0082] For example, the initial prompt text could be: "You are a near-field fulfillment timeliness data analysis intelligent assistant, dedicated to providing users with in-depth data insights and data-driven decision support services. The core idea is to iteratively design and innovate products based on data insights. Avoid code, prohibit associations, and only use the provided data for analysis. Do not draw conclusions from data of unknown origin. 'data' refers to the source of your analysis data. Input data: ${data}."
[0083] The second prompt text could be: 1. Data Analysis: You need to conduct a comprehensive analysis of the input data, using statistical methods and machine learning algorithms to identify key influencing factors, such as trends, patterns, correlations, or outliers, to help users understand the hidden information behind the data. Consider factors like weather, holidays, or major promotional events. 2. Metric Monitoring: Monitor key business metrics in real time to ensure data quality and provide timely warnings of potential problems or anomalies. 3. Visualization: Present the analysis results in charts, dashboards, or other visual formats to make complex data easier to understand and interpret. 4. Suggested Action Directions: Based on the data analysis results, provide targeted suggestions, including improvement measures, optimization strategies, or exploration of new opportunities, providing a scientific basis for users to develop action plans.
[0084] The third prompt text could be: 1. You rely on the quality of the data provided by the user. If the data is missing, incorrect, or biased, it will affect the accuracy of the analysis results. 2. You need the user to clearly define the target data and target scenario in order to conduct more accurate data analysis and interpretation of results. 3. Although you can provide data-driven insights and suggestions, when it comes to specific implementation, the user still needs to make the final decision based on their own judgment and experience. 4. When processing large-scale data or executing complex algorithms, it may take some time to complete the analysis task. Please be patient. In this implementation method, the prompt text template is broken down into prompt texts categorized by role, ability, and constraints. This achieves refined and standardized control over the behavior of large models, ensuring that the target model always maintains the correct identity and improving the accuracy of data analysis.
[0085] 404. Based on the data from the first and second phases, the target contribution level is determined.
[0086] The first-stage data refers to the data distribution at the beginning of the target time period. The second-stage data refers to the data distribution at the end of the target time period. The target contribution rate characterizes the degree of influence of the target data distribution information on the target data indicator. In some embodiments, the target contribution rate can be expressed as a percentage, and can be any suitable size, such as 19.85%, 30.27%, etc. There is at least one target contribution rate. For a single target data distribution information, multiple dimensions of information can be stored; therefore, a single target data distribution information can correspond to at least one target contribution rate.
[0087] In some embodiments, a larger target contribution indicates a greater influence of the target data distribution information on the target data indicator. Based on the first-stage data and the second-stage data, a contribution decomposition algorithm can be used to determine the degree of interference of the target data indicator with different target data distribution information within the target time period, thereby obtaining at least one target contribution.
[0088] In this implementation, the target time period in the relevant parameters is used, and the first stage data at the start time and the second stage data at the end time of the target time period are selected from the target data distribution information. Then, the target contribution is determined by using the first stage data and the second stage data. The data required to determine the target contribution can be extracted through two important time nodes in the target time period, which reduces the actual amount of data for data analysis and can accurately and quickly determine the target contribution.
[0089] In one possible implementation, a first ratio indicator value is determined based on data from the first phase. A second ratio indicator value is determined based on data from the second phase. The target contribution is determined based on the first and second ratio indicator values.
[0090] In this context, ratio indicators refer to core performance data used to measure data metrics within a specific analytical dimension, typically presented as percentages in the numerator and denominator. The first ratio indicator value corresponds to the data from the first stage. This first ratio indicator value can be any suitable size, such as 52.1%, 23.35%, etc. The second ratio indicator value corresponds to the data from the second stage. This second ratio indicator value can be any suitable size, such as 26.8%, 18.62%, etc. In some embodiments, the first and second ratio indicator values are different. In some embodiments, the difference between the first and second ratio indicator values can be determined as the target contribution.
[0091] In some embodiments, when determining the ratio index value (including the first ratio index value and the second ratio index value) using stage data (including first stage data and second stage data), sub-ratio index values can be calculated for data of different dimensions included in the stage data. Then, the sub-ratio index values are corrected by the weights corresponding to the data of different dimensions. Finally, all corrected sub-ratio index values are added together to obtain the ratio index value. The sum of the weights corresponding to the data of different dimensions is a fixed value (e.g., 1).
[0092] In this implementation, the target contribution is determined by the first ratio indicator value determined by the first stage data and the second ratio indicator value determined by the second stage data. This can quantify the overall fluctuation of the target business indicator into a clear target contribution, thus improving the accuracy of determining the target contribution.
[0093] In one possible implementation, the target data distribution information has corresponding influencing factors. The target percentage contribution is determined based on a second ratio indicator value, which characterizes the impact of changes in the weights of the influencing factors on the target contribution. The target ratio contribution is determined based on a first ratio indicator value and a second ratio indicator value, which characterizes the impact of changes in the performance level of the influencing factors on the target contribution. The sum of the target percentage contribution and the target ratio contribution is taken as the target contribution.
[0094] The target data distribution information has corresponding influence factors. In some embodiments, a target data distribution information has at least one corresponding influence factor (i.e., data from different dimensions). The second ratio indicator value is the ratio indicator value corresponding to the second-stage data. The second ratio indicator value can be any suitable size, such as 26.8%, 18.62%, etc. The target percentage contribution is used to characterize the impact of changes in the weight of the influence factor on the target contribution. The target percentage contribution can be any suitable size, such as 12.8%, 9.26%, etc. The first ratio indicator value is the ratio indicator value corresponding to the first-stage data. The first ratio indicator value can be any suitable size, such as 52.1%, 23.35%, etc. The target ratio contribution is used to characterize the impact of changes in the performance level of the influence factor itself on the target contribution. The target ratio contribution can be any suitable size, such as 24.53%, 6.25%, etc.
[0095] In some embodiments, the first ratio index value and the second ratio index value can be expanded, and then a factor with an actual value of zero can be introduced. Based on the expanded first ratio index value and the expanded second ratio index value, the target ratio contribution can be determined, and based on the expanded second ratio index value, the target percentage contribution can be determined.
[0096] It's important to note that in determining the target contribution, the target percentage contribution and the target ratio contribution represent two reasons driving fluctuations in the data indicators. Only by calculating the target percentage contribution and the target ratio contribution separately can we trace the specific root causes from the overall result, thereby generating targeted recommendations. Specifically, the target percentage contribution characterizes the impact of changes in the weights of influencing factors on the target contribution, reflecting the influence of changes in structure or resource distribution. The target ratio contribution characterizes the impact of changes in the performance level of the influencing factors themselves on the target contribution, revealing the impact of changes in their own efficiency levels on the target contribution.
[0097] In some embodiments, the target contribution can be determined using the following formula. :
[0098]
[0099]
[0100]
[0101]
[0102]
[0103] In the above formula, This is the ratio index value corresponding to the i-th dimension in the first ratio index value. The weight is the proportion of the denominator of the ratio index value corresponding to the i-th dimension in the first ratio index value to the total denominator. This refers to the ratio index value corresponding to the i-th dimension in the second ratio index value. This represents the proportion of the denominator of the ratio index value corresponding to the i-th dimension in the second ratio index value to the total denominator. ,therefore, T is an arbitrary constant. Contribute to the target percentage Contribute to the target ratio. The first ratio indicator value, This is the second ratio indicator value.
[0104] In this implementation, the target percentage contribution is determined by using the second ratio indicator value, and the target ratio contribution is determined by using the first ratio indicator value and the second ratio indicator value. Then, the sum of the target percentage contribution and the target ratio contribution is used as the target contribution degree. This can decompose the target contribution degree into two aspects: the weight change of the influencing factor and the change of the performance level of the influencing factor itself. This can clearly distinguish the reasons for the fluctuation of the target data indicator.
[0105] 405. Based on the target contribution, determine the target analysis results corresponding to the target input information.
[0106] The target analysis results are used to characterize the reasons for fluctuations in the target data indicators in the target input information. In some embodiments, different target input information corresponds to different target analysis results. The target analysis result corresponding to the target input information can be determined by selecting the target data distribution information with the highest target contribution based on the target contribution information corresponding to the target data distribution information.
[0107] In one possible implementation, the target data distribution information has corresponding influence factors. A target influence factor is determined from the influence factors corresponding to at least one target data distribution information, where the target contribution of the target influence factor is the maximum value among the target contributions. Based on the target influence factor, target recommendation information is determined. The target recommendation information is then output.
[0108] The target data distribution information has corresponding influencing factors. In some embodiments, a target data distribution information has at least one corresponding influencing factor. Influencing factors may include, but are not limited to, at least one of the following: delivery method, distance range, city, time period, pressure level, weather, product category, etc. The pressure level is used to characterize the order pressure on merchants. The pressure level can be represented by a specific numerical value; for example, 1 can represent normal pressure, 2 can represent low to medium pressure, etc.
[0109] The target impact factor is an impact factor determined from at least one impact factor corresponding to the target data distribution information, and the target contribution corresponding to the target impact factor is the maximum value among the target contributions. Target suggestion information is recommendation information proposed for the target impact factor. Target suggestion information can be of any suitable type, such as text, audio, or images. In some embodiments, target suggestion information can be generated using a first model and a second model in the target model. The process of determining the target contribution can be simultaneously output to the target user interface.
[0110] In some embodiments, a target display format for the target suggestion information is determined. The target suggestion information is adjusted based on the target display format to output the adjusted target suggestion information. Here, the target display format is the format in which the target suggestion information is displayed. The target display format may include, but is not limited to, one of the following: Hyper Text Markup Language (HTML), Rich Text Format (RTF), Graphical User Interface (GUI) format, audio format, video format, etc. HTML is a markup language that includes a series of tags that unify the document format on the web, connecting scattered Internet resources into a logical whole. RTF is a text and graphic document format that is easy to view on different devices and systems. GUI is a computer operating user interface that displays information graphically.
[0111] In this implementation, target influence factors are determined from the influence factors corresponding to at least one target data distribution information. Then, target recommendation information is determined and output using the target influence factors. Recommendations can be generated through the target influence factors with the greatest contribution, thus completing an automated closed loop from data analysis to decision-making recommendations. This enhances the guiding value of the target analysis results and provides an actionable solution for the questioner.
[0112] In one possible implementation, at least one retrieval enhancement content is obtained from the target knowledge base based on the target impact factor. Based on the retrieval enhancement content, first suggestion information and second suggestion information are determined. The first suggestion information is generated using a first model, and the second suggestion information is generated using a second model, which are different from each other. Target suggestion information is then determined based on the first and second suggestion information.
[0113] To provide a clearer explanation of the above implementation methods, the following section will describe the process of determining the query vector corresponding to the target question and determining the enhanced retrieval content in two parts.
[0114] Part 1: Based on the target impact factor, obtain at least one retrieval enhancement content from the target knowledge base.
[0115] The target impact factor is determined from at least one impact factor corresponding to the target data distribution information, and the target contribution of the target impact factor is the maximum value among the target contributions. The target knowledge base is a pre-defined knowledge base that stores relevant enhanced content for different scenarios in the food delivery field. The retrieved enhanced content is obtained by retrieving the target knowledge base using the target impact factor. In some embodiments, one target impact factor corresponds to at least one retrieved enhanced content.
[0116] In some embodiments, multiple augmented contents in the target knowledge base are stored in vector form. After determining the target impact factor, the target vector corresponding to the target impact factor is first determined. Then, the target vector is matched with multiple augmented contents in the target knowledge base to obtain at least one retrieval augmented content. When determining the target vector corresponding to the target impact factor, the target impact factor can be directly encoded to obtain the target vector, or the keywords of the target impact factor can be extracted first, and then the keywords of the target impact factor can be encoded to obtain the target vector corresponding to the target impact factor.
[0117] In some embodiments, the target impact factor is encoded into a query vector through vector transformation. Vector transformation can be implemented using embedding, a technique that transforms high-dimensional sparse feature vectors into low-dimensional dense word vectors. Embedding can be used for dimensionality reduction, dimensionality increase, similarity calculation, etc. Through embedding, the target impact factor can be encoded into a target vector. In some embodiments, vector transformation can be implemented using residual networks, a type of network capable of classification and object recognition. Through residual networks, the target impact factor can be encoded into a target vector.
[0118] In some embodiments, at least one target vector is obtained by using a target vector and multiple augmentations in a target knowledge base through vector search or keyword matching. Vector search is a search technique used to find similar items or data points in a large set. Keyword matching is a common text processing task, typically used for searching, replacing, and validating text. When processing a target vector and multiple augmentations in a target knowledge base using vector search, the similarity between the target vector and the multiple augmentations in the target knowledge base can be determined, and then the retrieved augmentations can be determined based on the similarity.
[0119] In some embodiments, enhanced content with a similarity greater than or equal to a target similarity threshold is used as retrieval enhanced content to obtain at least one retrieval enhanced content. The target similarity threshold can be any suitable value, such as 85%, 90%, etc. In some embodiments, the similarities are sorted, and the enhanced content with the top N similarities is used as retrieval enhanced content to obtain at least one retrieval enhanced content. N can be any suitable positive integer, such as 5, 3, etc.
[0120] In some embodiments, similarity is determined using a similarity calculation algorithm. This similarity calculation algorithm may include, but is not limited to, cosine similarity (CS) and Euclidean distance (ED). For example, the cosine similarity algorithm calculates the cosine of the angle between the target vector and multiple augmented content items in the target knowledge base, and uses this cosine value as the similarity. As another example, the Euclidean distance algorithm calculates the distance between the midpoints of the target vector and multiple augmented content items in the target knowledge base, and uses this distance value as the similarity.
[0121] Part Two: Based on the enhanced search content, determine the first and second recommended information.
[0122] In this system, the first suggestion information is generated using a first model, and the second suggestion information is generated using a second model. Both the first and second models can be used to generate target suggestion information. The first and second models are different. In some embodiments, at least one retrieval enhancement content and a target impact factor are concatenated. Then, the concatenated retrieval enhancement content and target impact factor are converted into vector form and input into the first and second models respectively. The first and second models first encode the data and then perform multiple rounds of iterative decoding to output the first and second suggestion information, thereby determining the first and second suggestion information based on the retrieval enhancement content.
[0123] In some embodiments, at least one retrieval enhancement content and a target impact factor are concatenated using methods such as point-by-point addition or vector concatenation. Point-by-point addition adds at least one retrieval enhancement content and a target impact factor with the same feature dimensions to obtain the concatenated at least one retrieval enhancement content and a target impact factor. Vector concatenation concatenates at least one retrieval enhancement content and a target impact factor in a specific order to obtain the concatenated at least one retrieval enhancement content and a target impact factor.
[0124] Part Three: Based on the first and second recommendation information, determine the target recommendation information.
[0125] The target suggestion information is the final output suggestion information. In some embodiments, either the first suggestion information or the second suggestion information can be used as the target suggestion information, or the first suggestion information and the second suggestion information can be merged and the merged information can be used as the target suggestion information.
[0126] In this implementation, at least one retrieval enhancement content is obtained from the target knowledge base using the target impact factor. Then, based on the retrieval enhancement content, different models are used to determine the first and second suggestion information, ultimately obtaining the target suggestion information to be output. The retrieval enhancement content provides a reliable data foundation for suggestion generation, avoiding random model generation. At the same time, parallel reasoning using two different architecture models can generate suggestion information with advantages from different perspectives, ultimately obtaining the target suggestion information, thus improving the accuracy and reliability of the target suggestion information.
[0127] In one possible implementation, a target similarity is determined between first suggestion information and second suggestion information. If the target similarity is greater than or equal to a calibrated similarity threshold, the first suggestion information and second suggestion information are fused to obtain target suggestion information. If the target similarity is less than the calibrated similarity threshold, a first confidence level corresponding to the first suggestion information and a second confidence level corresponding to the second suggestion information are determined based on the retrieved enhanced content, and the target suggestion information is determined based on the first confidence level and the second confidence level.
[0128] Here, the target similarity is the similarity between the first suggestion information and the second suggestion information. In some embodiments, the target similarity is expressed as a percentage, and the target similarity can be any suitable size, such as 62%, 93%, etc. The similarity threshold can be any suitable value, such as 80%, 95%, etc. In some embodiments, the target similarity is determined by a similarity calculation algorithm. This similarity calculation algorithm may include, but is not limited to, CS, ED, etc.
[0129] In some embodiments, when the target similarity is greater than or equal to a calibrated similarity threshold, the first suggestion information and the second suggestion information are fused to obtain the target suggestion information. Specifically: First, NLP technology is used to identify and remove completely repeated statements and highly overlapping semantic units in the first and second suggestion information. Then, unique key information points are extracted from the first and second suggestion information and integrated complementaryly according to logical relationships (e.g., causal, parallel, progressive, etc.). Finally, the language structure is reorganized based on the principles of topic consistency and language coherence to obtain the target suggestion information.
[0130] In some embodiments, when the target similarity is less than a calibrated similarity threshold, a first confidence level corresponding to the first suggested information and a second confidence level corresponding to the second suggested information are determined based on the retrieved enhanced content. The target suggested information is then determined based on the first and second confidence levels. Confidence level is an important concept in statistics and machine learning, reflecting the reliability and accuracy of an estimate or prediction, and is usually expressed as a percentage. Confidence level is commonly used in statistical analysis to help determine the reliability and significance level of results.
[0131] The first confidence level is the confidence level corresponding to the first suggestion information. The first confidence level can be any suitable size, such as 0.53, 0.81, etc. The second confidence level is the confidence level corresponding to the second suggestion information. The second confidence level can be any suitable size, such as 0.68, 0.76, etc. In some embodiments, the first and second suggestion information are evaluated using enhanced retrieval content to obtain the first and second confidence levels respectively. When determining the target suggestion information based on the first and second confidence levels, the suggestion information corresponding to the higher confidence level is selected as the target suggestion information.
[0132] In this implementation, the target similarity between the first and second suggestion information is determined. Then, based on the relationship between the target similarity and the calibrated similarity threshold, the first and second suggestion information are either fused or selected as the target suggestion information based on confidence level. When the suggestions of the two models are relatively consistent, reliable target suggestion information is quickly fused to improve response speed. When there is a discrepancy in the suggestions, a decision is made based on factual evidence, avoiding misjudgments caused by meaningless conflicts or illusions between models. This ensures that the output target suggestion information is always based on reliable information, significantly improving the system's decision robustness in complex scenarios.
[0133] In summary, the method provided in this application determines relevant parameters, including target data indicators, from the acquired target input information, then obtains at least one target data distribution information that is correlated with the target data indicators, and further determines the target analysis result corresponding to the target input information based on the target contribution degree corresponding to the determined target data distribution information. This method can accurately determine the contribution degree by referring to the information of merchants and users in time and space, accurately map unstructured natural language problems into specific relevant parameters and spatiotemporal data distributions, quantify the impact of each data distribution information on business indicators, accurately perform data analysis on complex input problems, quickly and accurately locate the cause of the problem, and reduce the threshold and time cost of data analysis.
[0134] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values or scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes based on the above examples, and such modifications or changes also fall within the scope of the embodiments of this application.
[0135] The above text combined Figures 1-4 The data analysis method provided in the embodiments of this application is described in detail below; the following will be combined with Figure 5 and Figure 6 The apparatus embodiments of this application are described in detail below. It should be understood that the apparatus in the embodiments of this application can perform the various methods described in the foregoing embodiments of this application, that is, the specific working processes of the various products described below can be referred to the corresponding processes in the foregoing method embodiments.
[0136] Figure 5 This is a schematic diagram of the structure of a data analysis device provided in an embodiment of this application. Figure 6 As shown, the data analysis device 500 includes: an acquisition module 510 and a determination module 520. Wherein: The acquisition module 510 is used to acquire target input information, determine relevant parameters, including target data indicators, and acquire at least one target data distribution information that is related to the target data indicators. The target data distribution information includes information on merchants and users in time and / or space. The determination module 520 is used to determine the target contribution degree corresponding to the target data distribution information. The target contribution degree is used to characterize the degree of influence of the target data distribution information on the target data indicators. Based on the target contribution degree, the target analysis result corresponding to the target input information is determined.
[0137] In one possible implementation, the relevant parameters also include a target time period. A determining module 520 is used to select first-stage data and second-stage data from target data distribution information based on the target time period. The first-stage data is used to characterize the data distribution at the beginning of the target time period, and the second-stage data is used to characterize the data distribution at the end of the target time period. Based on the first-stage data and the second-stage data, the target contribution is determined.
[0138] In one possible implementation, the determining module 520 is used to determine a first ratio indicator value based on the first stage data; determine a second ratio indicator value based on the second stage data; and determine the target contribution based on the first ratio indicator value and the second ratio indicator value.
[0139] In one possible implementation, the target data distribution information has corresponding influencing factors. The determination module 520 is used to determine the target percentage contribution based on the second ratio index value, which is used to characterize the impact of the weight change of the influencing factor on the target contribution; and to determine the target ratio contribution based on the first ratio index value and the second ratio index value, which is used to characterize the impact of the change of the performance level of the influencing factor on the target contribution; and to use the sum of the target percentage contribution and the target ratio contribution as the target contribution.
[0140] In one possible implementation, the target data distribution information has a corresponding influence factor. The determining module 520 is used to determine a target influence factor from the influence factors corresponding to at least one target data distribution information, wherein the target contribution degree corresponding to the target influence factor is the maximum value among the target contribution degrees. Based on the target influence factor, target recommendation information is determined. The device also includes an output module for outputting the target recommendation information.
[0141] In one possible implementation, the determining module 520 is configured to obtain at least one retrieval enhancement content from the target knowledge base based on the target impact factor; determine first suggestion information and second suggestion information based on the retrieval enhancement content, wherein the first suggestion information is generated by a first model and the second suggestion information is generated by a second model, and the first model and the second model are different; and determine target suggestion information based on the first suggestion information and the second suggestion information.
[0142] In one possible implementation, the determining module 520 is used to determine the target similarity between the first suggestion information and the second suggestion information; if the target similarity is greater than or equal to a calibrated similarity threshold, the first suggestion information and the second suggestion information are fused to obtain target suggestion information; if the target similarity is less than the calibrated similarity threshold, based on the retrieved enhanced content, a first confidence level corresponding to the first suggestion information and a second confidence level corresponding to the second suggestion information are determined, and the target suggestion information is determined based on the first confidence level and the second confidence level.
[0143] In one possible implementation, the acquisition module 510 is used to acquire a prompt text template; the device further includes a processing module, used to fill the prompt text template with target data distribution information to obtain target prompt text; and through the target model, based on the target prompt text, to perform the step of determining the target contribution corresponding to the target data distribution information.
[0144] In one possible implementation, the prompt text template includes at least one of the following: a first prompt text, a second prompt text, and a third prompt text, wherein the first prompt text is used to prompt the role corresponding to the target model, the second prompt text is used to prompt the capabilities of the target model, and the third prompt text is used to prompt the limitations of the target model.
[0145] The division of modules in the above-described data analysis device is for illustrative purposes only. In other embodiments, the data analysis device can be divided into different modules as needed to complete all or part of the functions of the above-described data analysis device.
[0146] The various modules in the data analysis apparatus provided in this application embodiment can be implemented in the form of computer programs. These computer programs can run on a server or client electronic device. The program modules constituted by these computer programs can be stored in the memory of the server or client electronic device. When the computer program is executed by a processor, it implements all or part of the steps of the methods described in this application embodiment.
[0147] It should be noted that the data analysis device 500 described above is embodied in the form of a functional unit. The term "module" here can be implemented in software and / or hardware, without specific limitations.
[0148] For example, a "module" can be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components that support the described functions.
[0149] Therefore, the units of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0150] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0151] For example, such as Figure 6 As shown, the electronic device 600 includes a memory 601 and a processor 602. The memory 601 stores executable program code 6011, and the processor 602 is used to call and execute the executable program code 6011 to perform a data analysis method.
[0152] For example, memory 601 can be used to store programs related to the data analysis methods provided in the embodiments of this application; processor 602 can call the programs related to the data analysis methods stored in memory 601 to execute the data analysis methods of the embodiments of this application; for example, acquiring target input information, determining relevant parameters, the relevant parameters including target data indicators; acquiring at least one target data distribution information that is correlated with the target data indicators, the target data distribution information including information about merchants and users in time and / or space; determining the target contribution degree corresponding to the target data distribution information, the target contribution degree being used to characterize the degree of influence of the target data distribution information on the target data indicators; and determining the target analysis result corresponding to the target input information based on the target contribution degree.
[0153] This embodiment can divide the device into functional modules based on the above method example. For example, each module can correspond to a separate function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0154] When the functional modules are divided according to their respective functions, the device may also include a processing module and a communication module. It should be noted that all relevant content regarding the steps involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here.
[0155] It should be understood that the apparatus provided in this embodiment is used to perform the above-described data analysis method, and therefore can achieve the same effect as the above-described implementation method.
[0156] When using integrated units, the device may include a processing module and a storage module. The processing module may be a processor or a controller that can implement or execute various exemplary logic blocks, modules, and circuits shown in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module may be a memory.
[0157] In addition, the apparatus provided in the embodiments of this application may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute a data analysis method provided in the above embodiments.
[0158] This application also provides a computer-readable storage medium storing computer program code, which, when run on a computer, causes the computer to execute the aforementioned method steps to implement a data analysis method provided in the above embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, Digital Video Discs (DVDs), Compact Disc Read-Only Memory (CD-ROMs), microdrives, and magneto-optical disks, read-only memory (ROMs), random access memory (RAMs), erasable programmable read-only memory (EPROMs), electrically erasable programmable read-only memory (EEPROMs), dynamic random access memory (DRAMs), video random access memory (VRAMs), flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of media or device suitable for storing instructions and / or data.
[0159] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a data analysis method provided in the above embodiments.
[0160] The computer-readable storage medium, computer program product, or chip provided in this application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0161] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0162] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0163] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data analysis method, characterized in that, include: Obtain target input information and determine relevant parameters, including target data indicators; Obtain at least one target data distribution information that is correlated with the target data indicator, the target data distribution information including information on merchants and users in time and / or space; Determine the target contribution degree corresponding to the target data distribution information, wherein the target contribution degree is used to characterize the degree of influence of the target data distribution information on the target data indicator; Based on the target contribution, the target analysis result corresponding to the target input information is determined.
2. The method according to claim 1, characterized in that, The relevant parameters also include the target time period, and determining the target contribution corresponding to the target data distribution information includes: Based on the target time period, first-stage data and second-stage data are selected from the target data distribution information. The first-stage data is used to characterize the data distribution at the beginning of the target time period, and the second-stage data is used to characterize the data distribution at the end of the target time period. The target contribution is determined based on the data from the first stage and the data from the second stage.
3. The method according to claim 2, characterized in that, The determination of the target contribution based on the data from the first stage and the data from the second stage includes: Based on the data from the first stage, determine the value of the first ratio indicator; Based on the data from the second phase, determine the value of the second ratio indicator; The target contribution is determined based on the first ratio indicator value and the second ratio indicator value.
4. The method according to claim 3, characterized in that, The target data distribution information has corresponding influence factors, and determining the target contribution based on the first ratio index value and the second ratio index value includes: The target percentage contribution is determined based on the second ratio index value, and the target percentage contribution is used to characterize the impact of the weight change of the influencing factor on the target contribution. Based on the first ratio index value and the second ratio index value, a target ratio contribution is determined, wherein the target ratio contribution is used to characterize the impact of the change in the performance level of the influencing factor on the target contribution. The sum of the target percentage contribution and the target ratio contribution is taken as the target contribution degree.
5. The method according to claim 1, characterized in that, The target data distribution information has corresponding influencing factors, and the method further includes: A target impact factor is determined from at least one impact factor corresponding to the target data distribution information, wherein the target contribution of the target impact factor is the maximum value among the target contributions; Based on the target impact factors, target recommendation information is determined; Output the target suggestion information.
6. The method according to claim 5, characterized in that, The determination of target recommendation information based on the target impact factor includes: Based on the target impact factor, at least one enhanced retrieval content is obtained from the target knowledge base; Based on the enhanced retrieval content, first suggestion information and second suggestion information are determined. The first suggestion information is generated by a first model, and the second suggestion information is generated by a second model. The first model and the second model are different. The target recommendation information is determined based on the first recommendation information and the second recommendation information.
7. The method according to claim 6, characterized in that, The step of determining the target suggestion information based on the first suggestion information and the second suggestion information includes: Determine the target similarity between the first suggestion information and the second suggestion information; If the target similarity is greater than or equal to the calibrated similarity threshold, the first suggestion information and the second suggestion information are fused to obtain the target suggestion information; If the target similarity is less than the calibrated similarity threshold, based on the enhanced retrieval content, a first confidence level corresponding to the first suggestion information and a second confidence level corresponding to the second suggestion information are determined, and the target suggestion information is determined based on the first confidence level and the second confidence level.
8. An electronic device, characterized in that, include: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 7.