An artificial intelligence-based tobacco data acquisition and analysis system
By combining multi-source data acquisition, processing, and intelligent analysis modules with large language models and random forest models, the problem of traditional systems being unable to handle complex tobacco data has been solved, achieving efficient and automated data acquisition and in-depth analysis, and generating accurate visualization results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2025-04-28
- Publication Date
- 2026-07-03
AI Technical Summary
Traditional data acquisition and analysis systems are unable to automatically collect tobacco data with complex structures and diverse types, and generate comprehensive and visualized analysis results.
By employing a multi-source data acquisition module, a data processing and transformation module, and an intelligent analysis module, combined with a large language model and a random forest model, the system achieves automated acquisition, cleaning, standardization, feature extraction, and in-depth analysis of tobacco data, generating visualization tools.
It enables efficient and automated collection and in-depth analysis of complex tobacco data, generating accurate visualization tools to support rapid decision-making.
Smart Images

Figure CN120631958B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tobacco data acquisition and analysis, and in particular to an artificial intelligence-based tobacco data acquisition and analysis system. Background Technology
[0002] In the tobacco industry, data typically refers to structured or semi-structured data encompassing both corpus data and web observation data. Corpus data includes consumer insights, market research, product evaluation, and social media sentiment data, while web observation data includes purchasing power data, component analysis data, and product description data scraped from various tobacco-related websites. The collection and analysis of tobacco data directly impacts production efficiency, supply chain optimization, and precise marketing. Therefore, comprehensive tobacco data collection and the provision of intuitive, visualized analytical results are crucial for understanding complex data relationships and making rapid decisions.
[0003] Traditional data acquisition and analysis systems typically only have basic data entry and simple data processing functions, limited to handling highly structured, single-type datasets, and still rely on manual analysis. Therefore, they cannot automatically collect complex and diverse tobacco data, nor generate comprehensive and visualized analysis results based on different types of tobacco data, which does not meet the processing needs of tobacco data.
[0004] Therefore, providing an artificial intelligence-based tobacco data acquisition and analysis system to ensure the automated acquisition of tobacco data with complex structures and diverse types, and to conduct intelligent in-depth analysis based on its data characteristics to generate accurate analysis results and comprehensive visualization tools is an urgent problem to be solved. Summary of the Invention
[0005] Based on the foregoing analysis, the main objective of this invention is to provide an artificial intelligence-based tobacco data acquisition and analysis system that solves the problem that traditional data acquisition and analysis systems cannot automatically acquire tobacco data with complex structures and diverse types, and obtain intelligent analysis results and visualization tools based on its data characteristics.
[0006] In response, the present invention provides an artificial intelligence-based tobacco data acquisition and analysis system, comprising:
[0007] The multi-source data acquisition module is used to automatically and incrementally collect tobacco corpus data and network observation data based on a pre-established polling mechanism, and output them as tobacco data after unifying the format. The data processing and conversion module is used to clean, verify, uniformly encode, and generate standardized data for the tobacco data. The intelligent analysis module is used to extract features from the standardized data, reduce the dimensionality to generate a structured dataset, perform deep analysis on the structured dataset based on the integrated large language model and random forest model, generate analysis results, and dynamically optimize the analysis results by optimizing the hyperparameter configuration of the large language model. A visualization tool is generated based on the optimized analysis results.
[0008] Preferably, the multi-source data acquisition module includes: a streaming computing unit, used to formulate the polling mechanism, collect and incrementally calculate the corpus data and the network observation data in real time, and restart the acquisition behavior within a preset number of times when the acquisition fails; an automatic questionnaire data acquisition unit, integrating an HTTP interface, used to connect to a third-party questionnaire platform, and collect the corpus data based on the polling mechanism; a website data intelligent extraction unit, used to load DOM analysis technology and XPath expressions, and collect the network observation data from the target website based on the polling mechanism; an acquisition task scheduling unit, used to create tasks for the acquisition behavior, allocate cycles for the tasks and track their progress in real time; and a data format adaptation unit, used to unify the corpus data and the network observation data into JSON format and output it as tobacco data.
[0009] Preferably, the data processing and conversion module includes: a data cleaning unit for cleaning null values, outliers, and duplicate data from the tobacco data; a data quality control unit for verifying the fields of the tobacco data after cleaning, outputting non-compliant data to the data cleaning unit, and retaining compliant data; and a data conversion unit for converting the data type of the compliant data, unifying the encoding, generating standardized data, and outputting it to the intelligent analysis module.
[0010] Preferably, the intelligent analysis module includes: a data preprocessing unit, used to extract features from the standardized data and generate feature data, and to reduce the dimensionality of the feature data to generate the structured dataset; a large model integration unit, used to interface with the API of the DeepSeek-R1 large language model, input the structured dataset to it via asynchronous requests, and receive the analysis results; an intelligent analysis unit, used to analyze the structured dataset through the DeepSeek-R1 large language model, and to dynamically optimize the analysis results by optimizing the hyperparameter configuration of the large language model, and then output the analysis results; and a result processing submodule, used to generate visualization tools based on the analysis results.
[0011] As a further preferred embodiment, the intelligent analysis module further includes: a questionnaire generation module, used to generate open-ended questions and corresponding answers based on the analysis results, so as to automatically generate a survey questionnaire; and a decision-making module, used to generate consumer behavior pattern reports based on the analysis results.
[0012] As a further preferred embodiment, the result processing submodule further includes: a chart display unit, used to generate interactive bar charts, line charts, or heat maps from the analysis results based on the integrated components; a table display unit, used to generate tables from the analysis results; and a map display unit, used to display the geographical distribution information of tobacco planting and consumption areas in the analysis results.
[0013] As a further preferred embodiment, the analysis steps of the intelligent analysis unit include: analyzing the structured dataset. Feature selection, feature transformation, and domain feature addition are performed to form an enhanced feature set. ; for the enhanced feature set Dimensionality reduction processing yields a dimensionality-reduced feature set. : and the dimensionality reduction feature set Input the DeepSeek-R1 large language model Combined with the aforementioned DeepSeek-R1 large language model With the aforementioned random forest model The analysis yielded the first-order analysis results. ; ; Generate the first-order analysis results Natural Language Report The analysis results are then integrated. and the analysis results , and perform iterative optimization.
[0014] As a further preferred option, the generation of the first-order analysis result Natural Language Report The analysis results are then integrated. and the analysis results Iterative optimization was performed, including: utilizing the DeepSeek-R1 large language model. By combining LIME and SHAP value interpretation algorithms, a natural language report is generated. : Bayesian optimization and genetic algorithms were applied to optimize all participants in the DeepSeek-R1 large language model. The hyperparameter configuration And find the optimal set of hyperparameters. Used to adjust the DeepSeek-R1 large language model Analysis speed: Establish a feedback mechanism: based on the natural language report. Combining the aforementioned optimal hyperparameter set Obtain the optimal analysis result : ,in This indicates dynamic adaptation to new data. Indicates changes in demand; adopts the optimal analysis results. Replace the original analysis results .
[0015] As a further preferred embodiment, it also includes a relational database for receiving and hierarchically storing the analysis results, the questionnaires, the behavior pattern reports, and the visualization tools; specifically, it includes: a data integration unit for grouping the analysis results, the questionnaires, the behavior pattern reports, and the visualization tools according to their relevance; a relational storage unit for establishing storage for the analysis results based on the groupings, and associating the corresponding questionnaires, behavior pattern reports, and visualization tools for storage; a search engine for creating indexes and associating them with the groupings, scoring the relevance of the analysis results under the groupings, and sorting the analysis results according to the scores from high to low; and a caching unit for reading the index counts and temporarily storing the analysis results with index counts > N. When the cache unit capacity is less than 5%, the analysis result with the lowest current index count is automatically removed.
[0016] Preferably, the system also includes a data security protection module to ensure the security of the tobacco data acquisition and analysis system during user operation. Specifically, this includes: a data encryption unit for key management, encryption, and decryption of the analysis results; an access control unit for establishing a fine-grained permission management mechanism and allocating permissions to restrict user access; an audit log unit for periodically recording user operation logs and monitoring abnormal behavior; a data backup unit for periodically backing up the analysis results; and a status monitoring unit for real-time tracking and monitoring of the load, response time, and throughput indicators of the tobacco data acquisition and analysis system.
[0017] The tobacco data acquisition and analysis system based on artificial intelligence of the present invention has the following beneficial effects:
[0018] First, unlike traditional systems that simply collect and input highly structured, single-type datasets, this system's multi-source data acquisition module automatically and incrementally collects tobacco corpus data and network observation data based on a defined polling mechanism. While automatically collecting data, the real-time polling mechanism improves the data update frequency and completeness.
[0019] Secondly, unlike traditional systems that repeatedly clean and verify complex data, this system first outputs tobacco data after the automatically collected data is initially formatted and then cleaned, verified and uniformly encoded using the data processing and conversion module to ensure data consistency. While improving data quality, the efficient preprocessing steps significantly shorten the data processing time.
[0020] Furthermore, unlike traditional systems that rely on a single model for data analysis, this system's intelligent analysis module employs a fusion of a large language model and a random forest model to analyze the structured datasets generated after feature extraction and dimensionality reduction. This allows for in-depth analysis of complex patterns and potential relationships within the structured datasets. Simultaneously, it can identify and extract key features from standardized data, ensuring the data used in the analysis is highly relevant and representative. Moreover, by dynamically optimizing the hyperparameters of the large language model and continuously adjusting the model output through iterative processes, it obtains optimal, accurate, and comprehensive analysis results. This ensures that the model can self-adjust based on the latest data and feedback for different structured datasets. Finally, it generates targeted visualization tools based on the analysis results, providing powerful data support and insights for decision-making. Attached Figure Description
[0021] Figure 1 This is a system topology diagram of an artificial intelligence-based tobacco data acquisition and analysis system according to an embodiment of the present invention;
[0022] Figure 2 This is an analysis flowchart of the intelligent analysis unit of an artificial intelligence-based tobacco data acquisition and analysis system according to an embodiment of the present invention. Detailed Implementation
[0023] The present invention will now be described in more detail with reference to the accompanying drawings. It should be noted that the following description of the present invention with reference to the accompanying drawings is merely illustrative and not restrictive.
[0024] Where possible, the various embodiments described below can be rearranged to form other embodiments not shown in the following description; the various technical features described below can also be rearranged to form other embodiments not shown in the following description.
[0025] Example 1:
[0026] Please refer to the appendix. Figure 1 and attached Figure 2 .
[0027] To address the problem that traditional data acquisition and analysis systems cannot automatically acquire tobacco data with complex structures and diverse types, and obtain intelligent analysis results and visualization tools based on its data characteristics, this embodiment provides an artificial intelligence-based tobacco data acquisition and analysis system. The system mainly includes a multi-source data acquisition module, a data processing and conversion module, and an intelligent analysis module. The above modules perform their respective functions to realize the entire process from automated acquisition and processing of tobacco data to intelligent in-depth analysis to generate analysis results and visualization tools.
[0028] The multi-source data acquisition module automatically and incrementally collects tobacco corpus data and network observation data based on a pre-established polling mechanism, outputting tobacco data after unifying the format. The data processing and transformation module cleans, verifies, uniformly encodes, and generates standardized data for tobacco data. The intelligent analysis module extracts features from the standardized data, reduces dimensionality to generate a structured dataset, performs deep analysis of the structured dataset based on a fusion random forest model with a pre-connected large language model, generates analysis results, and dynamically optimizes the analysis results by optimizing the hyperparameter configuration of the large language model. A visualization tool is generated based on the optimized analysis results.
[0029] It should be noted that the aforementioned tobacco data usually refers to tobacco industry-related data in various formats collected from different networks and platforms, which can provide comprehensive information on production, sales, market trends, and consumer behavior.
[0030] In this embodiment: The data processing and transformation module automatically standardizes and structures tobacco data by formulating refined data processing rules, ensuring that all data input to the intelligent analysis module has a unified format. The multi-source data acquisition module is pre-configured with interfaces and rule sets specifically for collecting tobacco-related data and supports automated data acquisition processes through simple target settings. The intelligent analysis module integrates the advantages of large-scale pre-trained language models and random forest algorithms, providing deep data analysis capabilities and significantly improving the accuracy and insight of the analysis results. This module not only relies on the powerful text understanding and generation capabilities of large language models for intelligent analysis but also dynamically improves the analysis model through continuous optimization of hyperparameter configurations, ensuring that the accuracy of the analysis results reaches the optimal level. Based on the optimized analysis results, the system can automatically generate highly customized visualization tools to intuitively and interactively display complex data relationships and key findings. It should be noted that the intelligent analysis module supports multi-dimensional and multi-perspective data analysis, covering tobacco industry-related time series analysis, market forecasting, and consumer behavior pattern recognition, thereby providing support for tobacco-related decision-making. Its efficient algorithms significantly improve data processing speed and effectively reduce manual intervention, thereby achieving automated processes.
[0031] The modules described above work closely together to not only efficiently process complex tobacco data, but also accurately extract valuable information from unstructured text, such as consumer feedback and market trends, thereby providing decision-makers with detailed and accurate analysis results and visualization tools.
[0032] In a preferred embodiment, a multi-source data acquisition module should be configured before tobacco data collection to ensure the accuracy and efficiency of data acquisition. First, basic website information should be configured, including website name, basic URL address, encoding format, User-Agent information, and necessary cookie information, with optional request header configuration. Furthermore, list pages and detail pages need to be configured according to page analysis rules: list page configuration includes URL template, pagination parameters, pagination parameter names and their ranges, list item selectors, and next-page button selectors; detail page configuration involves selectors for fields such as title, content, and date, metadata selectors for author, category, etc., and related link extraction rules.
[0033] It should be noted that, to ensure the stability of the data acquisition module and data quality, before data collection, it is necessary to define the acquisition task, schedule and configure the execution frequency, retry count, interval, and timeout, as well as the number of threads and their intervals. A failure handling strategy must also be established to achieve automatic tobacco data collection. Furthermore, the proxy pool type (static or dynamic) needs to be determined according to the proxy IP management strategy, the proxy API interface needs to be configured, a proxy availability testing mechanism needs to be established, and the proxy rotation strategy needs to be adjusted based on the number of accesses or the time required to maintain efficient tobacco data collection and provide a solid foundation for subsequent data analysis. These configurations collectively ensure system stability and data quality.
[0034] In a preferred embodiment, the multi-source data acquisition module includes: a streaming computing unit, used to establish a polling mechanism for real-time acquisition and incremental calculation of corpus data and network observation data, and to restart the acquisition behavior within a preset number of times if acquisition fails; an automatic questionnaire data acquisition unit, integrating an HTTP interface for connecting to a third-party questionnaire platform and acquiring the corpus data based on the polling mechanism; a website data intelligent extraction unit, used to load DOM analysis technology and XPath expressions, and acquire network observation data from target websites based on the polling mechanism; an acquisition task scheduling unit, used to create tasks for acquisition behavior, allocate task cycles, and track their progress in real time; and a data format adaptation unit, used to unify the corpus data and the network observation data into JSON format and output it as tobacco data.
[0035] In this embodiment: The questionnaire data automatic collection unit uses a polling mechanism to periodically synchronize questionnaire data, screen open-ended questions and filter invalid answers, and record the source and answer information. Text cleaning, content segmentation, key information extraction, and sentiment analysis are performed, and tags are automatically applied. Data quality is ensured by checking content length, duplicate detection, and sensitive word filtering, and the data integrity and usability are enhanced by associating it with basic questionnaire information, respondent information, and question context. The website data intelligent extraction unit uses a headless browser mode to access the target website, sets a page loading timeout limit, and handles exceptions based on timeout conditions to ensure stable access to the target website. After successful access to the target website, information is extracted from the webpage source code, the target content is analyzed according to preset rules, and data validity is verified.
[0036] It should be noted that the polling mechanism periodically accesses multiple data sources at set time intervals to ensure real-time information updates. This is suitable for scheduled synchronization tasks and allows for repeated data collection from target addresses that failed to collect data within a preset number of attempts until the relevant data is successfully collected. The questionnaire data automatic collection unit and the website data intelligent extraction unit utilize this mechanism to periodically obtain the latest corpus data and network observation data from third-party questionnaire platforms and target websites to ensure the timeliness and completeness of the data.
[0037] In this embodiment, the streaming computing unit processes tobacco data in real time, ensuring timeliness. The task scheduling unit allocates tasks, sets execution frequencies, manages the thread pool, tracks progress, and handles exceptions, ensuring orderly task execution. The data format adaptation unit unifies all data into JSON format, simplifying subsequent processing. These steps collectively guarantee the efficiency, accuracy, and stability of data acquisition, providing reliable data support and enhancing the system's flexibility and scalability.
[0038] In another preferred embodiment, the data processing and conversion module includes: a data cleaning unit for cleaning null values, outliers, and duplicate data in tobacco data; a data quality control unit for verifying fields of tobacco data after cleaning, outputting non-compliant data to the data cleaning unit, and retaining compliant data; and a data conversion unit for converting the data type of compliant data, unifying the encoding, generating standardized data, and outputting it to the intelligent analysis module.
[0039] In this embodiment: The data cleaning unit removes useless tags such as script and style from HTML, cleans up special characters and whitespace characters, standardizes line breaks, and removes blank lines; then, data standardization is performed, including standardizing the date format to yyyy-MM-dd HH:mm:ss, standardizing the number format, ensuring UTF-8 encoding, and converting special symbols. The data quality control unit focuses on data validity verification, duplicate data processing, and abnormal data marking. It ensures data quality through required field verification, data format verification, business rule verification, and duplicate data checks, and records source information, version history, and update records, maintaining operation logs to achieve comprehensive data tracking. The data transformation unit further performs standardization and duplicate data processing.
[0040] In another preferred embodiment, the intelligent analysis module includes: a data preprocessing unit for extracting features from standardized data and generating feature data, and reducing the dimensionality of the feature data to generate a structured dataset; a large model integration unit for connecting to the API of the DeepSeek-R1 large language model, inputting the structured dataset to it via asynchronous requests, and receiving the analysis results; an intelligent analysis unit for analyzing the structured dataset through the DeepSeek-R1 large language model, dynamically optimizing the analysis results by optimizing the hyperparameter configuration of the large language model, and then outputting the analysis results; and a result processing submodule for generating visualization tools based on the analysis results.
[0041] In this embodiment, the intelligent analysis module efficiently processes tobacco data through its various sub-units. The data preprocessing unit first extracts features from the standardized data, including identifying keywords, themes, and sentiment in the text, calculating statistical features such as frequency distribution and distribution parameters, and extracting trends and periodic changes from time-series data. Subsequently, these features are dimensionality-reduced to generate a structured dataset. The large model integration unit interfaces with the DeepSeek-R1 large language model API, using asynchronous requests to input the structured dataset and receive preliminary analysis results. The intelligent analysis unit then utilizes the DeepSeek-R1 large language model to deeply analyze this data, further improving analysis accuracy through dynamic optimization of hyperparameter configuration, ultimately outputting high-quality analysis results. The results processing sub-module automatically generates customized visualization tools based on these analysis results to intuitively display key findings. The entire process not only achieves full automation from data collection to analysis but also accurately extracts important information such as consumer feedback and market trends, providing enterprises with powerful data support and helping decision-makers quickly respond to market changes and formulate data-driven strategies.
[0042] It should be noted that the intelligent analysis unit also incorporates prompt word engineering technology to further improve the quality of text generation. This includes prompt word template management, providing customized input formats for different types of analysis tasks; context optimization strategies, dynamically adjusting input content based on specific application scenarios to enhance model understanding; a dynamic parameter adjustment mechanism, optimizing model parameters based on real-time feedback to improve output accuracy; and output quality control measures to ensure the accuracy and usability of the final analysis results.
[0043] In a further preferred embodiment, the intelligent analysis module further includes: a questionnaire generation module, used to generate open-ended questions and corresponding answers based on the analysis results, so as to automatically generate a survey questionnaire; and a decision-making module, used to generate consumer behavior pattern reports based on the analysis results.
[0044] In this embodiment: the questionnaire generation module utilizes the data insights and improvement suggestions of a large language model to generate more targeted questions. The decision-making module integrates the results of multi-source data analysis to help users comprehensively understand key information such as market trends and user behavior, thereby making more informed business decisions. The results processing submodule transforms the analysis results into visualization tools, presenting data through charts, tables, and maps to make complex analyses intuitive and easy to understand, and performs quality control to ensure the accuracy and applicability of the output.
[0045] In a further preferred embodiment, the results processing submodule further includes: a chart display unit for generating interactive bar charts, line charts, or heat maps based on the integrated components; a table display unit for generating tables from the analysis results; and a map display unit for displaying the geographical distribution information of tobacco planting and consumption areas in the analysis results.
[0046] In this embodiment: the chart display unit utilizes various visualization formats such as pie charts, donut charts, bar charts, column charts, and line charts to present the analysis results, making the data more intuitive and easy to understand. The table display unit lists various statistical data in detail in tabular form, facilitating precise viewing and comparative analysis. The map display unit uses geographic information visualization to showcase data analysis results related to geographical location, such as the distribution of tobacco sales areas.
[0047] In another preferred embodiment, the analysis steps of the intelligent analysis unit include: analyzing the structured dataset. Feature selection, feature transformation, and domain feature addition are performed to form an enhanced feature set. ; Enhanced feature set Dimensionality reduction processing yields a dimensionality-reduced feature set. : and the dimensionality reduction feature set Input DeepSeek-R1 large language model Combined with the DeepSeek-R1 large language model With Random Forest Model The analysis yielded the first-order analysis results. ; ; Generate first-order analysis results Natural Language Report Integrate to form analysis results and the analysis results , and perform iterative optimization.
[0048] It should be noted that the natural language report is generated by a large language model, including the distribution of options and answer trends in the corpus data, as well as the identification of abnormal data and extraction of key features in the network parsing data. Based on the content of the natural language report, the tobacco data collection and analysis system in this embodiment will provide intelligent suggestions, such as data insights, improvement and optimization, and decision support, and automatically generate analysis reports in various formats, supporting export functionality.
[0049] In a further preferred embodiment, first-order analysis results are generated. Natural Language Report Integrate to form analysis results and the analysis results Iterative optimization was performed, including: utilizing the DeepSeek-R1 large language model. By combining LIME and SHAP value interpretation algorithms, a natural language report is generated. : Bayesian optimization and genetic algorithms were applied to optimize all participants in the DeepSeek-R1 large language model. Hyperparameter configuration And find the optimal set of hyperparameters. Used to adjust the DeepSeek-R1 large language model Analysis speed: Establish a feedback mechanism: based on natural language reports. Combining the optimal set of hyperparameters Obtain the optimal analysis result : ,in This indicates dynamic adaptation to new data. Indicates changes in demand; uses optimal analysis results. Replace the original analysis results .
[0050] In this embodiment, the accuracy and depth of the analysis results are significantly improved through iterative optimization and feedback mechanisms. Advanced interpretable algorithms and optimization techniques are used to dynamically adjust the hyperparameter configuration of the large language model, ensuring optimal analysis speed and effectiveness. The feedback mechanism enables the system to quickly adapt to new data and changing needs, continuously improving the analysis results. This adaptive and optimization mechanism enhances the system's flexibility and responsiveness while ensuring the quality of the analysis results.
[0051] In a further preferred embodiment, a relational database is also included for receiving and hierarchically storing analysis results, questionnaires, behavioral pattern reports, and visualization tools; specifically, it includes: a data integration unit for grouping analysis results, questionnaires, behavioral pattern reports, and visualization tools according to their relevance; a relational storage unit for establishing storage for analysis results based on the groups and associating the corresponding questionnaires, behavioral pattern reports, and visualization tools for storage; a search engine for creating indexes and associating groups, scoring the relevance of analysis results under the groups, and sorting the analysis results according to the scores from high to low; and a caching unit for reading the index count and temporarily storing the analysis results with an index count > N. When the cache unit capacity is less than 5%, the analysis result with the lowest current index count is automatically removed.
[0052] In this embodiment: the data integration unit groups the analysis results, questionnaires, behavioral pattern reports, and visualization tools based on their relevance to ensure close connections between the data. The relational storage unit groups and stores the analysis results, while also storing the corresponding questionnaires and behavioral pattern reports for easy retrieval and analysis later. The search engine is responsible for index building, including document analysis, word segmentation, vectorized storage, and index updates, and employs strategies such as semantic similarity calculation, keyword matching, contextual understanding, and result ranking to achieve efficient retrieval. The search function covers full-text search, group filtering, category filtering, and advanced search to meet diverse query needs. In addition, the search engine also manages index optimization and search term correction to improve retrieval accuracy.
[0053] It's important to note that the number of indexes can be set according to the actual application. When the cache unit capacity is less than 5%, failure to clean it up in a timely manner can lead to a significant performance degradation, increased disk I / O operations, potential data loss, and service interruption. Therefore, when the cache unit capacity is less than 5%, the system will trigger a cleanup mechanism to free up space for more important data or new, frequently accessed data, thereby maintaining the system's efficient operation and responsiveness.
[0054] In another preferred embodiment, a data security protection module is further included to ensure the security of the tobacco data acquisition and analysis system during user operation. Specifically, this includes: a data encryption unit for key management, encryption, and decryption of analysis results; an access control unit for establishing a fine-grained permission management mechanism and assigning permissions to restrict user access; an audit log unit for periodically recording user operation logs and monitoring abnormal behavior; a data backup unit for periodically backing up analysis results; and a status monitoring unit for real-time tracking and monitoring of the tobacco data acquisition and analysis system's load, response time, and throughput metrics.
[0055] In this embodiment: The data encryption unit ensures secure transmission and storage, encompassing HTTPS configuration, certificate management, encryption algorithm selection, and transmission protocol optimization. It also implements sensitive data encryption, data anonymization, backup strategies, and data destruction mechanisms. The access control unit achieves authentication through JWT token generation, token validity management, and refresh mechanisms, and utilizes the RBAC model for granular permission control, dynamic permission allocation, and permission caching. The audit log unit records critical operations, supporting SQL injection protection, XSS attack protection, CSRF protection, and file upload protection to ensure system security. The data backup unit performs regular backups and ensures secure data recovery, enhancing system reliability. The status monitoring unit monitors system status in real time, including IP whitelisting, interface rate limiting, concurrency control, and session management, preventing potential threats and maintaining stable system operation. The entire process, through the collaborative work of these units, achieves comprehensive security management and efficient system protection.
[0056] It should be understood that the embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
Claims
1. A tobacco data acquisition and analysis system based on artificial intelligence, characterized in that, include: The multi-source data acquisition module is used to automatically and incrementally collect tobacco corpus data and network observation data based on the established polling mechanism, and output them as tobacco data after unifying the format. The data processing and conversion module is used to clean, verify, uniformly encode, and generate standardized data from the tobacco data. The intelligent analysis module is used to extract the features of the standardized data and reduce the dimensionality to generate a structured dataset. After deep analysis of the structured dataset based on the integrated large language model and random forest model, the analysis results are generated. The analysis results are dynamically optimized by optimizing the hyperparameter configuration of the large language model. A visualization tool is generated based on the optimized analysis results; The intelligent analysis module includes: The data preprocessing unit is used to extract features from the standardized data and generate feature data, and to reduce the dimensionality of the feature data to generate the structured dataset. The large model integration unit is used to interface with the API of the DeepSeek-R1 large language model, input the structured dataset into it via asynchronous requests, and receive the analysis results. The intelligent analysis unit is used to analyze the structured dataset based on the DeepSeek-R1 large language model, and after dynamically optimizing the analysis results by optimizing the hyperparameter configuration of the large language model, output the analysis results. The results processing submodule is used to generate visualization tools based on the analysis results; The analysis steps of the intelligent analysis unit include: For the structured dataset Feature selection, feature transformation, and domain feature addition are performed to form an enhanced feature set. ; For the enhanced feature set Dimensionality reduction processing yields a dimensionality-reduced feature set. : and the dimensionality reduction feature set Input the DeepSeek-R1 large language model ; Combined with the DeepSeek-R1 large language model With the aforementioned random forest model The analysis yielded the first-order analysis results. ; ; Generate the first-order analysis results Natural Language Report The analysis results are then integrated. and the analysis results Perform iterative optimization; The first-order analysis result is generated. Natural Language Report The analysis results are then integrated. and the analysis results Perform loop optimization, including: Using DeepSeek-R1 large language model By combining LIME and SHAP value interpretation algorithms, a natural language report is generated. : ; Bayesian optimization and genetic algorithms were applied to optimize all participants in the DeepSeek-R1 large language model. The hyperparameter configuration And find the optimal set of hyperparameters. Used to adjust the DeepSeek-R1 large language model Analysis speed: ; Establish a feedback mechanism: based on the natural language report. Combining the aforementioned optimal hyperparameter set Obtain the optimal analysis result : ,in This indicates dynamic adaptation to new data. Indicates changes in demand; Using the optimal analysis results Replace the analysis results .
2. The tobacco data acquisition and analysis system according to claim 1, characterized in that, The multi-source data acquisition module includes: The streaming computing unit is used to formulate the polling mechanism, collect and incrementally calculate the corpus data and the network observation data in real time, and restart the collection behavior within a preset number of times when the collection fails. The questionnaire data automatic collection unit integrates an HTTP interface for connecting to third-party questionnaire platforms and collects the corpus data based on the polling mechanism. The website data intelligent extraction unit is used to load DOM analysis technology and XPath expressions, and collect the network observation data from the target website based on the polling mechanism. The task scheduling unit is used to create tasks for the collection behavior, allocate periods for the tasks, and track their progress in real time. The data format adaptation unit is used to unify the corpus data and the network observation data into JSON format and output it as the tobacco data.
3. The tobacco data acquisition and analysis system according to claim 1, characterized in that, The data processing and conversion module includes: A data cleaning unit is used to clean null values, outliers, and duplicate data from the tobacco data. A data quality control unit is used to verify the fields of the tobacco data after cleaning, output non-compliant data to the data cleaning unit, and retain compliant data. The data conversion unit is used to convert the data type of the compliant data, unify the encoding, generate the standardized data, and output it to the intelligent analysis module.
4. The tobacco data acquisition and analysis system according to claim 1, characterized in that, The intelligent analysis module also includes: The questionnaire generation module is used to generate open-ended questions and corresponding answers based on the analysis results, so as to automatically generate a survey questionnaire; The decision-making module is used to generate consumer behavior pattern reports based on the analysis results.
5. The tobacco data acquisition and analysis system according to claim 1, characterized in that, The result processing submodule also includes: The chart display unit is used to generate interactive bar charts, line charts, or heatmaps from the analysis results based on the integrated components; The table display unit is used to generate a table from the analysis results; The map display unit is used to show the geographical distribution information of tobacco planting and consumption areas in the analysis results.
6. The tobacco data acquisition and analysis system according to claim 4, characterized in that, It also includes a relational database for receiving and hierarchically storing the analysis results, the questionnaires, the behavioral pattern reports, and the visualization tools; specifically including: The data integration unit is used to group the analysis results, the questionnaire, the behavior pattern report, and the visualization tool according to their correlation. A relational storage unit is used to store the analysis results based on the grouping, and to associate and store the corresponding questionnaires, behavioral pattern reports, and visualization tools. A search engine is used to create an index and associate the group, score the relevance of the analysis results under the group, and sort the analysis results from high to low according to the score. A cache unit is used to read the index count and temporarily store the analysis results with an index count > N. When the cache unit capacity is less than 5%, the analysis result with the lowest current index count is automatically removed.
7. The tobacco data acquisition and analysis system according to any one of claims 1 to 6, characterized in that, It also includes a data security protection module to ensure the security of the tobacco data acquisition and analysis system during user operation; specifically including: A data encryption unit is used for key management, encryption, and decryption of the analysis results; Access control units are used to establish fine-grained permission management mechanisms and assign permissions to restrict the scope of user access. The audit log unit is used to periodically record user operation logs and monitor their abnormal behavior. A data backup unit is used to periodically back up the analysis results; The status monitoring unit is used to track and monitor the load, response time, and throughput indicators of the tobacco data acquisition and analysis system in real time.
Citation Information
Patent Citations
Electronic cigarette production data management system based on big data
CN117035560A
Data analysis system based on artificial intelligence
CN119066423A