Big data resource management and analysis system
Through the big data resource management and analysis system, the problem of information dispersion and analysis problems in the field of food safety is solved, and the rapid and accurate identification and evaluation of food safety trends and risks is achieved.
Patent Information
- Application Number
- CN202411861679.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-23
AI Technical Summary
The current food safety field is facing the problem of information dispersion and inability to effectively analyze statistically, which leads to lagging handling of food safety issues and increasing public health risks.
Provide a big data resource management and analysis system, identify food safety trends through real-time data collection and preprocessing, text standardization and sentiment analysis, use statistical analysis and machine learning to identify food safety trends, and create risk assessment models to assign risk labels to different food products.
Real-time and comprehensive collection, analysis and evaluation of multi-source data is realized, which can identify food safety trends and problems more quickly and accurately, and provide more effective risk assessment and prediction capabilities.
Smart Images

Figure CN120030291A_ABST
Abstract
Claims
1. A big data resource management and analysis method, the method comprising the following steps: Step 1: Real-time data collection and preprocessing; Step 2: Perform text standardization and sentiment analysis on the collected text data; Step 3: Use statistical analysis and machine learning to identify food safety trends and predict future issues; Step 4: Creation of risk assessment models and label assignment.
2. The big data resource management and analysis method according to claim 1, characterized in that: The steps of real-time data collection and preprocessing include: Regularly scrape information from multiple data sources, including social media, news reports, consumer complaints, and food supply chain data; Use custom web crawlers and API technology to collect information, including real-time updates on specific keywords or topics; Perform data cleaning and standardization, including duplicate data tagging, missing value processing, text standardization, and entity recognition.
3. The big data resource management and analysis method according to claim 2, characterized in that: The data cleaning and standardization, including the steps of duplicate data marking, missing value processing, text standardization and entity recognition, includes: The collected data is processed by using hash functions to generate unique identifiers for the text. These identifiers are then compared to find duplicate data records. Once duplicates are detected, the system will automatically mark and remove them. Secondly, missing value processing is performed. For missing text data, natural language processing tools such as NLTK are used to fill in the missing text information, such as using contextual information to infer missing words or phrases. In addition, data interpolation techniques, such as interpolation based on mean or regression models, can be used to fill in missing values for numerical data. At the same time, text mining techniques are used to extract key information, such as dates, locations, food product names, and keywords related to food safety. Finally, natural language processing techniques, such as named entity recognition, are used to identify entities in food supply chain data, such as producers and distributors, to understand the source of food safety issues. The above data is ultimately stored in a distributed database system to ensure scalability and high availability.
4. The big data resource management and analysis method according to claim 1, characterized in that: The step of performing text standardization and sentiment analysis on the collected text data includes: Use NLP tools to normalize text, including stemming, lemmatization, and stop word removal, for subsequent analysis and comparison; Use deep learning models to perform sentiment analysis to determine whether text contains negative information.
5. The big data resource management and analysis method according to claim 4, characterized in that: The steps of performing sentiment analysis using a deep learning model to determine whether a text contains negative information include: Sentiment analysis is performed on the collected text data. A deep learning model is used, and recurrent neural networks and convolutional neural networks are selected. Recurrent neural networks are used for sentiment analysis of text contexts, while convolutional neural networks can effectively capture local features in texts. Before sentiment analysis, a labeled dataset needs to be prepared. Each text in the dataset is accompanied by a sentiment label, such as "positive", "neutral" or "negative". This dataset is used to train and verify the deep learning model. Subsequently, model training is performed, and the recurrent neural network and convolutional neural network models are trained on the prepared dataset. In this process, the model learns how to identify the sentiment in the text and classify the text into predefined sentiment labels. The model continuously adjusts weights and biases through the back propagation algorithm to improve its performance. Once the model training is completed, it is applied to unlabeled food data text. The model will automatically classify the text into "positive", "neutral" or "negative" sentiment categories, which helps to determine whether the text contains negative information related to food safety. By adopting a deep learning model, we can accurately classify and evaluate the intensity of sentiment in food data, providing a powerful tool for the identification and handling of food safety issues.
6. The big data resource management and analysis method according to claim 1, characterized in that: The steps described for using statistical analysis and machine learning to identify food safety trends and predict future problems include: Use sliding windows to observe the changing trends of events to identify seasonal and cyclical trends; Apply seasonal decomposition techniques, such as Holt-Winters seasonal decomposition, to decompose time series data to identify seasonal variations.
7. The big data resource management and analysis method according to claim 6, characterized in that: The steps of using the sliding window method to observe the changing trend of events to identify seasonal and cyclical trends include: Using the sliding window method, we first define a time window, such as a month or a quarter, and then calculate the frequency of specific food safety events within each window. If the frequency of an event increases significantly within a specific time window, the system can reasonably infer that there is a certain seasonal or cyclical trend, such as food safety issues are more prominent during specific seasons or holidays. In addition, seasonal decomposition techniques are applied, such as Holt-Winters seasonal decomposition, to decompose time series data into trend, seasonal and residual parts. By analyzing the results of seasonal decomposition, we can determine whether the frequency of food safety events in a specific season or time period shows a clear increasing or decreasing trend.
8. The big data resource management and analysis method according to claim 1, characterized in that: The steps of creating the risk assessment model and assigning labels include: In the feature engineering phase, trend data and other relevant features are used as inputs to the model; Models are evaluated and tuned to ensure performance.
9. The big data resource management and analysis method according to claim 8, characterized in that: The steps of using trend data and other relevant features as inputs to the model during the feature engineering phase include: In the feature engineering stage, these trend data and other relevant features, such as food type and production location, are used as input features of the model. Subsequently, the decision tree algorithm is selected and the prepared training data is used to train the model. The model learns how to assess the risk of food products based on different features and trend data. The splitting rules of the decision tree will be determined based on the importance of the features. When the model training is completed, it can be used to assign risk labels to different food products. These labels include low risk, medium risk and high risk. Based on the evaluation results of the model, finally, the model is evaluated to verify its performance and make necessary adjustments and optimizations.
10. Big data resource management and analysis system, characterized by: The system comprises: Data collection and preprocessing module, used to collect data and perform cleaning operations; Text data standardization and sentiment analysis module, used to standardize the collected text data and perform sentiment bias analysis; Statistical analysis and risk assessment modules are used to analyze the current food safety situation and conduct risk assessment.