Data processing method and system based on AI technology and storage medium
Through AI-based data processing methods, combined with deep learning, graph neural network and privacy protection technology, the problems of low efficiency, insufficient privacy and poor flexibility in bidding data processing are solved, and efficient and secure data analysis and personalized report generation are achieved.
Patent Information
- Application Number
- CN202510415921.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technology has problems such as low data processing efficiency, insufficient privacy protection, weak data mining capabilities and poor report generation flexibility in bidding data processing, which cannot meet the efficient processing and personalized needs of massive data.
Data processing methods based on AI technology are adopted, including data acquisition, preprocessing, intelligent analysis, report generation and privacy protection, and fully automated data processing and personalized report generation are achieved using deep learning, graph neural network, natural language processing, federated learning and blockchain technology.
It improves data processing efficiency, enhances data privacy protection, realizes accurate prediction of market trends and corporate credit ratings, provides personalized and interactive report display, and improves user experience and decision-making scientificity.
Smart Images

Figure CN120257330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tender data processing, and specifically provides a data processing method, system and storage medium based on AI technology. Background Art
[0002] With the continuous growth of data volume and the development of technology, the bidding industry has gradually entered the digital age. In the prior art, the analysis of bidding data usually relies on manual collation and basic statistical analysis tools. These tools can process structured data, but for large-scale complex data, especially market analysis involving multi-dimensional information, traditional methods are stretched. Data collection usually relies on web crawlers and manual input, and data cleaning is mostly carried out through simple rules, often unable to comprehensively identify potential errors and redundancies in the data. Intelligent analysis is usually based on rules and linear models, difficult to handle non-linear relationships and multi-level patterns in the data, and lacking the ability to mine deep associations in the data. Report generation is mostly static and single, unable to be customized according to different needs, lacking dynamic interactivity and personalized support.
[0003] However, the prior art has certain limitations in dealing with these problems. First, the high workload and error risk brought by manual operations lead to low efficiency in data analysis and cannot meet the processing requirements of massive data. Second, traditional rule matching and linear models cannot deeply mine the potential information in the data. Especially in multi-dimensional data analysis, they can often only reveal surface trends and cannot comprehensively reveal deep features such as market dynamics and enterprise credit. Third, the privacy protection measures in the prior art often focus on encryption in the data storage stage, lacking privacy protection during the data processing process, and prone to the risk of data leakage. Finally, the fixed format and static display of the traditional report generation method make it lack flexibility and interactivity and cannot meet the personalized needs of users in a changing business environment. Therefore, how to efficiently, securely and deeply process bidding data and provide users with personalized and real-time updated reports has become an urgent technical problem to be solved. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a data processing method, system and storage medium based on AI technology, which solves the problems of low data processing efficiency, insufficient privacy protection, weak data deep mining ability and poor flexibility of report generation in bidding data analysis.
[0005] To achieve the above object, the present invention is realized through the following technical solutions: A data processing method based on AI technology, including: S1. Data collection: Collect bidding-related data from the target data source through web crawler technology and store the collected data in a distributed database; S2. Data preprocessing: Clean, standardize, and complete the collected data, removing abnormal data and filling in missing data; S3. Intelligent analysis: Based on a deep learning model, perform pattern recognition, trend prediction, and enterprise credit rating on the preprocessed data, and mine the correlation relationships between data; S4. Report generation: Generate a data analysis report based on the results of intelligent analysis and display the analysis results in a dynamic visualization manner; S5. Privacy protection: Encrypt sensitive data during the data processing process and combine federated learning technology to ensure data security.
[0006] Preferably, the data collection step includes: Use a distributed web crawler framework to concurrently collect data from bidding websites, enterprise databases, social media, and news media data sources; Through a dynamic update scheduling mechanism, adjust the data collection frequency according to the update frequency and content change rate of the data source, thereby optimizing the data collection efficiency; Store the collected data in a distributed database to support large-scale data storage and high-concurrency access.
[0007] Preferably, the data preprocessing step includes: Clean the collected data using natural language processing technology to remove redundant and invalid data; Detect outliers based on a clustering algorithm and remove data beyond the preset threshold range; Complete the missing data through a generative adversarial network, and the generative adversarial network generates data that meets expectations according to the data feature distribution.
[0008] Preferably, the intelligent analysis step includes: Perform pattern recognition on historical bidding data based on a deep neural network model to extract market trend features; Construct an enterprise knowledge graph with enterprises as nodes and the cooperation and competition relationships between enterprises as edges; Based on a graph neural network, iteratively update the node features of the enterprise association relationship graph, and generate a comprehensive feature representation of the enterprise through feature propagation between nodes; Use a time series analysis model to predict the future demand and competition dynamics of the bidding market.
[0009] Preferably, the iterative update of the graph neural network includes the following steps: Encode the initial features of each node; Update the state of each node through the feature propagation of neighbor nodes, where the node feature update formula is: ; Among them, : The feature representation of node i in the k-th layer; : The neighbor set of node i; : The weight matrix of the k-th layer; : The degree of node k; Finally, the comprehensive feature representation of the enterprise in the market environment is generated.
[0010] Preferably, the report generation step includes: Using natural language generation technology to generate an analysis report, which includes enterprise credit rating, market trend prediction, and competitiveness analysis; Providing a dynamic dashboard to display the analysis results through a visualization tool, and supporting users to filter and update the analysis content according to different dimensions; Preferably, the privacy protection step includes: Performing dynamic encryption on sensitive data to ensure the security of data during transmission and storage; Training the model in a distributed environment through federated learning technology to avoid sharing raw data; Recording the entire data processing process based on blockchain technology to achieve data traceability and immutability; Preferably, the data collection, data preprocessing, intelligent analysis, and report generation steps are all implemented based on a distributed computing architecture to support large-scale data processing and high-concurrency task execution.
[0011] A data processing system based on AI technology, comprising: A data collection module for collecting bidding data from a target data source through a distributed web crawler technology; A data preprocessing module for cleaning, standardizing, and complementing the collected data; An intelligent analysis module for pattern recognition, trend prediction, and enterprise credit rating based on a deep learning model; A report generation module for generating an analysis report including enterprise credit rating and market trend prediction, and displaying the analysis results through a dynamic visualization tool; A privacy protection module for encrypting sensitive data and protecting data privacy in combination with federated learning and blockchain technology.
[0012] A computer-readable storage medium stores a computer program, and the computer program, when executed by a computer, implements the data processing method based on AI technology according to any one of claims 1-8.
[0013] The present invention provides a data processing method, system, and storage medium based on AI technology. It has the following beneficial effects: 1. By introducing deep neural networks and graph neural networks, the present invention can deeply explore the market trends and enterprise credit in bidding data. Compared with the prior art that only makes market predictions through simple statistical analysis and rule matching, the present invention can accurately predict the future changes in the market and the credit performance of enterprises through the self-optimization of the deep learning model, effectively improving the scientific nature of decision-making and the accuracy of prediction.
[0014] 2. The present invention adopts an automated data processing solution based on AI technology. By integrating web crawlers, deep learning, and natural language processing technologies, it realizes a fully automated process from data collection to report generation, achieving the technical effect of significantly improving data processing efficiency. Compared with the manual operation and traditional rule matching methods in the prior art, the present invention can process a large amount of complex data in a short time, reducing the errors and time delays caused by manual intervention and greatly improving work efficiency.
[0015] 3. The present invention combines homomorphic encryption, federated learning, and blockchain technology to ensure the privacy protection of data during the processing process. Compared with the prior art that stores data centrally or protects data through traditional encryption technologies, the present invention allows operations to be performed on encrypted data through homomorphic encryption technology, and all data processing behaviors are traceable. This not only protects data privacy but also ensures the integrity and security of data, reducing the risk of data leakage.
[0016] 4. Through dynamic dashboards and visualization tools, the present invention supports the customization and interactive functions of reports. Different from the static reports with fixed formats in the prior art, the present invention can dynamically adjust the report content and presentation form according to user needs, providing a more personalized report display. This flexible report generation and display method enables users to quickly obtain accurate analysis information according to different business needs, enhancing the user experience and the practical application value of reports. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of the method steps of the present invention; Figure 2 is a schematic diagram of the modules of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Embodiment
[0019] Please refer to the appendix Figure 1 , an embodiment of the present invention provides a data processing method based on AI technology, including: S1. Data collection: Collect tendering and bidding related data from the target data source through web crawler technology, and store the collected data in a distributed database; S2. Data preprocessing: Clean, standardize, and complete the collected data, remove abnormal data, and complete missing data; S3. Intelligent analysis: Perform pattern recognition, trend prediction, and enterprise credit rating on the preprocessed data based on a deep learning model, and mine the correlation relationships between data; S4. Report generation: Generate a data analysis report based on the intelligent analysis results, and display the analysis results through dynamic visualization; S5. Privacy protection: Encrypt sensitive data during the data processing process, and combine federated learning technology to ensure data security; The data collection step includes: Use a distributed web crawler framework to concurrently collect data from tendering and bidding websites, enterprise databases, social media, and news media data sources; Through a dynamic update scheduling mechanism, adjust the data collection frequency according to the update frequency and content change rate of the data source, so as to optimize the data collection efficiency; Store the collected data in a distributed database to support large-scale data storage and high-concurrency access; The data preprocessing step includes: Clean the collected data using natural language processing technology to remove redundant and invalid data; Detect outliers based on a clustering algorithm, and remove data that exceeds the preset threshold range; Complete missing data through a generative adversarial network, and the generative adversarial network generates data that meets expectations according to the data feature distribution; The intelligent analysis step includes: Perform pattern recognition on historical tendering and bidding data based on a deep neural network model to extract market trend features; Construct an enterprise knowledge graph, with enterprises as nodes and the cooperation and competition relationships between enterprises as edges; Iteratively update the node features of the enterprise association relationship graph based on a graph neural network, and generate a comprehensive feature representation of the enterprise through feature propagation between nodes; Use a time series analysis model to predict the future demand and competition dynamics of the tendering and bidding market; The iterative update of the graph neural network includes the following steps: Encode the initial features of each node; Update the state of each node through feature propagation of neighbor nodes, where the node feature update formula is: ; where, : The feature representation of node i in the k-th layer; : The neighbor set of node i; : The weight matrix of the k-th layer; : The degree of node k; Finally, generate the comprehensive feature representation of the enterprise in the market environment; The report generation steps include: Use natural language generation technology to generate an analysis report, which includes enterprise credit rating, market trend prediction, and competitiveness analysis; Provide a dynamic dashboard to display the analysis results through a visualization tool, supporting users to filter and update the analysis content according to different dimensions; The privacy protection steps include: Dynamically encrypt sensitive data to ensure the security of data during transmission and storage; Train the model in a distributed environment through federated learning technology to avoid sharing raw data; Based on blockchain technology, record the entire process of data processing to achieve data traceability and immutability; The data collection, data preprocessing, intelligent analysis, and report generation steps are all implemented based on a distributed computing architecture to support large-scale data processing and high-concurrency task execution.
[0020] Specifically, in the present invention, the core objective of step S1 is to collect tendering and bidding-related data from multiple data sources through web crawler technology and store the data in a distributed database suitable for further processing. Data collection not only needs to ensure the comprehensiveness and timeliness of data, but also requires efficient and accurate crawling capabilities to meet the needs of large-scale data processing.
[0021] In this embodiment, the data collection module collects data from different data sources through a distributed web crawler framework. The data sources include but are not limited to tendering and bidding websites, enterprise databases, social media, news media, and relevant industry report databases, etc. In actual operation, the system configures a crawler task scheduling system to send requests to the target data sources regularly or in real time to collect tendering and bidding-related data. The system sets different crawling strategies according to the characteristics of different data sources, thus ensuring the breadth and depth of data.
[0022] In a possible implementation, to ensure the timeliness and coverage of data, the system adopts a concurrent collection strategy. This strategy can simultaneously capture data from multiple data sources, effectively avoiding bottleneck problems during the data collection process. By using a distributed crawler framework, multiple crawler instances work in parallel on multiple nodes, thus greatly improving the speed of data crawling. For example, when collecting bidding announcements, the system will initiate crawling tasks from multiple bidding websites simultaneously, not only ensuring the integrity of information but also shortening the time for data collection.
[0023] Specifically, when collecting data, the system will set corresponding crawling rules according to the characteristics of the data sources. For example, for static web pages, the system uses XPath-based parsing technology to extract bidding information from the web pages; for dynamic web pages, the system will use automated testing tools such as Selenium to simulate user operations and capture dynamically loaded data. In this way, the system can adapt to the characteristics of different types of web pages and comprehensively collect bidding-related data.
[0024] Generally, during the data collection process, the system will regularly check the update frequency of each data source to ensure that the latest bidding information is captured. When a data source changes, the system will automatically adjust the data collection frequency according to the dynamic update scheduling mechanism. Specifically, the system determines the collection tasks of each data source dynamically according to the update frequency
[0025] and the content change rate of each data source. This mechanism is scheduled through the following optimization formula: where is the dynamic weight of the i-th data source, which determines the collection priority of this data source at time t;
[0026] is the update frequency of the i-th data source at time t;
[0027] is the rate of content change of the i-th data source at time t; α and β are adjustment factors used to balance the weights of the update frequency and the content change rate.
[0028] This formula enables the system to adjust the collection tasks according to the changes in the data sources, thereby improving the collection efficiency while ensuring the timeliness of the data.
[0027] As an option, to avoid over-collection and the generation of invalid data, the system also introduces a data quality monitoring mechanism. Under this mechanism, the collected data will be preliminarily screened to eliminate records with format errors, invalidity, or duplicates. The screening rules are based on preset data validity criteria (such as data integrity, accuracy, etc.) to ensure the data quality in subsequent processing links.
[0028] In a possible implementation, the system stores the collected data in a distributed database to support large-scale data storage and the execution of high-concurrency tasks. Specifically, the data storage solution adopts a distributed storage framework such as Hadoop or Spark. By storing the data distributively on multiple nodes, it realizes the efficient management and rapid reading of massive data. This approach not only meets the data storage requirements but also provides a convenient storage environment for subsequent processing and analysis.
[0029] Specifically, by storing the data in a distributed system, the system can still maintain a high processing efficiency in a large-data-volume environment. When the data volume increases, the storage system can automatically expand to ensure the high availability and high performance of storage.
[0030] In some embodiments, to further improve the scalability of the system, the data storage module can select an appropriate storage method according to different data types. For example, for structured data (such as enterprise information, bidding history data), the system can use a relational database (such as MySQL, PostgreSQL, etc.); while for unstructured data (such as text data, news reports), a distributed file system (such as HDFS) can be used.
[0031] In this way, the system can ensure the efficient storage and processing of different types of data, providing reliable support for subsequent data preprocessing, intelligent analysis, and other steps. The main purpose of step S2 is to perform processing such as cleaning, standardization, outlier detection, and data completion on the raw data collected from multiple data sources to ensure the accuracy and consistency of the data and provide high-quality input data for subsequent data analysis and decision-making. Data preprocessing not only improves the quality of the data but also effectively avoids biases caused by data problems in subsequent analysis.
[0032] In this embodiment, step S2 is divided into the following main links: data cleaning, outlier detection and processing, data completion, and data standardization. The specific technical details of each link will be described in detail in combination with the technical solution of the present invention.
[0033] First of all, data cleaning is the first step to ensure the smooth progress of subsequent processing. Since the data collected from multiple data sources may be duplicate, invalid, or in the wrong format, redundant and incorrect information must be removed through the cleaning link to improve the quality of the data. Natural language processing (NLP) technology plays a crucial role in this process, especially when processing text data.
[0034] In this embodiment, the first step of data cleaning is to use word segmentation technology to split the text data and break each text paragraph into processable vocabulary units. Next, part-of-speech tagging and named entity recognition (NER) technology are used to identify the meaningful parts of the data. Specifically, named entity recognition technology removes irrelevant text content by identifying information such as company names, time, and location in the text.
[0035] In general, the text data cleaning process also includes deleting stop words, punctuation marks, web page tags and other redundant content to reduce unnecessary information interference. For bidding data containing fields such as date, amount, location, etc., the system will also format them into standardized date and value formats to ensure data consistency.
[0036] As an option, if there are duplicate records in the data source, the system will detect and delete duplicate data by calculating text similarity. Specifically, the system will determine whether two records are duplicates by calculating the Jaccard similarity or Cosine similarity of the text, and delete the duplicate data based on the set threshold. The specific formula is as follows:
[0037] Among them, A and B are word sets of two text data to be compared; is the number of co-occurring words in the text data; is the number of unique words in two text data.
[0038] During the data collection process, due to various reasons, some data may contain abnormal values that deviate from the normal range. Abnormal values will affect subsequent data analysis and lead to inaccurate results, so they need to be detected and processed.
[0039] In this embodiment, outlier detection uses a clustering algorithm, such as K-means clustering, to analyze the data and identify outliers that are significantly different from other data points. The system determines whether each data point is an outlier based on its distance from the cluster center. If a data point is too far from the center of its cluster, it is considered an outlier and will be removed.
[0040] The formula for outlier detection is as follows:
[0041] in, It is a data point With cluster center The Euclidean distance of Represents the distance between two points.
[0042] Specifically, the system calculates the distance of each data point to the nearest cluster center and sets a threshold. When the distance is greater than the threshold, the data point is considered an outlier and is removed from the dataset. The threshold can be adjusted according to the distribution of the dataset and is usually set as the average distance from the data points to the cluster center plus twice the standard deviation.
[0043] In some embodiments, to improve the accuracy of outlier detection, the system may also combine the Local Outlier Factor (LOF) algorithm to identify outliers in multi-dimensional data. The LOF algorithm determines whether a data point is an outlier by analyzing the density of the data point within its local neighborhood. If the local density of a data point is significantly lower than that of its neighbor nodes, the point is considered an outlier.
[0044] For the data missing during the acquisition process, the system uses a Generative Adversarial Network (GAN) to complete it. The training process of the Generative Adversarial Network consists of two parts: a generator and a discriminator. The generator is responsible for generating the completed data, while the discriminator judges the authenticity of the generated data. During the training process, the generator is gradually optimized to generate more realistic data, and the discriminator continuously improves its ability to distinguish between generated data and real data.
[0045] Specifically, the working principle of the Generative Adversarial Network is that the generator generates the missing data from the noise input z and is evaluated by the discriminator. The generator and the discriminator play against each other, and the generator adjusts its weights to generate as realistic data as possible. The formula of the Generative Adversarial Network is as follows:
[0046] Among them, is the completed data generated by the generator; is the generator function; z is the noise vector sampled from the standard normal distribution .
[0047] As an option, during the completion process, the generator not only generates based on the local features of the data but also combines the global features of the data to ensure that the generated data is consistent with the real data in the overall structure. Through multiple rounds of training, the generator can generate higher-quality completed data that is consistent with the features of the original data.
[0048] After data cleaning, outlier detection, and data completion, the system standardizes the data to ensure that the scales of various features are consistent. Usually, standardization includes two methods: min-max normalization and Z-score normalization. Min-max normalization maps the values of the data to a fixed range (e.g., [0,1]), while Z-score normalization makes the mean of the data 0 and the standard deviation 1.
[0049] In a possible implementation, the system first calculates the minimum value of each feature and the maximum value , and then normalizes the data to the range [0, 1] through the following formula:
[0050] Generally, for features with large deviations, the system uses Z-score standardization to calculate the mean μ and standard deviation of each feature , and then converts the data into a standard normal distribution:
[0051] Through the above standardization process, different dimensions and features of the data will become consistent, avoiding analysis biases caused by dimensional differences.
[0052] Step S3 is a key step in deeply analyzing the data that has been cleaned, standardized, and completed. This step uses various artificial intelligence technologies, such as deep learning, graph neural networks (GNNs), and time series analysis models, combines different data features, extracts important patterns in the data, and conducts market trend prediction, enterprise credit assessment, and association mining between data. The results of the intelligent analysis will provide the core basis for subsequent report generation and decision support.
[0053] The purpose of this step is to enable the data to self-learn through deep learning models and graph neural networks, extract potential patterns in the data, and at the same time predict market trends through time series analysis techniques. This process improves the accuracy of the analysis by leveraging the features obtained from the data and generates accurate market insights. The following details the specific implementation process of step S3.
[0054] In this embodiment, the intelligent analysis module first performs pattern recognition on the bidding data through a deep neural network (DNN). The training process of the deep neural network is based on a large amount of historical bidding data. Through a multi-layer network structure, the system can extract the key factors affecting the success of bidding from it and predict market trends based on these factors.
[0055] Specifically, the structure of a deep neural network usually includes multiple hidden layers, and each layer of neurons is fully connected to the neurons in the previous and next layers. Through iterative training, the network can gradually adjust the parameters to minimize the error, enabling the model to learn the potential laws in the data. The output y of the model is usually calculated from the input data X and the model parameters through the following formula:
[0056] Among them, X is the input feature data, representing various data related to bidding (such as enterprise information, market characteristics, etc.); are the parameters of the model, including the weight matrix and the bias term ; is an activation function, usually ReLU or Sigmoid, which is used to enhance the nonlinear expression ability of the model; y is the output of the model, which represents the predicted market trend or corporate credit score.
[0057] Specifically, deep neural networks can automatically extract important features from bidding data. For example, by learning from historical data, the system can identify which corporate characteristics and market factors have the greatest impact on the probability of winning a bid, thereby providing scientific decision-making support for companies or governments.
[0058] As an option, in order to improve the expressiveness of the model, the system can use a network architecture such as a convolutional neural network (CNN) to process data with spatial features (such as patterns in image data or text data). This approach can help the model better understand the local and global features in the data and improve the accuracy of the analysis.
[0059] In one possible implementation, the intelligent analysis step further introduces a graph neural network (GNN) to enhance the ability to mine relationships between data. GNN models the complex relationships between enterprises by viewing enterprises as nodes of a graph and the cooperative or competitive relationships between enterprises as edges of the graph. In this network, the feature representation of a node is updated through feature propagation of neighboring nodes, thereby generating a comprehensive feature representation of each node.
[0060] The feature propagation formula in GNN can be expressed as:
[0061] in, is the feature representation of node i in the kth layer; is the set of neighbor nodes of node i; and is the degree of node i and node j (i.e. the number of connections between each node); is the activation function; is the weight matrix of the kth layer.
[0062] Through this formula, the graph neural network transmits and updates information between nodes, thereby generating comprehensive characteristics of each enterprise, including its competitive and cooperative relationships with other enterprises. Ultimately, based on these characteristics, the system can generate market competitiveness scores and credit ratings for each enterprise.
[0063] In general, to further improve the prediction effect, the system may combine time series analysis to dynamically predict market demand and trends. Time series analysis can capture the time dependence in the data and predict future trends based on past market data. For example, the system uses Long Short-Term Memory (LSTM) networks or Transformer models for time series analysis, and these models can effectively process time series data with long-term dependencies.
[0064] The formula of the LSTM model is usually expressed as:
[0065] Where is the hidden state at time t; is the input data at time t (such as market data, historical bidding results, etc.); is the hidden state of the previous time; are the parameters of the model.
[0066] The LSTM network effectively retains long-term dependency information through a gating mechanism, and processes long-term trends and short-term fluctuations, providing a reliable basis for future predictions of the market. In this way, the system can predict key indicators such as future market demand and the probability of an enterprise winning a bid based on historical data, helping decision-makers formulate effective strategies.
[0067] As an option, if the time series characteristics of the bidding data are relatively complex, the system can adopt the Transformer model, which is a deep learning model based on self-attention mechanism, capable of processing longer time series data and capturing the global dependencies between data.
[0068] In some embodiments, to further improve the comprehensive effect of the model, the system may combine multiple algorithms and adopt an ensemble learning method to weight and aggregate the results of different models to obtain the final prediction result. Through this method, the advantages of multiple models can be complementary, thus providing more accurate prediction and analysis results.
[0069] The main objective of step S4 is to automatically generate a structured market analysis report based on the results of intelligent analysis and display it to the user in an easy-to-understand manner. Report generation not only includes the summary and insights of the data, but also should provide visualization tools to help users better understand the analysis results and make decisions.
[0070] Step S4 is closely connected to the aforementioned steps. In particular, the results output by the intelligent analysis module provide the basic data for report generation. These analysis results include market trends, enterprise credit scores, competitiveness analysis, etc. The system will automatically compile a report based on these results and conduct visual display. Through this link, the system can present complex data analysis to the end users in a concise and easy-to-interpret form.
[0071] In this embodiment, the report generation module first receives the data output by the intelligent analysis module. These data include enterprise credit scores, market trend forecasts, industry competitiveness analysis, etc. The system automatically organizes and generates a detailed analysis report based on these data. The generated report is presented not only in the form of text descriptions but also includes data tables, charts, and other visual elements to help users understand the analysis results from multiple perspectives.
[0072] Specifically, the generation of the report first uses natural language generation (NLG) technology to automatically generate text content. The NLG technology generates logical text based on the results of intelligent analysis, describing the enterprise's credit rating, market trends, competitive landscape, etc. in natural language form. For example, in the enterprise credit rating section, the system will generate relevant descriptions based on the enterprise's historical data, market behavior, and the scoring model of the intelligent analysis module. The content of the description includes the enterprise's credit status, past bidding history, and possible market risks.
[0073] Generally, based on the results of intelligent analysis, the NLG technology will generate text similar to the following: "The credit rating of Enterprise A is B, indicating that it has shown a relatively stable credit record in past bidding processes. According to the market trend forecast, Enterprise A has the potential to improve its credit rating, and it is expected that its winning bid probability will increase within the next 12 months." As an option, if the requirements for report generation are more complex, the system also supports personalized customization of the report content according to user requests. Through user profiling and requirement modeling, the system can generate reports with different focuses for different users. For example, some users may be more concerned about the enterprise's credit situation, while others may be concerned about market trends or competitiveness analysis.
[0074] In a possible implementation, to improve the readability of the report, the system integrates a dynamic dashboard. This dashboard can not only display the analysis results but also provide interactive functions, allowing users to filter data according to different dimensions and view the corresponding analysis results. Users can select different time ranges, industry categories, or regions according to their needs to obtain personalized market analysis. The visual elements such as charts and tables in the report will also be dynamically updated according to the user's selection.
[0075] Specifically, during the report generation process, the system generates charts through visualization tools. For example, in the market trend section, the system may display a time series data chart of bidding activities, clearly presenting the changes in market demand over the past few months. Meanwhile, the competitiveness analysis may show the market performance of multiple enterprises in the form of bar charts or radar charts, providing intuitive comparison data for users.
[0076] In some embodiments, to further enhance the practicality of the report, the system can also provide additional decision support functions for each report. For example, the report can automatically provide suggestions for future market changes to help enterprises or government departments make more informed decisions. In this way, the report is not just a data summary tool but a decision support platform with strong practical value.
[0077] As another option, if the user needs to view the analysis results of a specific part in detail, the system provides an interactive report generation function. The user can click on a certain part of the report, and the system will further expand the detailed analysis content of that part. For example, after the user clicks on the credit score part of a certain enterprise, the system will display the detailed historical bidding data of that enterprise and provide corresponding analysis basis.
[0078] In a possible implementation, in addition to the text report, the system also provides an export function that allows users to export the report in PDF or Excel format. In this way, users can conveniently save the report and view it without an Internet connection. The content of the exported report will include all charts, data, and relevant analysis text, ensuring that users can access the complete analysis results offline.
[0079] Generally, the report generation time is very short, almost real-time. Through automated generation and interactive functions, users can quickly obtain the required market insights and analysis results, greatly improving the decision-making efficiency.
[0080] Step S5 focuses on protecting data privacy through various technical means when storing, processing, and transmitting data. In particular, it combines homomorphic encryption, federated learning, and blockchain technologies to ensure that sensitive information is fully protected throughout the data analysis and processing process.
[0081] This step is closely related to the previous data preprocessing and intelligent analysis links. Each link of data cleaning and analysis may involve the processing of sensitive information, and privacy protection technology is the key to ensuring that information is not leaked during these links.
[0082] In this embodiment, privacy protection is first achieved through homomorphic encryption technology. Different from traditional encryption technologies, homomorphic encryption technology allows certain operations to be performed on encrypted data, and the operation results after decryption are the same as those obtained by directly operating on plaintext data. Therefore, the system can process and analyze data without decrypting it, thus avoiding data exposure to unauthorized personnel.
[0083] The basic idea of homomorphic encryption is that after encrypting data, necessary mathematical operations can still be performed on the encrypted data, and the decryption of the calculation results will obtain the correct plaintext results. For example, for some numerical data, the system can perform operations such as addition or multiplication in the encrypted state to ensure the privacy of the data.
[0084] Specifically, the addition operation of homomorphic encryption technology is as follows:
[0085] Among them, and are the encrypted forms of data and respectively; is the homomorphic encryption operator, indicating that an addition operation is performed on the encrypted data.
[0086] Through this formula, users can calculate data without exposing the data content, and the calculation results remain encrypted until they need to be decrypted to obtain the correct results.
[0087] Generally, homomorphic encryption is used to protect data privacy without affecting subsequent analysis and calculations. Especially in cloud computing and distributed systems, dynamic encryption is widely used to protect sensitive data and prevent data leakage during transmission and processing. In the application scenarios of the present invention, sensitive data such as bidding data, enterprise information, and credit scores can all be processed through dynamic encryption to ensure data security.
[0088] As an option, to further enhance the privacy protection effect, the system combines federated learning technology. Federated learning is a method of distributed learning that allows multiple participating parties to jointly train a machine learning model without exchanging data. Each participating party only trains its own model locally and uploads the updated model parameters calculated locally to the central server. The server is responsible for aggregating the updates from each participating party to form a global model.
[0089] Specifically, the process of federated learning is distributed, and the original data of each data holder always remains local. In each round of training, the participants calculate gradients or updates based on their own datasets and send this information to the server without exchanging the original data. The server then performs a weighted average of all the update results to obtain the global model update result. This process can be represented by the following formula:
[0090] where, are the global model parameters; is the local model parameter of the th participant after the th round of training; is the data volume of the th participant; is the total data volume of all participants.
[0091] Through federated learning, the system can utilize the characteristics of different data sources for model training while ensuring data privacy, thereby obtaining more accurate market analysis and enterprise credit assessment.
[0092] In some embodiments, to enhance data security, the system also incorporates blockchain technology. Blockchain provides a decentralized data storage and verification mechanism, ensuring the immutability and traceability of data during storage and use. During the data processing and analysis process, all data operations (such as data reading, updating, analysis, etc.) will be recorded on the blockchain. Each operation is stored in a distributed ledger, and once the data is written to the blockchain, any unauthorized operation cannot modify this data.
[0093] Specifically, each block in the blockchain contains a timestamp and a hash value pointing to the previous block. In this way, once the data is tampered with, the entire blockchain structure will be damaged, thereby promptly detecting the data tampering behavior. Therefore, blockchain can effectively prevent data from being maliciously modified or deleted, ensuring the integrity and reliability of the data. The formula representation of blockchain technology is as follows:
[0094] where, is the hash value of the i-th block; is the hash value of the previous block; is the data stored in block i; is the timestamp of block i.
[0095] In this way, blockchain ensures the transparency of all data processing and calculations, enhances data security, and enables each data operation to be traceable.
[0096] In general, by combining homomorphic encryption, federated learning, and blockchain technology, the system can provide security for each data processing link and ensure the privacy of data during the processing. Especially in the process of bid evaluation data analysis involving multiple parties, privacy protection technology provides strong support for the credibility and legality of the system.
[0097] As an option, in terms of data access, the system can also combine permission management technology to further strengthen privacy protection. By managing different user roles and permissions, it is ensured that only authorized personnel can access specific types of data or perform specific data operations. For example, the system can be set so that only authenticated users can access specific corporate financial data or can only perform model training in an authorized environment.
[0098] In a possible implementation, each operation of data access will be recorded and stored to ensure the transparency and compliance of the operation. By integrating audit logs and monitoring functions, the system can detect and track data access behaviors in real time, ensure the compliance of the data processing process, and promptly discover potential security threats.
[0099] Embodiment 2: Please refer to the appendix Figure 2 , a data processing system based on AI technology, including: A data collection module, used to collect bid evaluation data from the target data source through distributed web crawler technology; A data preprocessing module, used to clean, standardize, and complete the collected data; An intelligent analysis module, used to perform pattern recognition, trend prediction, and enterprise credit rating based on a deep learning model; A report generation module, used to generate an analysis report including enterprise credit rating and market trend prediction, and display the analysis results through a dynamic visualization tool; A privacy protection module, used to encrypt sensitive data and combine federated learning and blockchain technology to protect data privacy.
[0100] Specifically, in this embodiment, first, the data collection module collects relevant bid evaluation data from various data sources (such as bid evaluation websites, social media, news media, etc.) through distributed web crawler technology. The collected data will be initially stored and continuously updated by the system to ensure the timeliness and integrity of the data. The collected raw data is then transferred to the data preprocessing module for further processing.
[0101] The data preprocessing module cleans, standardizes, and fills in missing data for the collected raw data. Data cleaning includes removing irrelevant information, duplicate data, and data with format errors. Standardization converts the data into a unified format and uses natural language processing techniques to parse text data. Missing data is filled in using the Generative Adversarial Network (GAN) method to ensure the integrity and usability of the data. The preprocessed data will be transmitted to the intelligent analysis module as the basis for analysis.
[0102] Next, the intelligent analysis module performs deep learning and pattern recognition based on the preprocessed data. This module first extracts key features from historical bidding and tendering data through a Deep Neural Network (DNN) to identify the main factors affecting market changes. Then, it constructs an association network between enterprises through a Graph Neural Network (GNN) to analyze the competitive relationships and credit status of enterprises in the market. At the same time, the intelligent analysis module also uses time series analysis to predict market trends and evaluate future market demand and competition situations. This process ensures a comprehensive analysis of the data and generates analytical results with predictive capabilities.
[0103] The analysis results will be passed to the report generation module, which automatically generates a structured analysis report based on Natural Language Generation (NLG) technology. The report not only includes text descriptions but also displays key information such as market trends and enterprise credit ratings through charts and dynamic visualization tools. The report generation module supports personalized customization of reports, generating reports with different dimensions according to the needs of different users. For example, government departments focus more on market trends and policy impacts, while enterprises are more concerned with competitiveness evaluation and risk analysis. In addition, users can further filter and analyze the data through an interactive dashboard to customize the content of their reports.
[0104] Finally, the privacy protection module plays a crucial role throughout the data processing process. Whether it is data collection, storage, analysis, or report generation, the privacy protection module uses technologies such as homomorphic encryption, federated learning, and blockchain to ensure the privacy and security of the data. During data storage and calculation, homomorphic encryption enables encrypted data to be calculated without revealing the data content. Federated learning allows different data holders to jointly train models, avoiding data leakage. Blockchain technology ensures the transparency and immutability of all data operations, further enhancing the security and compliance of the system.
[0105] Through the collaborative work of these modules, the system realizes the full - process automation of data from collection, preprocessing to analysis and report generation, and ensures the privacy and security of the data in each link, ultimately providing users with accurate market analysis and decision - making support.
[0106] Example Three: A computer-readable storage medium stores a computer program which, when executed by a computer, implements a data processing method based on AI technology according to any one of claims 1-8.
[0107] Specifically, in this embodiment, the storage medium part is implemented by a computer-readable storage medium which stores a computer program. When this program is executed by a computer, it can implement the method in the above steps. In this embodiment, the storage medium can be various types of storage devices, such as hard disks, solid-state drives, optical discs, flash memories, cloud storages or any storage medium with data storage functions. The computer program includes the implementation logics of a data acquisition module, a data preprocessing module, an intelligent analysis module, a report generation module and a privacy protection module. When these programs are executed on a computer or a server, they complete the entire data processing process from data acquisition to report generation in sequence.
[0108] Specifically, the computer program stored in the storage medium first guides the data acquisition module to scrape bidding data from various data sources through web crawler technology and store it in a distributed database. The data preprocessing module then extracts valid information from the stored raw data, performs cleaning, standardization and completion to ensure the quality and integrity of the data. The processed data will be passed to the intelligent analysis module which analyzes the data through models such as deep neural networks and graph neural networks for pattern recognition, trend prediction and credit assessment, etc.
[0109] In the report generation stage, the computer program will automatically use natural language generation technology to convert the intelligent analysis results into a structured report and generate corresponding charts and data displays through visualization tools. These reports are stored and managed through the storage medium for users to access at any time. The privacy protection module ensures the privacy and security of all data during the entire data processing process, and uses dynamic encryption, federated learning and blockchain technology to encrypt and protect the data.
[0110] In some embodiments, the storage medium is not just a simple storage device, but may also have intelligent storage management functions. For example, in a cloud storage system, data may be stored distributively and automatically synchronized between different nodes to ensure high availability and data redundancy. The computer program in the storage medium can be dynamically scheduled and updated according to system requirements to adapt to the increasing data processing needs and achieve efficient management of large-scale data sets.
[0111] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data processing method based on AI technology, characterized in that, Including: S1, Data collection: Collect tendering and bidding related data from the target data source through web crawler technology, and store the collected data in a distributed database; S2, Data preprocessing: Clean, standardize and complete the collected data, remove abnormal data and complete missing data; S3, Intelligent analysis: Based on deep learning models, perform pattern recognition, trend prediction and enterprise credit rating on the preprocessed data, and mine the correlation relationships between data; S4, Report generation: Generate a data analysis report based on the intelligent analysis results, and display the analysis results through dynamic visualization; S5, Privacy protection: Encrypt sensitive data during the data processing process, and combine federated learning technology to ensure data security.
2. The data processing method based on AI technology according to claim 1, wherein, The data collection step includes: Use a distributed web crawler framework to perform concurrent collection on tendering and bidding websites, enterprise databases, social media and news media data sources; Through a dynamic update scheduling mechanism, adjust the data collection frequency according to the update frequency and content change rate of the data source, so as to optimize the data collection efficiency; Store the collected data in a distributed database to support large-scale data storage and high-concurrency access.
3. The data processing method based on AI technology according to claim 1, characterized in that The data preprocessing step includes: Clean the collected data using natural language processing technology to remove redundant and invalid data; Detect outliers based on clustering algorithms, and remove data beyond the preset threshold range; Complete missing data through a generative adversarial network, and the generative adversarial network generates data that meets expectations according to the data feature distribution.
4. The data processing method based on AI technology according to claim 1, wherein The intelligent analysis step includes: Perform pattern recognition on historical tendering and bidding data based on a deep neural network model, and extract market trend features; Construct an enterprise knowledge graph, with enterprises as nodes and the cooperation and competition relationships between enterprises as edges; Based on a graph neural network, iteratively update the node features of the enterprise association relationship graph, and generate a comprehensive feature representation of the enterprise through feature propagation between nodes; Use a time series analysis model to predict the future demand and competition dynamics of the tendering and bidding market.
5. A data processing method based on AI technology according to claim 4, characterized in that, The iterative update of the graph neural network includes the following steps: Encode the initial features of each node; Update the state of each node through feature propagation of neighbor nodes, where the node feature update formula is: ; Among them, : The feature representation of node i in the k-th layer; : The neighbor set of node i; : The weight matrix of the k-th layer; : The degree of node k; Finally generate a comprehensive feature representation of the enterprise in the market environment.
6. The data processing method based on AI technology according to claim 1, wherein The report generation step includes: Use natural language generation technology to generate an analysis report, and the analysis report includes enterprise credit rating, market trend prediction and competitiveness analysis; Provide a dynamic dashboard, display the analysis results through visualization tools, and support users to filter and update the analysis content according to different dimensions.
7. The data processing method based on AI technology according to claim 1, characterized in that, The privacy protection step includes: Dynamically encrypt sensitive data to ensure the security of data during transmission and storage; Train models in a distributed environment through federated learning technology to avoid sharing raw data; Record the entire data processing process based on blockchain technology to achieve data traceability and immutability.
8. A data processing method based on AI technology according to claim 1, characterized in that, The data collection, data preprocessing, intelligent analysis and report generation steps are all implemented based on a distributed computing architecture to support large-scale data processing and high-concurrency task execution.
9. A data processing system based on AI technology, according to any one of claims 1-8, a data processing method based on AI technology, characterized in that, Including: A data acquisition module for collecting bidding data from a target data source through distributed web crawler technology; A data preprocessing module for cleaning, standardizing, and completing the collected data; An intelligent analysis module for pattern recognition, trend prediction, and enterprise credit rating based on a deep learning model; A report generation module for generating an analysis report including enterprise credit rating and market trend prediction, and displaying the analysis results through a dynamic visualization tool; A privacy protection module for encrypting sensitive data and protecting data privacy by combining federated learning and blockchain technology.
10. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a computer, it implements a data processing method according to any one of claims 1-8 based on AI technology.