Big data enterprise environment public service system
The big data enterprise environmental public service system has solved the problems of untimely information updates and single data sources on existing platforms, enabling timely updates of enterprise environmental information and multi-source data collection, thereby improving the sustainable development and environmental protection level of enterprises.
Patent Information
- Application Number
- CN202411877011.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing corporate environmental information service platforms suffer from untimely information updates, limited data sources, and a lack of targeted services, failing to meet society's demand for corporate environmental information.
This invention provides a public service system for big data enterprise environments, including a data acquisition module, a data processing module, a user interface module, and a website function module. Through web crawling algorithms, multi-source API calls, real-time monitoring, and machine learning models, it dynamically adjusts the priority and frequency of data acquisition, providing comprehensive and convenient enterprise environment information management.
It enables timely updates of enterprise environmental information and multi-source data collection, providing a comprehensive, convenient, and reliable public service platform to promote sustainable enterprise development and ecological environmental protection.
Smart Images

Figure CN119809897B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a big data enterprise environment public service system. BACKGROUND
[0002] In the prior art, with the rapid development of industrialization, environmental pollution problems are increasingly serious, and the demand for enterprise environmental information of all sectors of society is increasing. The existing enterprise environmental information service platform has many deficiencies, such as information update not timely, single data source, lack of targeted services, etc. SUMMARY
[0003] The present application aims at the defects of the prior art, and provides a big data enterprise environment public service system to solve the problems in the prior art.
[0004] To achieve the above-mentioned purpose, the present application provides a big data enterprise environment public service system, comprising:
[0005] A data acquisition module is configured to process unstructured data through a web scraping algorithm and a multi-source API calling mechanism to obtain enterprise environmental information.
[0006] A data processing module is configured to dynamically adjust the task priority and collection frequency of data acquisition by monitoring the changes of data sources and user query records in real time.
[0007] A user interface module is configured to receive request information from multiple users and return corresponding results according to the enterprise environmental information. The request information includes API calling requests through a web browser, a mobile application, a desktop client, and an API calling request.
[0008] A website function module is configured to receive request information sent by the user interface module and perform certificate query, enterprise evaluation, and points exchange according to the web browser, mobile application, desktop client, API calling request, and enterprise environmental information.
[0009] In a possible implementation manner, the processing of unstructured data to obtain enterprise environmental information specifically includes:
[0010] An enterprise information is obtained from an enterprise website, a public database, or social media through a web crawler tool.
[0011] Environment monitoring data and enterprise annual reports are obtained by calling an API interface provided by an enterprise.
[0012] File information uploaded by a user is obtained.
[0013] The obtained information is preprocessed, and the preprocessing includes text cleaning, image processing, and audio-to-text processing.
[0014] Identify key entities in the text using NLP tools, determine the positive or negative nature of the enterprise environmental information according to the sentiment tendency of the text information; extract keywords and phrases in the text;
[0015] Extract text information from images using OCR technology;
[0016] Use deep learning models to identify specific objects in images, including pollution sources and environmental protection facilities;
[0017] Use table parsing tools to extract table data from PDF and images;
[0018] Store the processed data in a relational database and create an index for the data to generate a big data index library.
[0019] In one possible implementation, the user interface module is also used to define different priority labels for the first type of data and the second type of data, and to perform sorting processing before sending; the priority of the first type of data is higher than that of the second type of data; wherein the first type of data is data with strong real-time performance; the second type of data is data with low delay sensitivity.
[0020] In one possible implementation, the user interface module is also used to adjust the sending rate according to the network feedback traffic and according to an adaptive flow control algorithm; wherein the adaptive flow control algorithm includes the slow start mechanism and the congestion avoidance mechanism of TCP.
[0021] In one possible implementation, the data processing module is also used to dynamically adjust the caching strategy according to the request frequency of the user request and the update frequency of the data; the caching strategy includes keeping data with high frequency requests in the cache and removing data with low frequency requests from the cache.
[0022] In one possible implementation, the data processing module is specifically used to:
[0023] Set up a message queue to monitor changes in the data source, receive notification messages sent by the message queue when the data source is updated, or configure Webhooks of the data source provider, receive HTTP request notifications sent by the data source provider when the data changes, or set up a timing task to periodically check the status of the data source and trigger the data collection task when there is a change;
[0024] Record each query request of the user, including query time, query content, and query result; record the query history of the user in the database, including query keywords and query frequency; push the query record of the user to the message queue;
[0025] According to the change frequency of the data source, the collection frequency is dynamically adjusted; wherein the data source with high change frequency has high priority and high collection frequency;
[0026] According to the keywords and frequency in the user query record, the priority of the related data source is dynamically adjusted; the data source with high user query frequency has high priority and high collection frequency;
[0027] The machine learning model is trained to predict the change trend of the data source and the user query trend, and the task priority and the collection frequency are adjusted according to the prediction result;
[0028] A rule is preset, and the task priority and the collection frequency are dynamically adjusted according to the rule.
[0029] In a possible implementation, the website function module includes a production enterprise management and a dealer management submodule, which are used to receive information of the production enterprise and information of the dealer; and when a certificate query is received, a certificate confirmation message is generated according to the information of the production enterprise and the information of the dealer.
[0030] The integral module is used to obtain the integral through the activities such as evaluation and sharing, and the integral is used to exchange the goods or the coupon;
[0031] The enterprise evaluation module receives the evaluation information of the enterprise sent by the user terminal;
[0032] The social module receives the green product information shared by the user terminal.
[0033] By applying the big data enterprise environment public service system provided by the embodiment of the present application, a comprehensive, convenient and reliable enterprise environment public service platform is provided, and through comprehensive and accurate enterprise environment information management, the sustainable development of enterprises can be promoted, the ecological environment can be protected, and the overall environmental protection level of the society can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 The big data enterprise environment public service system provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0036] The technical scheme of the present application will be further described in detail below with reference to the drawings and embodiments.
[0037] Figure 1 The structure diagram of the big data enterprise environment public service system provided by the embodiment of the present application is shown below. Figure 1 The technical solutions of the present application are described in specific embodiments. As shown in the figure, the big data enterprise environment public service system includes a data acquisition module 1, a data processing module 2, a user interface module 3, and a website function module 4. Figure 1
[0038] The data acquisition module 1 is used to process unstructured data through web scraping algorithms and multi-source API calling mechanisms to obtain enterprise environment information.
[0039] Among them, the enterprise environment information refers to various information related to environmental protection generated by enterprises in the process of production and operation. These information covers the impact of enterprises on the environment, environmental management measures, environmental performance, etc., and is an important basis for evaluating the environmental protection performance and sustainable development ability of enterprises. Enterprise environmental information includes emission information, energy consumption information, resource utilization information, and environmental management information. Emission information includes air emissions, including waste gas emissions, pollutant types (such as sulfur dioxide, nitrogen oxides, particulate matter, etc.) and their concentrations, water emissions, including wastewater emissions, pollutant types (such as chemical oxygen demand, ammonia nitrogen, heavy metals, etc.) and their concentrations, solid waste emissions, including solid waste generation, types, and treatment methods (such as recycling, landfilling, incineration, etc.). Energy consumption information includes power consumption, including enterprise electricity consumption and its source (such as renewable energy, fossil fuels, etc.), fuel consumption, enterprise use of various fuels (such as coal, natural gas, diesel, etc.), and their purposes, heat consumption, including enterprise use of heat (such as steam, hot water, etc.) and its source. Resource utilization information includes water resource utilization, including enterprise water consumption and its use (such as production water, domestic water, etc.), raw material utilization: types and amounts of main raw materials used by enterprises, and renewable resource utilization: conditions of renewable resources (such as waste paper, waste plastic, etc.) used by enterprises. Environmental management information includes environmental management system: whether the enterprise has established an environmental management system (such as ISO 14001) and its operation, environmental policy: including the enterprise's environmental protection policy and target, environmental training: including the situation of employees receiving environmental training, including emergency plan: the emergency plan of the enterprise to deal with sudden environmental events and its drilling situation. Environmental performance includes: environmental performance indicators: environmental performance indicators set by the enterprise and their achievement, such as emission reduction targets, energy consumption targets, etc., environmental audit: results of environmental audit conducted by the enterprise and its rectification measures, environmental report: annual environmental report published by the enterprise, including environmental performance, environmental management measures, etc.
[0040] Specifically, the enterprise information is obtained from the enterprise website, public database or social media through a web crawler tool; for example, Scrapy, BeautifulSoup, etc. are used to capture relevant information from enterprise websites, government public databases, social media and other sources.
[0041] The environmental monitoring data and enterprise annual reports are obtained by calling the API interface provided by the enterprise.
[0042] The file information uploaded by the user is obtained; the file information includes but is not limited to PDF, Word;
[0043] The obtained information is preprocessed; the preprocessing includes text cleaning, image processing and audio-to-text processing; wherein the data preprocessing includes text cleaning, removing HTML tags, special characters, stop words, etc. and retaining useful information. Image processing: using image processing libraries (such as OpenCV) to crop, scale, denoise, etc. Audio-to-text: using speech recognition technology such as Speech-to-Text API to convert audio to text.
[0044] The key entities in the text are identified using NLP tools, and the positive or negative nature of the enterprise environmental information is judged according to the sentiment tendency of the text information; the key words and phrases in the text are extracted; wherein the entity recognition is to identify the key entities in the text using NLP tools (such as spaCy, Stanford NER), such as enterprise name, date, location, etc. Sentiment analysis is to analyze the sentiment tendency of the text to judge the positive or negative nature of the enterprise environmental information.
[0045] OCR technology is used to extract text information from images; such as extracting text information from images through Tesseract;
[0046] A deep learning model is used to identify specific objects in the image, including pollution sources and environmental protection facilities.
[0047] A table parsing tool is used to extract table data from PDF and images;
[0048] The extracted data is converted into a unified format such as CSV, JSON, etc. The units, date formats, etc. in the data are unified to ensure the consistency and comparability of the data, and the duplicate data records are removed to avoid redundancy.
[0049] The processed data is stored in a relational database, and an index is created for the data to generate a big data index library.
[0050] Subsequently, statistical data such as mean, maximum, minimum, etc. of each index can be calculated. Using time series analysis methods, the changing trend of enterprise environmental information is analyzed. Using data visualization tools (such as Tableau, Power BI, Matplotlib) to display data in the form of charts, it is convenient for users to understand and analyze.
[0051] The data processing module 2 dynamically adjusts the task priority and collection frequency of data collection by monitoring the changes of the data source and the user query records in real time.
[0052] Further, according to the request frequency of user requests and the update frequency of data, the cache strategy is dynamically adjusted; the cache strategy includes keeping the data of high-frequency requests in the cache and removing the data of low-frequency requests from the cache.
[0053] The data processing module 2 is specifically used for:
[0054] Set up a message queue to monitor the changes of the data source, when the data source has updates, the receiving message queue will send a notification message, or configure the Webhooks of the data source provider, when the data changes, receive the HTTP request notification sent by the data source provider; Or, set up a timing task to check the status of the data source regularly, and trigger the data collection task when there is a change;
[0055] Record each query request of the user, the query request including query time, query content, query result; record the query history of the user in the database, the query history including query keyword, query frequency; push the query record of the user to the message queue;
[0056] According to the change frequency of the data source, dynamically adjust the collection frequency; wherein the data source with high change frequency has high priority and high collection frequency;
[0057] According to the keywords and frequency in the user query record, dynamically adjust the priority of the related data source; the data source with high user query frequency has high priority and high collection frequency;
[0058] Train a machine learning model to predict the change trend of the data source and the user query trend, and adjust the task priority and collection frequency according to the prediction result;
[0059] A rule is preset, and the task priority and collection frequency are dynamically adjusted according to the rule. For example, if a certain data source has changed multiple times in the past 24 hours and the user query frequency is high, the task priority and collection frequency of the data source are increased.
[0060] A user interface module 3 for receiving request information from a plurality of users and returning corresponding results according to enterprise environment information; the request information includes requests through a web browser, a mobile application, a desktop client, an API call;
[0061] The request message can be a request message of a dealer terminal, an enterprise terminal, or a public user terminal.
[0062] Specifically, the user interface module uses a binary format instead of a traditional text format (such as JSON). At the same time, a compression algorithm can be introduced to compress the message body, further reducing the amount of data transmitted. Reducing the size of the data packet reduces the parsing overhead.
[0063] Further, the user interface module is the basic architecture of a high-performance serialization / deserialization library, including core interfaces and data structures. Any data object is serialized into a byte stream, and then the byte stream is deserialized into the original data object.
[0064] Define common serialization formats such as binary format, JSON, Protocol Buffers, etc. Support extension of custom data types. Use SIMD instruction sets to optimize performance. Modern processors support SIMD (Single Instruction Multiple Data) instruction sets, which can process multiple data points in a single instruction, greatly improving performance. Common SIMD instruction sets include Intel's SSE, AVX, and ARM's NEON.
[0065] Further, borrowing from the multiplexing mechanism of HTTP / 2, multiple data streams are parallelized on one TCP connection. This can effectively avoid the "head-of-line blocking" problem and improve concurrent processing capability. Thus, the present application can transmit multiple requests or responses simultaneously on one connection, reducing the time cost of establishing and disconnecting connections.
[0066] Specifically, each request or response is treated as an independent data stream. The data stream is divided into smaller units, called frames. Frame header: contains frame type, length, stream identifier, etc. Frame data: actual data content. Reduce the amount of header information transmitted to improve transmission efficiency. The client and server establish a TCP connection. Stream identifier: each data stream has a unique identifier. Priorities can be set for data streams to ensure that critical data is transmitted first. Window update frame controls the transmission rate of the data stream to avoid congestion.
[0067] In one example, session initialization is first performed, the client sends an initial frame (such as setting the maximum frame size, the maximum number of concurrent streams, etc.), and the server replies with an acknowledgement frame. The client and the server can send multiple data streams concurrently, each data stream consisting of multiple frames. Errors in processing a single data stream do not affect other data streams. Errors in processing the entire connection can require re-establishment of the connection.
[0068] Further, the user interface module 3 is further configured to define different priority tags for the first type of data and the second type of data, and perform sorting processing before sending; the priority of the first type of data is higher than the priority of the second type of data; wherein the first type of data is real-time data; the second type of data is delay-sensitive data.
[0069] Further, the user interface module 3 is further configured to adjust the sending rate according to the network feedback flow and according to an adaptive flow control algorithm; wherein the adaptive flow control algorithm includes a slow start mechanism of TCP, a congestion avoidance mechanism.
[0070] The slow start mechanism is to rapidly increase the sending rate at the initial stage until network congestion is detected. The congestion avoidance mechanism is to gradually increase the sending rate after detecting network congestion to avoid congestion again. The fast retransmission mechanism is to immediately retransmit the lost data segment when three repeated ACKs are received. The fast recovery mechanism is to gradually increase the sending rate after fast retransmission to recover to the level before congestion. In one example, the parameters are first initialized, the initial congestion window (CWND) is set to 1 MSS (Maximum Segment Size), the slow start threshold (SSTHRESH) is set to a large value such as 64K, and the receiver window (RWND) is initialized to the maximum receiving window of the receiver.
[0071] For the slow start mechanism, the congestion window is increased by 1 MSS for each received ACK. The congestion window is checked: if the congestion window is greater than the slow start threshold, the congestion avoidance stage is entered.
[0072] For the congestion avoidance mechanism, the congestion window is increased: for each received ACK, the congestion window is increased by 1 / CWND MSS. The congestion window is checked: if the congestion window reaches the receiver window or the network congestion flag is triggered, the congestion window is reduced.
[0073] For the fast retransmission mechanism, when three repeated ACKs are received, the lost data segment is immediately retransmitted.
[0074] For fast recovery mechanism, the slow start threshold is set to half of the congestion window, and then the congestion window is set to the slow start threshold, and the congestion window is continued to be increased.
[0075] Then a prediction model is established. Network status information such as RTT (Round-Trip Time), packet loss rate, bandwidth, etc. is collected regularly. Machine learning models such as linear regression, decision tree, neural network, etc. are used to predict future network status. According to the prediction result, the sending rate is adjusted in advance to avoid congestion.
[0076] The website function module 4 is used to receive the request information sent by the user interface module, and to perform certificate query, enterprise evaluation and points exchange according to the Web browser, mobile application, desktop client, API call request and enterprise environment information.
[0077] The website function module 4 includes a production enterprise management and a dealer management submodule, which is used to receive information of the production enterprise and information of the dealer; and when receiving the certificate query, a certificate confirmation message is generated according to the information of the enterprise and the information of the dealer.
[0078] The points module is used to obtain points through activities such as evaluation and sharing, and to exchange goods or coupons using the points;
[0079] The enterprise evaluation module receives the evaluation information of the enterprise sent by the user terminal;
[0080] The social module receives the green product information shared by the user terminal.
[0081] By applying the big data enterprise environment public service system provided by the embodiment of the application, a comprehensive, convenient and reliable enterprise environment public service platform is provided, and through comprehensive and accurate enterprise environment information management, the sustainable development of enterprises can be promoted, the ecological environment can be protected, and the overall environmental protection level of the society can be improved.
[0082] Those skilled in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0083] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and
[0084] The above detailed description describes the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A big data enterprise environment common service system, characterized in that, The system comprises: a data acquisition module for processing unstructured data to obtain enterprise environmental information through web scraping algorithms and multi-source API calling mechanisms; a data processing module for dynamically adjusting the task priority and collection frequency of data acquisition by monitoring the changes in the data source and user query records in real time; a user interface module for receiving request information from multiple users and returning corresponding results according to the enterprise environmental information; the request information includes requests through a web browser, a mobile application, a desktop client, and an API call; a website function module for receiving request information sent by the user interface module and performing certificate queries, enterprise evaluations, and point exchanges according to the web browser, mobile application, desktop client, API call request, and enterprise environmental information; wherein the data processing module is further configured to dynamically adjust the cache strategy according to the request frequency of user requests and the update frequency of data; the cache strategy includes keeping high-frequency requested data in cache and removing low-frequency requested data from cache; the data processing module is specifically configured to: set up a message queue to monitor changes in the data source, receive notification messages sent by the message queue when the data source is updated, or configure Webhooks of the data source provider to receive HTTP request notifications sent by the data source provider when the data changes; or set up a timing task to periodically check the status of the data source and trigger data collection tasks when there are changes; record each query request of the user, including query time, query content, and query results; record the user's query history in the database, including query keywords and query frequency; and push the user's query records to the message queue; dynamically adjust the collection frequency according to the change frequency of the data source; data sources with high change frequency have high priority and high collection frequency; dynamically adjust the priority of related data sources according to the keywords and frequency in the user query records; data sources with high user query frequency have high priority and high collection frequency; train a machine learning model to predict the change trend of the data source and the user query trend, and adjust the task priority and collection frequency according to the prediction results; preset a rule to dynamically adjust the task priority and collection frequency according to the rule.
2. The system of claim 1, wherein, The processing of unstructured data to obtain enterprise environmental information specifically includes: obtaining enterprise information from enterprise websites, public databases, or social media through web crawler tools; obtaining environmental monitoring data and enterprise annual reports by calling the API interface provided by the enterprise; obtaining file information uploaded by the user; preprocessing the obtained information; the preprocessing includes text cleaning, image processing, and audio-to-text processing; using NLP tools to identify key entities in the text and judging the positive or negative nature of the enterprise environmental information according to the sentiment tendency of the text information; extracting keywords and phrases from the text; using OCR technology to extract text information from images; using deep learning models to identify specific objects in images, including pollution sources and environmental protection facilities; Extracting table data from PDF and image using table parsing tool; Storing the processed data in a relational database and creating an index for the data to generate a big data index library.
3. The system of claim 1, wherein, The user interface module is further configured to define different priority tags for the first type of data and the second type of data, and perform sorting processing before sending; the priority of the first type of data is higher than that of the second type of data; wherein the first type of data is data with strong real-time performance; the second type of data is data with low delay sensitivity.
4. The system of claim 1, wherein, The user interface module is further configured to adjust the sending rate according to the network feedback flow and according to an adaptive flow control algorithm; wherein the adaptive flow control algorithm includes the slow start mechanism and the congestion avoidance mechanism of TCP.
5. The system of claim 1, wherein, The website function module includes production enterprise management and distributor management sub-modules, configured to receive information of production enterprises and information of distributors; and when a certificate query is received, generate a certificate confirmation message according to the information of the enterprises and the information of the distributors; The credit module is configured to obtain credits through evaluation and sharing activities, and use the credits to exchange goods or coupons; The enterprise evaluation module; Receive the evaluation information of the enterprise sent by the user terminal; The social module receives the green product information shared by the user terminal.
Citation Information
Patent Citations
Data interaction system for archive management
CN117827743A
Tin new material product database scheduling method based on B / S architecture
CN118656391A
Data acquisition method and device, equipment and storage medium
CN118820015A