A data product production management and publishing system
Through the data product production management and release system, the crawled data is processed and displayed to generate intuitive and clear target data products, solving the problem of messy data and improving the efficiency and accuracy of data acquisition.
Patent Information
- Application Number
- CN202210509681.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-05-11
AI Technical Summary
The large amount of data obtained in the prior art has not been processed, resulting in the data presentation being not intuitive and chaotic enough, making it difficult for users to effectively obtain the desired information.
Provide a data product production management and publishing system, including data crawling, processing, entry and display subsystems, through which crawled data is processed and displayed to generate target data products.
It realizes the complete process from data crawling to display, generates intuitive and clear target data products, and improves the efficiency and accuracy of data acquisition.
Smart Images

Figure CN114896233B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a data product production management and publishing system. Background Art
[0002] With the advent of the big data era, how to use big data to "predict" the future and analyze the laws and truths behind the data has received widespread attention.
[0003] Solutions in the existing technology only focus on how to quickly obtain a large amount of data, but do not involve processing the large amount of data obtained, resulting in technical problems such as the data presented to users is not intuitive enough and the data is disorganized.
[0004] Therefore, there is an urgent need to provide a data product production management and publishing system for providing data products to users so that users can intuitively and clearly obtain the technical effects of the desired data information. Summary of the Invention
[0005] In view of this, it is necessary to provide a data product production management and publishing system to solve the technical problems existing in the existing technology that the large amount of acquired data is not processed, resulting in the data presented to users being not intuitive enough and the data being disorganized.
[0006] In order to solve the above technical problems, the present invention provides a data product production management and publishing system, including: a data capture subsystem, a data processing subsystem, a data entry subsystem and a data display subsystem;
[0007] The data crawling subsystem is used to crawl data from the website, obtain crawled data, and store the crawled data in the target database;
[0008] The data processing subsystem is used to obtain the crawled data from the target database and process the crawled data to obtain data to be entered;
[0009] The data entry subsystem is used to obtain attribute information of the data to be entered, and generate a target data product based on the attribute information and the data to be entered;
[0010] The data display subsystem is used to display the target data product.
[0011] In some possible implementations, the target data product includes visual data, tabular data, data reports, data packets, and data illustrations. The formats of the visual data and the tabular data are spreadsheet formats, the format of the data reports is portable file format, the format of the data packets is compressed file format, and the format of the data illustrations is portable network graphics format.
[0012] In some possible implementations, the data crawling subsystem includes a crawler management module, an operation management module, a monitoring management module, and a database management module;
[0013] The crawler management module is used to configure the crawler file and crawl the website based on the crawler file to obtain the crawled data;
[0014] The operation management module is used to configure a timing file and / or a termination file, and to implement a scheduled data crawling of the website based on the timing file, and to stop the data crawling of the website based on the termination file;
[0015] The monitoring management module is used to monitor the data crawling process of the crawler file and generate a monitoring file;
[0016] The database management module is used to configure storage rules and store the crawled data in a target database based on the storage rules.
[0017] In some possible implementations, the data processing subsystem includes a data source management module, a task scheduling management module, and a data cleaning module;
[0018] The data source management module is used to determine the database to be retrieved from the target database and obtain the crawled data from the database to be retrieved.
[0019] The task scheduling management module is used to determine the data processing task, and the data processing task includes the data processing time;
[0020] The data cleaning module is used to clean the crawled data based on the data processing time to obtain the data to be entered.
[0021] In some possible implementations, the data entry subsystem includes a data entry module and a data review module;
[0022] The data entry module is used to obtain the attribute information of the data to be entered, and generate a data product to be reviewed by combining the attribute information and the data to be entered;
[0023] The data review module is used to review the data product to be reviewed. If the data product to be reviewed passes the review, the data product to be reviewed becomes the target data product.
[0024] In some possible implementations, the data entry subsystem further includes a chart matching module;
[0025] The chart matching module is used to match the target chart form for the visual data;
[0026] The data display subsystem is used to display the visual data in the form of the target chart.
[0027] In some possible implementations, the data entry subsystem further includes a search module; the search module is configured to obtain search keywords, obtain the data to be selected from the data to be entered based on the search keywords, and generate the data packet based on the data to be selected.
[0028] In some possible implementations, the data to be entered is in a first language, and the data product production management and publishing system further includes a data entry intelligent collaboration subsystem, wherein the data entry intelligent collaboration subsystem includes a translation function module, a duplicate detection function module, an error correction function module, a classification function module, and a keyword generation function module;
[0029] The translation function module is used to convert the data to be entered into a second language;
[0030] The duplicate detection module is used to determine the duplication of the data to be entered and delete the duplicate data to be entered;
[0031] The error correction function module is used to determine whether the data to be entered is erroneous, and when the data to be entered is erroneous, generate an error reminder message;
[0032] The classification function module is used to classify the data to be entered;
[0033] The keyword generation function module is used to determine at least one keyword of the data to be entered.
[0034] In some possible implementations, the data to be entered includes a plurality of portable file data in a portable file format, the portable file data includes a first portable file sub-data and a second portable file sub-data, the duplicate detection function module includes a first duplicate detection unit for performing duplicate detection on the portable file data, the first duplicate detection unit includes a conversion sub-unit, a hash processing sub-unit, and a first duplicate detection sub-unit;
[0035] The conversion sub-unit is used to extract first sub-data to be judged and second sub-data to be judged from the first portable file sub-data and the second portable file sub-data respectively;
[0036] The hash processing subunit is used to perform hash processing on the first sub-data to be judged and the second sub-data to be judged respectively to obtain a first hash string and a second hash string;
[0037] The first duplicate detection sub-unit is used to perform duplication judgment on the first portable file sub-data and the second portable file sub-data based on the first hash string and the second hash string, and when the first portable file sub-data and the second portable file sub-data are duplicated, delete the first portable file sub-data or the second portable file sub-data.
[0038] In some possible implementations, the data to be entered further includes a plurality of spreadsheet file data in a spreadsheet format, the spreadsheet file data includes a first spreadsheet file sub-data and a second spreadsheet file sub-data, the duplicate detection function module includes a second duplicate detection unit for performing duplicate detection on the spreadsheet file data, the second duplicate detection unit includes a data processing sub-unit and a second duplicate detection sub-unit;
[0039] The data processing sub-unit is used to extract first data information and second data information of the first electronic spreadsheet file sub-data and the second electronic spreadsheet file sub-data respectively;
[0040] The second duplicate detection sub-unit is used to perform duplication detection on the first spreadsheet file sub-data and the second spreadsheet file sub-data based on the first data information and the second data information, and when the first spreadsheet file sub-data and the second spreadsheet file sub-data are duplicated, delete the first spreadsheet file sub-data or the second spreadsheet file sub-data.
[0041] The beneficial effect of the present invention is that the data product production management and publishing system provided by the present invention realizes the complete process from crawling data on the website to generating target data products based on the crawled data by setting up a data capture subsystem, a data processing subsystem, a data entry subsystem and a data display subsystem. Compared with directly displaying crawled data to users, by generating and displaying target product data, users can obtain the desired data information intuitively and clearly.
[0042] Furthermore, the present invention realizes the seamless connection of data capture, data processing, data entry and data display, making the entire production process of the target data product clearer, more efficient and standardized, thereby improving the generation efficiency of the target data product. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0044] Figure 1 A schematic diagram of the structure of an embodiment of the data product production management and publishing system provided by the present invention;
[0045] Figure 2 A schematic structural diagram of an embodiment of the data entry subsystem provided by the present invention;
[0046] Figure 3 This is a schematic structural diagram of an embodiment of the duplicate detection function module provided by the present invention. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0048] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowcharts may be implemented out of sequence, and steps that do not have a logical contextual relationship may be reversed or performed simultaneously. In addition, those skilled in the art, guided by the present disclosure, may add one or more additional operations to the flowcharts or remove one or more operations from the flowcharts.
[0049] In the description of the embodiments of the present invention, unless otherwise specified, "plurality" means two or more. "And / or" describes the association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.
[0050] Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor systems and / or microcontroller systems.
[0051] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0052] The present invention provides a data product production management and publishing system, which is described below.
[0053] Figure 1 This is a schematic diagram of an embodiment of the data product production management and publishing system provided by the present invention, as shown in FIG. Figure 1 As shown, the data product production management and publishing system 10 provided by the embodiment of the present invention includes: a data capture subsystem 100, a data processing subsystem 200, a data entry subsystem 300 and a data display subsystem 400;
[0054] The data crawling subsystem 100 is used to crawl data from the website, obtain crawled data, and store the crawled data in the target database;
[0055] The data processing subsystem 200 is used to obtain crawled data from the target database and process the crawled data to obtain data to be entered;
[0056] The data entry subsystem 300 is used to obtain attribute information of the data to be entered, and generate a target data product based on the attribute information and the data to be entered;
[0057] The data display subsystem 400 is used to display the target data product.
[0058] Compared with the prior art, the data product production management and publishing system 10 provided in the embodiment of the present invention, by setting up a data capture subsystem 100, a data processing subsystem 200, a data entry subsystem 300 and a data display subsystem 400, can realize the complete process from crawling data on the website to generating a target data product based on the crawled data. Compared with directly displaying the crawled data to the user, by generating and displaying the target product data, the user can obtain the desired data information intuitively and clearly.
[0059] Furthermore, the embodiments of the present invention achieve seamless connection between data capture, data processing, data entry and data display, making the entire production process of the target data product clearer, more efficient and standardized, thereby improving the generation efficiency of the target data product.
[0060] It should be understood that in a specific embodiment of the present invention, the data presentation subsystem 400 may be a Dysprosium data aggregation platform.
[0061] It should be noted that the target database in the embodiment of the present invention may be at least one of a MySQL database, a MongoDB database, and a Redis database.
[0062] It should also be noted that the data processing subsystem 200 in the embodiment of the present invention can also obtain data from an external database or an external file, and can specifically obtain data from an external database through an API data structure.
[0063] In some embodiments of the present invention, the target data product includes visual data, tabular data, data reports, data packets, and data illustrations. The format of the visual data and tabular data is a spreadsheet format (xlsx), the format of the data report is a portable document format (PDF), the format of the data packet is a compressed file format (zip), and the format of the data illustration is a portable network graphics format (png).
[0064] It should be understood that the data package may include visual data, tabular data, data reports, and graphic information.
[0065] In some embodiments of the present invention, Figure 1 As shown, the data crawling subsystem 100 includes a crawler management module 110, an operation management module 120, a monitoring management module 130 and a database management module 140;
[0066] The crawler management module 110 is used to configure crawler files and crawl data from the website based on the crawler files to obtain crawled data;
[0067] The operation management module 120 is used to configure a timing file and / or a termination file, and to implement a scheduled data crawling of the website based on the timing file, and to stop the data crawling of the website based on the termination file;
[0068] The monitoring management module 130 is used to monitor the data crawling process of the crawler file and generate a monitoring file;
[0069] The database management module 140 is used to configure storage rules and store the crawled data into the target database based on the storage rules.
[0070] The embodiment of the present invention can adjust the method and time of obtaining crawled data from the website by setting the crawler management module 110, the operation management module 120, and the monitoring management module 130, thereby improving the flexibility of obtaining crawled data.
[0071] Specifically, the crawler file may include a crawler parameter setting subfile, a crawler name maintenance subfile, and a crawler list display subfile, etc. Parameters such as the crawling speed of data can be adjusted through the crawler parameter setting subfile.
[0072] Furthermore, when an error occurs in crawling data by the data capture subsystem, the specific time of the error can be defined through the monitoring file generated by the monitoring management module 130, which facilitates maintenance by maintenance personnel.
[0073] In some embodiments of the present invention, Figure 1 As shown, the data processing subsystem 200 includes a data source management module 210, a task scheduling management module 220, and a data cleaning module 230;
[0074] The data source management module 210 is used to determine the database to be retrieved from the target database and obtain crawled data from the database to be retrieved;
[0075] The task scheduling management module 220 is used to determine data processing tasks, which include data processing time;
[0076] The data cleaning module 230 is used to clean the crawled data based on the data processing time to obtain the data to be entered.
[0077] Specifically, the data source management module 210 is used to determine the database to be retrieved from the MySQL database, the MongoDB database, and the Redis database according to the database type, the maximum number of connections, and the connection timeout limit.
[0078] Among them, data cleaning of crawled data includes but is not limited to operations such as extraction and conversion of data.
[0079] It should be noted that: Figure 1 As shown, the data processing subsystem 200 may further include a task management module 240, which is used to manage data processing tasks, for example, canceling data processing tasks, or adjusting data processing time in data processing tasks.
[0080] In some embodiments of the present invention, Figure 1 As shown, the data entry subsystem 300 includes a data entry module 310 and a data review module 320;
[0081] The data entry module 310 is used to obtain attribute information of the data to be entered, and to generate a data product to be reviewed by combining the attribute information and the data to be entered;
[0082] The data review module 320 is used to review the data product to be reviewed. If the data product to be reviewed passes the review, the data product to be reviewed becomes the target data product.
[0083] The attribute information includes but is not limited to data title, data format, etc.
[0084] Specifically, the data review module 320 is used to review whether the input data has been tampered with during the input process and whether the data product to be reviewed complies with the specifications.
[0085] The embodiment of the present invention can improve the accuracy and reliability of the target data product by providing the data review module 320.
[0086] It should be understood that the data product production management and publishing system 10 can also include a manual entry method, that is, in addition to generating the target data product through the data entry subsystem 300, it can also receive entry from manual operations to obtain the target data product.
[0087] In some embodiments of the present invention, Figure 2 As shown, the data entry module 310 includes a visual data entry unit 311, a table data entry unit 312, a data report entry unit 313 and a data packet entry unit 314. The visual data entry unit 311, the table data entry unit 312, the data report entry unit 313 and the data packet entry unit 314 are respectively used to generate visual data to be reviewed, table data to be reviewed, data reports to be reviewed and data packets to be reviewed.
[0088] Likewise, if Figure 2 As shown, the data review module 320 includes a visual data review unit 321, a tabular data review unit 322, a data report review unit 323 and a data packet review unit 324. The visual data review unit 321, the tabular data review unit 322, the data report review unit 323 and the data packet review unit 324 are respectively used to review the visual data to be reviewed, the tabular data to be reviewed, the data report to be reviewed and the data packet to be reviewed. If the review is passed, the visual data, tabular data, data report and data packet are obtained.
[0089] In a specific application scenario, the data entry user enters the data to be entered to generate a data product to be reviewed, and the reviewer reviews the data product to be reviewed. When the reviewer believes that the data product to be reviewed has not passed the review, in order to improve the modification speed of the data product to be reviewed, in some embodiments of the present invention, such as Figure 1 As shown, the data entry subsystem 300 further includes an information communication management module 330, which is used to convey the reasons for the reviewer's failure to the data entry user for the reference of the data entry user, thereby improving the modification speed of the data product to be reviewed.
[0090] Furthermore, in order to increase the production speed of the target data product, multiple data entry and multiple reviewers may be allowed to operate at the same time. In order to avoid the reviewer doing the work of the data entry person or the data entry person doing the reviewer's work when operating at the same time, resulting in the final generated target data product not meeting the requirements, in some embodiments of the present invention, different permissions are assigned to the reviewer and the data entry person. Through permission restrictions, multiple lines of data entry and data review can be carried out simultaneously without any confusion, thereby improving the efficiency of target data product generation and ensuring the reliability of the generated target data product.
[0091] Since there are many ways to display charts, such as bar charts, line charts, pie charts, etc., in order to display visual data in the best chart display method, in some embodiments of the present invention, Figure 1 As shown, the data entry subsystem 300 also includes a chart matching module 340;
[0092] The chart matching module 340 is used to match the target chart form to the visual data;
[0093] The data display subsystem 400 is used to display visual data in the form of target charts.
[0094] Specifically, the chart matching module 340 is specifically used to obtain a mapping relationship between preset data features and chart forms, obtain data features of visual data, and match the target chart form of the visual data based on the data features and the mapping relationship.
[0095] For example, if the number of data items of a visual data does not exceed 5 and the data is in percentage form, a pie chart is selected to display the visual data.
[0096] The present invention can improve the readability and intuitiveness of visual data by matching the target chart form to the visual data in real time.
[0097] In order to facilitate the generation of data packets, in some embodiments of the present invention, such as Figure 1 As shown, the data entry subsystem 300 further includes a search module 350 ; the search module 350 is used to obtain search keywords, obtain the data to be selected from the data to be entered according to the search keywords, and generate a data package according to the data to be selected.
[0098] It should be understood that the search keywords may include at least one of data source, data subject, data year, and data creation time, etc., and the weight of each keyword can be set.
[0099] In some embodiments of the present invention, the data to be entered is in the first language, such as Figure 1 As shown, the data product production management and publishing system 10 further includes a data entry intelligent collaboration subsystem 500, which includes a translation function module 510, a duplicate detection function module 520, an error correction function module 530, a classification function module 540, and a keyword generation function module 550;
[0100] The translation function module 510 is used to convert the data to be entered into a second language;
[0101] The duplicate detection module 520 is used to determine the duplication of the data to be entered and delete the duplicate data to be entered;
[0102] The error correction module 530 is used to determine whether the data to be entered is erroneous and generate an error reminder message when the data to be entered is erroneous;
[0103] The classification function module 540 is used to classify the input data;
[0104] The keyword generation module 550 is used to determine at least one keyword for the data to be entered.
[0105] By setting up a data entry intelligent collaborative subsystem 500, the embodiment of the present invention can improve the generation efficiency and accuracy of the target data product through the translation function module 510, further improve the accuracy and reliability of the target data product through the duplicate detection function module 520 and the error correction function module 530, and facilitate the classification of the target data product through the classification function module 540 and the keyword generation function module 550 for users to view by category, thereby improving the user's browsing comfort and convenience.
[0106] It should be understood that the first language and the second language are different, and the first language and the second language may be any one of Chinese, English, German, Japanese, etc.
[0107] In a specific embodiment of the present invention, the error correction module 530 is specifically used to identify errors in the data to be entered. If there are typos in the data to be entered, the data to be entered is highlighted to prompt the user.
[0108] In some embodiments of the present invention, the data to be entered includes a plurality of portable document data in a portable document format (PDF), the portable document data including a first portable document sub-data and a second portable document sub-data, such as Figure 3 As shown, the duplicate detection function module 520 includes a first duplicate detection unit 521 for detecting duplicates of portable document data. The first duplicate detection unit 521 includes a conversion subunit 5211, a hash processing subunit 5212, and a first duplicate detection subunit 5213.
[0109] The conversion sub-unit 5211 is used to extract the first sub-data to be judged and the second sub-data to be judged from the first portable file sub-data and the second portable file sub-data respectively;
[0110] The hash processing sub-unit 5212 is configured to perform hash processing on the first sub-data to be determined and the second sub-data to be determined, respectively, to obtain a first hash string and a second hash string;
[0111] The first duplicate detection sub-unit 5213 is used to perform duplicate detection on the first portable file sub-data and the second portable file sub-data according to the first hash string and the second hash string, and when the first portable file sub-data and the second portable file sub-data are duplicated, delete the first portable file sub-data or the second portable file sub-data.
[0112] It should be noted that the formats of the first sub-data to be determined and the second sub-data to be determined are picture formats.
[0113] In order to improve the reliability of duplicate detection, in some embodiments of the present invention, the conversion sub-unit 5211 is used to extract three first sub-data to be judged and three second sub-data to be judged from the first portable file sub-data and the second portable file sub-data, respectively. Duplicate detection is performed using the three first sub-data to be judged and the three second sub-data to be judged, which can reduce the possibility of misjudgment and improve the reliability of duplicate detection.
[0114] In some embodiments of the present invention, the data to be entered also includes a plurality of electronic table file data in a spreadsheet (xlsx) format, and the electronic table file data includes a first electronic table file sub-data and a second electronic table file sub-data, such as Figure 3 As shown, the duplicate detection function module 520 includes a second duplicate detection unit 522 for detecting duplicate data in the electronic spreadsheet file. The second duplicate detection unit 522 includes a data processing subunit 5221 and a second duplicate detection subunit 5222.
[0115] The data processing sub-unit 5221 is used to extract first data information and second data information of the first electronic spreadsheet file sub-data and the second electronic spreadsheet file sub-data respectively;
[0116] The second duplicate detection sub-unit 5222 is used to perform duplicate detection on the first spreadsheet file sub-data and the second spreadsheet file sub-data based on the first data information and the second data information, and delete the first spreadsheet file sub-data or the second spreadsheet file sub-data when the first spreadsheet file sub-data and the second spreadsheet file sub-data are duplicated.
[0117] The first data information and the second data information may include a data title, a data source, and data content.
[0118] The above is a detailed introduction to the data product production management and publishing system provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, based on the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A data product production management and publishing system, characterized by: include: Data capture subsystem, data processing subsystem, data entry subsystem and data display subsystem; The data crawling subsystem is used to crawl data from the website, obtain crawled data, and store the crawled data in the target database; The data processing subsystem is used to obtain the crawled data from the target database and process the crawled data to obtain data to be entered; The data entry subsystem is used to obtain attribute information of the data to be entered, and generate a target data product based on the attribute information and the data to be entered; the attribute information includes a data title and a data format; The data display subsystem is used to display the target data product; The data entry intelligent collaboration subsystem includes a duplicate detection function module, which includes a first duplicate detection unit for detecting duplicates of portable file data and a second duplicate detection unit for detecting duplicates of electronic spreadsheet file data; the first duplicate detection unit and the second duplicate detection unit have different duplicate detection logics.
2. The data product production management and publishing system according to claim 1, characterized in that: The target data product includes visual data, tabular data, data reports, data packets and data illustrations. The formats of the visual data and the tabular data are spreadsheet formats, the format of the data reports is portable file format, the format of the data packets is compressed file format, and the format of the data illustrations is portable network graphics format.
3. The data product production management and publishing system according to claim 1, characterized in that: The data crawling subsystem includes a crawler management module, an operation management module, a monitoring management module and a database management module; The crawler management module is used to configure the crawler file and crawl the website based on the crawler file to obtain the crawled data; The operation management module is used to configure a timing file and / or a termination file, and to implement a scheduled data crawling of the website based on the timing file, and to stop the data crawling of the website based on the termination file; The monitoring management module is used to monitor the data crawling process of the crawler file and generate a monitoring file; The database management module is used to configure storage rules and store the crawled data in a target database based on the storage rules.
4. The data product production management and publishing system according to claim 1, characterized in that: The data processing subsystem includes a data source management module, a task scheduling management module and a data cleaning module; The data source management module is used to determine a database to be retrieved from the target database, and obtain the crawled data from the database to be retrieved; The task scheduling management module is used to determine the data processing task, and the data processing task includes the data processing time; The data cleaning module is used to clean the crawled data based on the data processing time to obtain the data to be entered.
5. The data product production management and publishing system according to claim 2, characterized in that: The data entry subsystem includes a data entry module and a data review module; The data entry module is used to obtain the attribute information of the data to be entered, and generate a data product to be reviewed by combining the attribute information and the data to be entered; The data review module is used to review the data product to be reviewed. If the data product to be reviewed passes the review, the data product to be reviewed becomes the target data product.
6. The data product production management and publishing system according to claim 5, characterized in that: The data entry subsystem also includes a chart matching module; The chart matching module is used to match the target chart form for the visual data; The data display subsystem is used to display the visual data in the form of the target chart.
7. The data product production management and publishing system according to claim 5, characterized in that: The data entry subsystem further includes a search module; the search module is used to obtain search keywords, obtain the data to be selected from the data to be entered according to the search keywords, and generate the data packet according to the data to be selected.
8. The data product production management and publishing system according to claim 2, characterized in that: The data to be entered is in a first language, and the data product production management and publishing system further includes a data entry intelligent collaboration subsystem, which includes a translation function module, an error correction function module, a classification function module, and a keyword generation function module; The translation function module is used to convert the data to be entered into a second language; The duplicate detection module is used to determine the duplication of the data to be entered and delete the duplicate data to be entered; The error correction function module is used to determine whether the data to be entered is erroneous, and when the data to be entered is erroneous, generate an error reminder message; The classification function module is used to classify the data to be entered; The keyword generation function module is used to determine at least one keyword of the data to be entered.
9. The data product production management and publishing system according to claim 8, characterized in that: The data to be entered includes a plurality of portable file data in a portable file format, the portable file data includes a first portable file sub-data and a second portable file sub-data, and the first duplicate detection unit includes a conversion sub-unit, a hash processing sub-unit and a first duplicate detection sub-unit; The conversion sub-unit is used to extract first sub-data to be judged and second sub-data to be judged from the first portable file sub-data and the second portable file sub-data respectively; The hash processing subunit is used to perform hash processing on the first sub-data to be judged and the second sub-data to be judged respectively to obtain a first hash string and a second hash string; The first duplicate detection sub-unit is used to perform duplicate judgment on the first portable file sub-data and the second portable file sub-data based on the first hash string and the second hash string, and when the first portable file sub-data and the second portable file sub-data are duplicated, delete the first portable file sub-data or the second portable file sub-data.
10. The data product production management and publishing system according to claim 8, characterized in that: The data to be entered further includes a plurality of electronic table file data in an electronic table format, wherein the electronic table file data includes a first electronic table file sub-data and a second electronic table file sub-data, and the second duplicate detection unit includes a data processing sub-unit and a second duplicate detection sub-unit; The data processing sub-unit is used to extract first data information and second data information of the first electronic spreadsheet file sub-data and the second electronic spreadsheet file sub-data respectively; The second duplicate detection sub-unit is used to perform duplication detection on the first spreadsheet file sub-data and the second spreadsheet file sub-data based on the first data information and the second data information, and when the first spreadsheet file sub-data and the second spreadsheet file sub-data are duplicated, delete the first spreadsheet file sub-data or the second spreadsheet file sub-data.
Citation Information
Patent Citations
Web service publishing and visualization combined system for unstructured data
CN110134776A
Enterprise analysis system based on big data
CN114443924A