News publishing statistical analysis system
By designing a B-S architecture news publishing statistical analysis system, using React and FastAPI to achieve front-end integration, and combining PostgreSQL and Pandas for data processing, the existing system's inefficiency and single functions are solved, efficient data analysis and intelligent report generation are realized, and real-time decision-making needs are supported.
Patent Information
- Application Number
- CN202510290699.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120218035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of press and publication statistical analysis, and specifically provides a press and publication statistical analysis system. Background Art
[0002] Press and publication statistics is an important part of the government's statistical work, and the accuracy of its data and the analysis efficiency directly affect the formulation of cultural industry policies and industry supervision. In the existing technology, traditional statistical systems mainly rely on a basic database to store original publication data (such as the print run of periodicals, book ISBN numbers, financial and tax fields), provide data retrieval functions through a simple query interface, and rely on manual export of Excel tables for secondary processing. For example, periodical data is stored annually but not classified and aggregated, book publication records only retain details without cross-year association of ISBN numbers, and the financial module supports the entry of basic fields but does not integrate automated calculation logic. Although such systems can achieve basic data storage, their functional design is limited to static data management and cannot meet the needs of dynamic analysis and real-time decision-making.
[0003] However, the existing technology has significant defects: firstly, low efficiency. The outdated system architecture leads to slow data processing. For example, when searching for periodical details, it takes too long due to lag and is difficult to meet the real-time requirements of on-site defense during joint review; secondly, single function. It only provides the display of original data and lacks key calculation capabilities (such as aggregating quarterly publication data by ISBN and automatically identifying the status of new publications / reprints), and complex formula calculations need to be performed manually, which is time-consuming and error-prone; thirdly, insufficient intelligence. It cannot generate text descriptions of data changes, and the joint review report relies on manual writing, resulting in the disconnection between analysis conclusions and data. For example, due to the lack of ISBN number aggregation function in the book module, it is difficult to analyze the reasons for cross-year changes of the same book; the financial module does not implement automatic verification of tax items (such as the comparison of the sum of income tax and business tax with the total amount), and manual line-by-line verification is required, greatly increasing the work complexity. These problems seriously restrict the efficiency and decision-making support ability of press and publication statistics work, and for this reason, a press and publication statistical analysis system is proposed. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a press and publication statistical analysis system to solve the problems in the background art.
[0005] To achieve the above object, the present invention provides the following technical solution: A press and publication statistical analysis system, which is implemented based on the B-S architecture and includes:
[0006] A front-end interaction module: constructs a user interface based on the React framework, supports users to upload original Excel files, select business modules and index values, and displays analysis results through web tables;
[0007] Backend processing module: Built with FastAPI, including:
[0008] Data cleaning engine:
[0009] Perform column name mapping, special character removal, missing value filling, and error data correction on the uploaded Excel file to generate a standardized structured data table
[0010] Classification calculation engine: Based on preset algorithms, implement data classification aggregation and index calculation for periodicals, newspapers, books, audio-visual electronics, and financial modules;
[0011] Data storage module: Based on the PostgreSQL database, store the original Excel file in the form of BYTEA (Blob), and serialize the cleaned data and store it in the JSONB (JSON) format through Pandas to support efficient querying. When requesting data through the API, deserialize it from JSONB to a Pandas DataFrame for calculation and processing.;
[0012] Report generation module: Dynamically generate explanatory text containing differences and year-on-year changes based on the metric values selected by the user, and export a downloadable Blob file.
[0013] Preferably, the front-end interaction module further includes:
[0014] Real-time verification of the format and integrity of required fields of the uploaded file, and return an error prompt when the verification fails;
[0015] Provide dynamic filtering and sorting controls to support users in selecting analysis requirements by business module (periodicals, newspapers, books, audio-visual electronics, finance) and metric values (average number of copies printed per issue, total number of copies printed, total pricing amount, etc.);
[0016] Integrate "View Details", "Copy Description", and "Download Comparison File" interactive buttons to support real-time filtering and sorting operations at the review site.
[0017] Preferably, the data cleaning engine further includes:
[0018] Perform regular expression matching verification on the fields of the original Excel file to ensure that the ISBN number conforms to the international standard format and the date field conforms to the YYYY-MM-DD format;
[0019] Automatically identify and delete duplicate data entries, missing data entries, and unsubmitted status data entries, and fill the missing fields with the mean value of similar data or mark them for manual review.
[0020] Adjust the abnormal data that does not conform to the specifications appropriately. For example: If the same ISBN code corresponds to multiple book titles, merge the book titles and add identifiers for subsequent manual review.
[0021] Preferably, the classification calculation engine further includes:
[0022] Journal analysis: Automatically classify and aggregate data according to the first letter category of CIP (comprehensive category Z, science and technology category NOPQRSTUVX, social philosophy category ABCDEFK, literature and art category IJ), calculate and analyze the differences and year-on-year change percentages of key business indicators such as the average quarterly print run, total print run, total print sheets, and total pricing amount in the past two years, generate corresponding reports and tables sorted by indicators according to user needs;
[0023] Newspaper analysis: Screen data according to newspaper categories (professional, comprehensive, life service, reader target, digest), calculate and analyze the differences and year-on-year change percentages of key business indicators such as the average quarterly print run, total print run, total print sheets, and total pricing amount in the past two years, generate corresponding reports and tables sorted by indicators according to user needs;
[0024] Book analysis: Aggregate the quarterly publication details of each book according to the publisher name and publication mode (new publication, reprint, leased type) selected by the user, including: the print run, print sheets, and total amount in each quarter, automatically identify the status of new publication to reprint / reprint to new publication of books through algorithms, generate a detailed Excel table and an overview description text to assist in material writing;
[0025] Audio-visual and electronic analysis: Classify and aggregate data according to the carrier form (audio-visual and electronic such as CD, DVD, etc.), supplement and summarize the publication quantity, distribution actual amount, list price amount and other business indicators in each quarter of this year and the previous year into the same table; add corresponding auxiliary analysis columns to view whether the data of publication and distribution (a total of 4 data sources) in this year and the previous year are consistent; add auxiliary columns to analyze whether there are differences in the publication and distribution data in the past two years (such as whether the languages and distribution discounts are unified); generate ranking tables and text descriptions of the core business indicators of each unit and the overview data.
[0026] Financial analysis: Analyze the data in the enterprise, institution, and non-independent legal entity tables, and automatically calculate the difference between the total sum of per capita salary and tax items and the total tax amount.
[0027] Preferably, the journal analysis function further includes:
[0028] After the user clicks the "View Details" button, display all the publication details and historical change trend charts of the specified journal;
[0029] Through the "Copy Description" function, generate standardized description text containing key business indicators with one click.
[0030] Preferably, the book analysis function further includes:
[0031] Generating detailed business indicators such as the quarterly print runs, print sheets, and total amounts of books aggregated by ISBN number (Excel details table), and generating a two-year comparative analysis report of indicators such as the number of general books (books other than teaching aids, children's books, and textbooks) under the 22 classifications of the Chinese Library Classification, the total number of print sheets, the total pricing amount, and the quarterly publication details;
[0032] Exporting the publication index ranking file grouped by publisher through the "Download Publisher Group Ranking File" button.
[0033] Exporting the two-year publication details comparison file of this publisher through the "Download Two-Year Comparison File of This Publisher" button.
[0034] Automatically generating an overview text description to assist in writing corresponding materials. Preferably, the audio-visual and electronic analysis function further includes:
[0035] Dynamically generating a sorted or reverse-sorted ranking table according to the sorting criteria (number of varieties, number of publications, difference in total pricing amount) selected by the user;
[0036] Generating a two-year comparison file of publication and distribution containing unit codes and ISBN numbers.
[0037] Generating explanatory text materials for the core business indicators of the corresponding units to assist in writing.
[0038] Preferably, the financial analysis function further includes:
[0039] Generating a complete explanatory document containing the difference in total tax amount and the difference in financial expenses and interest expenses through the "One-Click Download Explanation File" function;
[0040] Classifying and counting the changes in the number of units within two years by enterprise units, public institutions, and non-independent legal person units, and calculating the year-on-year increase or decrease percentage of per capita salary.
[0041] Preferably, the report generation module further includes:
[0042] Based on natural language generation technology, dynamically splicing predefined templates and real-time data to generate analysis conclusions that meet the requirements of joint review;
[0043] Supporting users to customize report templates, including table styles, text paragraph structures, and key indicator preference settings.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] The present invention processes the news and publication statistical review work in an informatized, intelligent, and automated manner, significantly improving the review efficiency and the efficiency of material writing. It realizes full-process data management through a modular architecture and intelligent algorithms, and ensures the traceability of the original data and the efficient invocation of the cleaned data based on the Blob storage of PostgreSQL and the JSON serialization and deserialization technologies of Pandas. The front-end React dynamically loads enumeration values and supports multi-dimensional real-time filtering and sorting. The back-end FastAPI classification engine automatically completes the aggregation of the first letters of journal CIP, the identification of ISBN status, the analysis of the difference in tax items, and the statistical classification of the Chinese Library Classification. Combined with natural language generation technology, it can generate a standardized report template with one key and export Excel / PDF comparison files. The system improves stability through seamless migration of Alembic and multi-process deployment of Gunicorn, solves the pain points of the traditional system with single function, slow response, and high manual dependence, and provides real-time data support and accurate decision-making basis for the review.
[0046] Other features and advantages of the present invention will be described in the following specification, and in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. Brief Description of the Drawings
[0047] Figure 1 is the overall flowchart of the news and publication statistical analysis system of the present invention;
[0048] Figure 2 is the audio-visual and electronic business flowchart of the present invention;
[0049] Figure 3 is the book business flowchart of the present invention;
[0050] Figure 4 is the general financial flowchart of the present invention;
[0051] Figure 5 is the journal business flowchart of the present invention;
[0052] Figure 6 is the newspaper business flowchart of the present invention. Detailed Description of the Preferred Embodiments
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0054] Please refer toFigure 1-6 , a news publishing statistical analysis system in the present invention, through informatization and intelligent technologies, solves the problems of low data processing efficiency, single function, and lack of automated analysis in the existing system, and realizes the rapid cleaning, intelligent calculation, dynamic display, and automatic report generation of news publishing statistical data.
[0055] System Architecture and Full Process Implementation
[0056] This system adopts a B / S architecture, with the front-end built based on the React framework, the back-end using FastAPI to provide RESTful interfaces, and the database using PostgreSQL. The data flow is as follows:
[0057] User uploads the original Excel file: The front-end receives the file through the upload component, and the back-end verifies the format and stores it as a Blob;
[0058] Data cleaning and preprocessing: The cleaning engine performs field mapping, missing value filling, and regular expression verification;
[0059] Structured storage: The cleaned data is serialized into JSON format for storage;
[0060] User interaction and analysis: The user selects the business module and metrics, and the back-end calls the calculation engine to generate results;
[0061] Dynamic report generation: Generate text descriptions and downloadable files through the template engine.
[0062] 1. Data Storage Solution
[0063] 1. Storage and Cleaning of Original Files
[0064] Blob Storage Mechanism:
[0065] After the original Excel file uploaded by the user is submitted through the front-end interface, the back-end stores it as binary data using the BYTEA type field in PostgreSQL. When the file is uploaded, the system automatically records the file name, upload time, and user information into the raw_files table to ensure the traceability of the original data. For example, after uploading the file "2023 Journal Data.xlsx", its binary content is directly stored in the database, retaining the original format and metadata.
[0066] Data Cleaning and Structured Storage:
[0067] Format Verification: The system checks whether the file extension is.xlsx. If it is not in the standard Excel format, an error message "File format error, only.xlsx is supported" is returned.
[0068] Field Mapping and Cleaning:
[0069] Map the user-defined column name (such as "print run") to the system preset field (such as print_volume).
[0070] Verify the ISBN number through regular expressions. For example, 978-3-16-148410-0 needs to meet the combination rule of 13 digits and hyphens. If the verification fails, it is marked as "invalid ISBN".
[0071] Missing value handling:
[0072] When a numerical field (such as "average print run") is missing, fill it with the mean value of data from similar journals. For example, if a science and technology journal is missing the print run, take the average value of all science and technology journals to fill it.
[0073] When a categorical field (such as "publisher name") is missing, mark it as "awaiting manual review" and highlight it on the front-end interface.
[0074] JSON serialization and storage: The cleaned data is converted into a DataFrame through Pandas and serialized into a JSON string, which is stored in the VARCHAR field of the cleaned_data table. For example, the cleaned journal data contains fields {"journal_name": "Science and Technology Journal A", "print_volume": 120000}.
[0075] 2. Database version management
[0076] Alembic migration process:
[0077] Initialize the migration environment: Execute alembic init migrations in the project root directory to generate the migration script directory.
[0078] Generate migration scripts: When the data model changes (such as adding a "journal category" field), execute alembic revision --autogenerate -m "add journal_category field" to automatically generate the difference script.
[0079] Execute the production environment migration: Synchronize the script to the production database through alembic upgrade head to ensure that the table structures in the development, testing, and production environments are consistent.
[0080] 2. Front-end technical solution
[0081] 1. Interface architecture and dynamic interaction
[0082] Integration of React and Vite:
[0083] The front end is built using the Vite toolchain, which supports module hot replacement (HMR). In the development environment, the page will be refreshed in real time after code modification. For example, after modifying the style of the filtering control, the interface will take effect without refreshing.
[0084] Chakra UI component library:
[0085] Upload component: Supports drag-and-drop upload and batch file selection, and the upload progress bar is displayed in real time.
[0086] Dynamic table: Integrates the Material Table component, and users can click on the table header to sort fields such as "average print run" and "total print run" in ascending / descending order.
[0087] Dynamic loading of enumeration values:
[0088] After the user uploads a file, the front end calls the / api / metadata interface to extract options such as the list of publishers and report types from the cleaned data and dynamically updates them to the dropdown selection box. For example, after uploading a file containing "Publisher A" and "Publisher B", the corresponding options will be automatically added to the "Publisher" selection box on the front end.
[0089] 2. Real-time filtering and review support
[0090] Multi-condition composite filtering: Users can combine and select categories (such as "science and technology periodicals"), time ranges (such as "2022 - 2023"), and index thresholds (such as "print run difference > 10%"), and the system will return matching data in real time.
[0091] Details viewing function: Click on a journal row in the table, and a modal box will pop up to display all its quarterly detailed data and render an ECharts trend chart to intuitively show the monthly fluctuations in the print run.
[0092] 3. Back-end technical solutions
[0093] 1. Interface architecture and deployment
[0094] FastAPI modular design:
[0095] Business module interfaces:
[0096] / api / journals: Handles journal classification, index calculation, and detailed data query.
[0097] / api / books: Supports ISBN aggregation, publisher analysis, and Chinese Library Classification statistics.
[0098] Finance module interfaces:
[0099] / api / finance / tax: Calculates the sum and difference of tax items.
[0100] / api / finance / salary: Generate the per capita salary report.
[0101] Gunicorn multi-process deployment:
[0102] By configuring 4 worker processes and a timeout (timeout = 120), it can support high-concurrency requests. For example, when 50 people operate simultaneously at the review site, the interface response time remains stable within 1 second.
[0103] 2. Implementation of the journal business module
[0104] 21. Classification rules and indicator calculation
[0105] Initial letter classification logic:
[0106] The system automatically classifies by the initial letter of the journal CIP:
[0107] Comprehensive category (Z): The initial letter of the journal CIP starts with "Z", such as "China Newsweek".
[0108] Science and technology category (N, O, P, etc.): Journals with the initial letter of the journal CIP being N, O, P, etc. are classified into this category, such as the Chinese version of "Nature".
[0109] Social philosophy category (A - K): The initial letter of the journal CIP is A - K, covering fields such as sociology and philosophy, such as "Philosophical Research".
[0110] Literature and art category (I, J): The initial letter of the journal CIP is I or J
[0111] Culture and education category (G, H): The initial letter of the journal CIP is G or H
[0112] Difference calculation and ranking:
[0113] Extract the data of the current year and the previous year, and calculate the indicator difference (Δ value):
[0114] ΔP = P current -P previous
[0115] Generate a ranking table according to the sorting basis selected by the user (Δ value or display value). For example, sort in descending order by "print run difference", and the top 10 journals in the ranking are displayed in green highlight.
[0116] 22. Report generation and details display
[0117] Dynamic text template:
[0118] The system automatically fills in data with the preset template and generates content such as: "
Technology
[0119] Trend visualization:
[0120] Call ECharts to generate a line chart to show the quarterly print run changes of a certain periodical in the past three years, and support exporting as a PNG image for embedding in the PDF report.
[0121] 3. Implementation of the book business module
[0122] 31. ISBN aggregation and status identification
[0123] Newly published / reprinted logic:
[0124] The system compares the current year's data with historical data. If an ISBN number appears for the first time, it is marked as "newly published"; otherwise, it is marked as "reprinted". For example, if ISBN "978-3-16-148410-0" already existed in 2022 and appears again in 2023, it is marked as "reprinted".
[0125] Classification statistics by the Chinese Library Classification (CLC) 22:
[0126] Count the number of titles, total print runs, and total printed sheets of books in each category according to the classification code (e.g., "I" represents literature).
[0127] Calculate the difference and percentage change between the two years.
[0128] 32. One-click download of the comparison table (partial)
[0129]
[0130] Conditional formatting: Rows with a print run difference exceeding 10% are marked with a yellow background, and rows below -5% are marked with a red background.
[0131] 4. Implementation of the financial module
[0132] 4.1 Tax calculation and difference analysis
[0133] Analysis of tax items:
[0134] Extract the following fields from the enterprise unit table:
[0135] Income tax (line 43)
[0136] Business tax (lines 50 and 28)
[0137] Value-added tax (line 51)
[0138] Administrative expense tax (line 31)
[0139] Calculation logic:
[0140] Sum of tax items = Line 43 + Line 50 + Line 51 + Line 31 + Line 28
[0141] Tax difference = Sum of items - Total tax (Line 49)
[0142] Exception handling: If the absolute value of the difference exceeds the threshold (e.g., 10,000 yuan), the system marks it as "Needs manual review" at the front end.
[0143] 4. Per capita salary statistics for 2 people
[0144] Formulas and data sources:
[0145]
[0146] Classification and comparison: Calculate the per capita salary by classifying into enterprises, institutions, and non-independent legal entities, and generate a bar chart for comparative analysis.
[0147] 4. Dynamic text generation technology
[0148] 1. Dynamic text generation technology
[0149] Natural Language Generation (NLG):
[0150] The system extracts key indicators (such as Δ value, year-on-year change rate) from the database.
[0151] Generate descriptive statements according to predefined logic:
[0152] If Δ value > 0, it is described as "growing"; if Δ value < 0, it is described as "declining".
[0153] Automatically match the reason library, such as "Reason for growth: Increased market demand" "Reason for decline: Impact of policy adjustment".
[0154] 2. Generation of downloadable files
[0155] Excel comparison table:
[0156] Generate a formatted table through the openpyxl library, support multiple sheets (such as "Journal analysis", "Financial statistics"), and set the freeze pane for easy viewing.
[0157] PDF report:
[0158] Use the ReportLab library to dynamically generate a PDF file containing text, tables, and charts, with automatic pagination and addition of headers and footers.
[0159] 5. Verification of example effects
[0160] 1. Efficiency comparison
[0161] Task Traditional system time consumption This system time consumption Efficiency improvement Data cleaning 2 hours 10 minutes 91.7% Review report generation 5 hours 40 minutes 86.7%
[0162] 2. Accuracy Verification
[0163] Data verification: Regular expressions reduce the ISBN error rate from 15% to 0.5%.
[0164] Automated calculation: The consistency rate between the analysis results of tax differences and manual accounting reaches 100%.
[0165] 3. User Experience Enhancement
[0166] Dynamic interaction: The real-time filtering and sorting functions shorten the review decision-making time by 70%.
[0167] Report readability: Visual charts and standardized text improve the readability of reports, and the user satisfaction survey reaches 95%.
[0168] This embodiment discloses a news publishing statistical analysis system, which realizes the Blob storage of the original Excel file and the JSON structured storage of the cleaned data through the PostgreSQL database, and uses Alembic for database version management; the front end is built based on React and Vite, integrating Chakra UI and Material Table to achieve dynamic filtering, sorting and enumeration value loading; the back end is built with FastAPI to build modular interfaces, and the business modules cover the classification aggregation and index calculation of periodicals, newspapers, audio-visual electronics, and books (such as the difference in average periodic print runs, the sum of small tax items and the difference), and the financial module supports per capita salary analysis and automated report generation, combined with dynamic text templates and downloadable files (Excel / PDF) output, to realize the full-process informatization of data cleaning, analysis, and display, significantly improving statistical efficiency and accuracy, and meeting the real-time decision-making needs of the review.
Claims
1. A news publishing statistical analysis system, which is implemented based on the BS architecture and is characterized by: include: Front-end interactive module: The user interface is built based on the React framework, which supports users to upload original Excel files, select business modules and indicator values, and display analysis results through web tables and Excel binary files; Backend processing module: built using FastAPI, including: Data cleaning engine: performs column name mapping, special character removal, missing value filling, and erroneous data correction on the uploaded Excel files to generate standardized structured data tables; Classification calculation engine: realizes data classification aggregation and index calculation of journals, newspapers, books, audio-visual electronics, and financial modules based on preset algorithms; Data storage module: Based on PostgreSQL database, the original Excel file is stored in Blob format, and the cleaned data is serialized into JSON format through Pandas; Report generation module: Dynamically generates explanatory text containing difference and year-on-year changes based on the indicator values selected by the user, and exports a downloadable Blob file.
2. A news publication statistical analysis system according to claim 1, characterized in that: The front-end interaction module further includes: Real-time verification of the format, required field integrity and data compliance of uploaded files, and return an error message if the verification fails; Provide dynamic filtering and sorting controls to support users to select analysis needs according to business modules (periodicals, newspapers, books, audio-visual electronics, finance) and indicator values (average issue print run, total print run, total pricing amount, etc.); Integrates interactive buttons of "View Details", "Copy Instructions", and "Download Comparison Files" to support real-time screening and sorting operations during material writing and review.
3. A news publication statistical analysis system according to claim 1, characterized in that: The data cleaning engine further comprises: Perform regular expression matching verification on the fields of the original Excel file to ensure that the ISBN number conforms to the international standard format and the date field conforms to the YYYY-MM-DD format; Automatically identify and delete duplicate data entries, missing data entries, and unsubmitted data entries, and fill missing fields with the average value of similar data or mark them for manual review. For abnormal data that does not meet the standards, appropriate adjustments are made. For example, for books with ISBN numbers corresponding to multiple titles, the title display operation is adopted, that is, all the titles corresponding to the ISBN numbers in the data are displayed on each piece of data.
4. A news publication statistical analysis system according to claim 1, characterized in that: The classification calculation engine further comprises: Journal analysis: automatically classify and aggregate data by CIP first letter category (Z for comprehensive category, NOPQRSTUVX for science and technology category, ABCDEFK for social philosophy category, and IJ for literature and art category), and calculate and analyze the difference and year-on-year change percentage of key business indicators within two years; Newspaper analysis: filter data by professional, comprehensive, life service, reader target, and abstract categories, extract and sort the index values of all newspapers, and generate reports on the difference between the number of categories and the total print run after classification and aggregation, as well as the year-on-year change percentage of key business indicators; Book analysis: Based on the publisher name and publishing model (new, reprint, rental) selected by the user, quarterly publishing details of each book are aggregated by ISBN number, including: number of prints, sheets, and total amount for each quarter. The algorithm automatically identifies the status of books from new to reprint / reprint to new, generates detailed Excel tables and overview text, and assists in writing materials; Audiovisual and electronic analysis: Aggregate data by carrier form (CD, DVD, etc.), aggregate the publication quantity and business indicators such as actual sales and code sales of this year and last year in each quarter into the same table; add corresponding auxiliary analysis columns to check whether the publication and distribution data (a total of 4 data sources) of this year and last year are consistent; add auxiliary columns to analyze whether there are discrepancies in the publication and distribution data of the two years (such as whether the language and distribution discount are unified); generate ranking tables and text descriptions of core business indicators for each unit and overview data. Financial analysis: Analyze the data of enterprises, institutions and non-independent legal entities, and automatically calculate the difference between the sum of per capita salary, tax sub-items and the total tax amount.
5. A news publication statistical analysis system according to claim 4, characterized in that: The journal analysis function further includes: After the user clicks the "View Details" button, all publication details and historical change trend charts of the specified journal are displayed; Use the "Copy Description" function to generate standardized description text containing key business indicators with one click.
6. A news publication statistical analysis system according to claim 4, characterized in that: The book analysis function further includes: Generate detailed information (Excel details table) on business indicators such as the number of prints, sheets, and total amount of books aggregated by ISBN number in each quarter, and generate a two-year comparative analysis report on indicators such as the number of types, total sheets, total pricing amount, and quarterly publishing details of general books (except for teaching aids, children's books, and textbooks) under the 22 categories of the Chinese Library Classification; Click the "Download publisher group ranking file" button to export the publishing index ranking file grouped by publisher. Click the "Download this publisher's two-year comparison file" button to export the publisher's two-year publication details comparison file. Automatically generate overview data and a text description of the overall situation of all categories of data from each publishing unit to assist in the writing of relevant materials.
7. A news publication statistical analysis system according to claim 4, characterized in that: The audio and video electronic analysis function further includes: Dynamically generate a ranking table in ascending or descending order based on the sorting criteria selected by the user (number of varieties, number of publications, difference in total pricing amount); Generate a two-year comparison file of publication and distribution containing unit code and ISBN number. Generate textual explanations of the core business indicators of the corresponding units to assist in writing the materials.
8. A news publication statistical analysis system according to claim 4, characterized in that: The financial analysis function further includes: Generate a complete explanation document including the difference in total tax, financial expenses and interest expense through the "one-click download explanation document" function; The change in the number of units within two years is counted by classification of enterprises, institutions and non-independent legal entities, and the year-on-year increase or decrease percentage of per capita salary is calculated.
9. A news publication statistical analysis system according to claim 1, characterized in that: The report generation module further comprises: Based on natural language generation technology, predefined templates and real-time data are dynamically combined to generate analysis text that meets the review requirements; Supports user-defined report templates, including table styles, text paragraph structure and key indicator priority settings.