Publication market demand analysis system based on data mining
Through a data mining-based publishing market demand analysis system, users and publications data are collected and analyzed in real time, and demand prediction models are built, which solves the problems of insufficient sample representativeness and inaccurate future demand prediction in the existing technology, and accurately analyzes and forecasts of market demand, helping authors adjust their publishing strategies.
Patent Information
- Application Number
- CN202510031916.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
When investigating the demand for the publishing market, the sample selection or insufficient representation is made inappropriate, resulting in the analysis results that cannot accurately reflect the overall market demand, and analysis based on historical data cannot accurately predict future demand, which requires a lot of time and capital investment.
Provide a publishing market demand analysis system based on data mining, including data acquisition module, user analysis module, publication analysis module, prediction model module and display module. Through big data algorithms and NLP technology, users and publication data can be collected and analyzed in real time, demand prediction models are built, and detailed visual reports are generated.
By collecting and analyzing data in real time, we can timely capture changes in the market's mainstream reading preferences, help authors keep up with market trends, adjust publishing strategies, provide accurate market demand analysis and forecasts, and reduce time and capital investment.
Smart Images

Figure CN119941307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of demand analysis, and in particular to a publishing market demand analysis system based on data mining. Background Art
[0002] With the continuous expansion of the readership and the growing demand for reading, novels and publications are emerging in an endless stream, covering a variety of themes such as ancient and modern stories, urban workplace, fantasy, etc. The development of digital technology and the popularization of online reading platforms have made the dissemination channels of novels more extensive, and readers can easily obtain and read various novels. Publishers and authors are also constantly exploring new content forms and presentation methods to meet the diverse needs of readers.
[0003] Then the existing survey methods for the publishing market mainly collect information about demand, preference, purchasing habits, etc. from target consumers through designed questionnaires, which results in improper sample selection or insufficient representativeness. The analysis results may not accurately reflect the overall market demand. The market environment and consumer preferences change rapidly. Analysis based on historical data may not be able to accurately predict future demand and requires a lot of time and financial investment.
[0004] In view of the above technical defects, a solution is now proposed. Summary of the invention
[0005] In view of the deficiencies of the prior art, the present invention provides a publishing market demand analysis system based on data mining.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a publishing market demand analysis system based on data mining, including: a data collection module, a user analysis module, a publication analysis module, a prediction model module and a display module.
[0007] The data collection module includes collection and preprocessing units. The collection unit collects publication types, sales information and user data, and transmits them to the corresponding analysis module. The preprocessing unit is used to clean the data, remove duplicate erroneous items, unify the format and process missing values.
[0008] The user analysis module receives the user behavior information from the data collection module, splits it, and divides the users into three levels: level one (core users), level two (less engaged users), and level three (less engaged users) based on age, reading time, and number of comments.
[0009] The publication parsing module includes a feature extraction unit and an analysis unit. The feature extraction unit uses big data algorithms and NLP to extract feature data and types from publications, and uses NLP technology to parse the content, classify features such as themes, plots, etc., and collect evaluation information. The analysis unit performs sentiment analysis to obtain a sentiment value R in the range of 0-1. The feature analysis unit calculates the feature heat value F based on sales volume M, reading volume C, and comment volume D.
[0010] The prediction model module will build a demand prediction model based on key factors such as feature heat value F, feature classification, author popularity, user market trends, etc. The model is continuously applied to new publications to generate feature predicted reading volume and sales volume. By comparing the difference between the predicted and actual reading volume, the model is adjusted and optimized.
[0011] The display module uses a big data analysis algorithm to comprehensively classify feature categories, popularity values and reading volumes, as well as feature-predicted reading volumes, to screen a category of user comments, and based on this, generates a detailed and intuitive visualization report, covering chart data, comment content, distribution characteristics, and user feedback trends.
[0012] The data collection module includes a collection unit and a preprocessing unit. The collection unit is used to collect and classify the types of publications, record the sales information of each publication in different time periods and different sales channels, collect user information and transmit the information to the preprocessing unit.
[0013] The preprocessing unit is used to preprocess the original data, clean the publication type information, sales information and user information, and remove duplicate, erroneous or invalid data; unify data in different formats and process missing values, fill in the missing values using reasonable methods, and then transmit the publication type sales and other information to the publication analysis module, and package the user registration information, purchase history, browsing behavior, social media interaction and other information into a user behavior information package and transmit it to the user analysis module.
[0014] The user analysis module is used to receive the user behavior information package transmitted by the data acquisition module, split the user information package, and use an algorithm to classify and label the users according to multiple dimensions such as user age, reading time and number of comments, and divide the users into level one, level two and level three. The level one users are a highly engaged core user group with moderate age, long reading time and large number of comments; the level two users are a user group with slightly lower engagement than the level one users; and the level three users are a user group with lower engagement.
[0015] The publication parsing module includes a feature extraction unit and an analysis unit. The feature extraction unit uses natural language processing technology to extract feature types and feature sales M, feature reading volume C and feature comment volume D from the publication content and evaluation. The feature analysis unit calculates the feature heat value F based on the publication feature sales M, feature reading volume C and feature comment volume D.
[0016] The feature extraction unit uses natural language processing technology (NLP) to parse the content of the publication, extracts and classifies feature information, and collects evaluation information from readers. The analysis unit performs sentiment analysis on the evaluation information to determine its positive, negative or neutral attitude, associates the evaluation with the feature data, and obtains a sentiment value R. The sentiment value range is 0 to 1, and the sentiment value of the positive comment is greater than the sentiment value of the negative comment.
[0017] The characteristic analysis unit calculates the characteristic heat value F according to the formula, and the specific formula is as follows:
[0018]
[0019] Among them, M represents the characteristic sales volume, C represents the characteristic reading volume, R represents the sentiment value, D represents the characteristic comment volume, X1 represents the weight coefficient of the characteristic sales volume, X2 represents the weight coefficient of the characteristic reading volume, X3 represents the weight coefficient of the sentiment value, and X4 represents the weight coefficient of the characteristic comment volume.
[0020] The prediction model module constructs a demand prediction model based on features, feature heat values F and multiple key factors. The demand prediction model is intermittently applied to new publications to generate corresponding feature predicted reading volumes and feature predicted sales volumes. The demand prediction model is adjusted and optimized by comparing the difference between feature predicted reading volumes and actual reading volumes and between feature predicted sales volumes and actual sales volumes.
[0021] The display module uses a big data analysis algorithm, combined with dimensions such as feature classification, feature heat value F, and feature reading volume, and then filters out some comments of a category of users, and generates a detailed visualization report based on these key information. The visualization report includes chart data and the filtered comment content, distribution characteristics, and user feedback trend chart.
[0022] The specific process of the demand analysis system is as follows:
[0023] S1. The data collection module collects data related to the types and sales volume of publications in the publishing market and user data, and performs preprocessing to obtain user information data and publication information data.
[0024] S2. Analyze user information data using the user analysis module and divide users according to their age, usage time and number of comments.
[0025] S3, the publication analysis module uses big data and NIP to divide and summarize the content and evaluation of publications and extract feature data and feature types, and calculates the feature heat value F based on data such as the number of readings and comments.
[0026] S4. Establish a demand forecasting model based on the characteristic heat value F and key factors, use the forecasting model to predict new publications and generate predicted reading volume, and verify and improve the forecasting model based on actual market feedback.
[0027] S5. The display module uses a big data analysis algorithm to filter out some comments of a category of users, and generates a visualization report based on feature classification, feature heat value F, feature reading volume, and comments of a category of users, and displays it to the author.
[0028] The present invention provides a publishing market demand analysis system based on data mining. Compared with the prior art, it has the following beneficial effects:
[0029] The present invention collects newer data and performs preprocessing through the collection module, so as to timely capture the changes in the mainstream reading preferences of the market, which helps authors to keep up with market trends and adjust publishing strategies. The work analysis module conducts in-depth analysis on the publication feature data to obtain the feature sales volume M, feature reading volume C, feature comment volume D and sentiment value R, and calculates the feature heat value F of these features, and quantitatively analyzes the high and low feature heat values F to intuitively show which features are popular in the current market.
[0030] The present invention divides users through a user analysis module, and then screens out comments from first-level users, which are more representative and have reference value. Combined with the prediction data given by the prediction model, the prediction data can provide a good reference. A detailed visualization report is produced using this information. The report charts the data and draws a trend change chart, and incorporates the commentary opinions in the first-level user comments, providing a reliable reference for publication authors. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of the principle framework of the present invention. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] Embodiment 1:
[0034] See also Figure 1 , the present application provides a publishing market demand analysis system based on data mining, including a data acquisition module, a user analysis module, a publication analysis module, a prediction model module and a display module.
[0035] The data collection module includes collection and preprocessing units. The collection unit collects publication types, sales information and user data, and transmits them to the corresponding analysis module. The preprocessing unit is used to clean the data, remove duplicate erroneous items, unify the format and process missing values.
[0036] The user analysis module receives the user behavior information from the data collection module, splits it, and divides the users into three levels: level one (core users), level two (less engaged users), and level three (less engaged users) based on age, reading time, and number of comments.
[0037] The publication parsing module includes a feature extraction unit and an analysis unit. The feature extraction unit uses big data algorithms and NLP to extract feature data and types from publications, and uses NLP technology to parse the content and classify features such as themes, plots, etc. It also collects evaluation information. The analysis unit performs sentiment analysis to obtain a sentiment value R in the range of 0-1. The feature analysis unit calculates the feature heat value F based on sales volume M, reading volume C, and comment volume D.
[0038] The prediction model module will build a demand prediction model based on key factors such as feature heat value F, feature classification, author popularity, user market trends, etc. This model is continuously applied to new publications to generate feature predicted reading volume and sales volume. By comparing the difference between the predicted and actual reading volume, the model is adjusted and optimized.
[0039] The display module uses big data analysis algorithms to comprehensively consider feature classification, popularity value and reading volume as well as feature predicted reading volume, screen a category of user comments, and generate a detailed and intuitive visualization report based on this, covering chart data, comment content, distribution characteristics and user feedback trends.
[0040] The data collection module includes a collection unit and a preprocessing unit. The collection unit is used to collect and classify the types of publications, record the sales information of each publication in different time periods and different sales channels, collect user information and transmit the information to the preprocessing unit.
[0041] The preprocessing unit is used to preprocess the original data, clean the publication type information, sales information and user information, and remove duplicate, erroneous or invalid data; unify data in different formats and process missing values, fill in the missing values using reasonable methods, and then transmit the publication type sales and other information to the publication parsing module, and package the user registration information, purchase history, browsing behavior, social media interaction and other information into a user behavior information package and transmit it to the user parsing module.
[0042] The publication parsing module includes a feature extraction unit and an analysis unit. The feature extraction unit uses natural language processing technology to extract feature types and feature sales M, feature reading volume C and feature comment volume D from the publication content and evaluation. The feature analysis unit calculates the feature heat value F based on the publication feature sales M, feature reading volume C and feature comment volume D.
[0043] The feature extraction unit uses natural language processing technology (NLP) to parse the content of the publication, extract and classify feature information, and collect evaluation information from readers. The analysis unit performs sentiment analysis on the evaluation information to determine its positive, negative or neutral attitude, and associates the evaluation with the feature data to obtain the sentiment value R. The sentiment value range is 0 to 1, and the sentiment value of positive comments is greater than that of negative comments.
[0044] The feature analysis unit calculates the feature heat value F according to the formula. The specific formula is as follows:
[0045]
[0046] Among them, M represents the characteristic sales volume, C represents the characteristic reading volume, R represents the sentiment value, D represents the characteristic comment volume, X1 represents the weight coefficient of the characteristic sales volume, X2 represents the weight coefficient of the characteristic reading volume, X3 represents the weight coefficient of the sentiment value, and X4 represents the weight coefficient of the characteristic comment volume.
[0047] Specific workflow:
[0048] The collection module collects the latest reading data in real time from major reading platforms, social media, bookstore sales systems and other channels, covering various types of publications. The collected raw data is cleaned to remove duplicate, erroneous and invalid data. The pre-processed data is analyzed in depth through the work analysis module to extract key features related to publications, including feature sales M (sales volume in a certain period of time), feature reading volume C (number of readings in a certain period of time), feature comment volume D (number of comments in a certain period of time) and sentiment value R. Based on feature data such as feature sales M, feature reading volume C, feature comment volume D and calculated sentiment value R, the feature heat value F is calculated using the formula, and a visual report is generated using visualization tools such as charts and reports, and the report is presented to the author.
[0049] Furthermore, the present invention collects newer data and performs preprocessing through the collection module, so as to timely capture the changes in the mainstream reading preferences of the market, which helps authors to keep up with market trends and adjust publishing strategies. The work analysis module conducts in-depth analysis of the publication feature data to obtain the feature sales volume M, feature reading volume C, feature comment volume D and sentiment value R, and calculates the feature heat value F of these features, and quantifies them through the high and low feature heat value F, intuitively showing which features are popular in the current market.
[0050] Embodiment 2:
[0051] The user analysis module is used to receive the user behavior information package transmitted by the data collection module, split the user information package, and use algorithms to classify and label users according to multiple dimensions such as user age, reading time and number of comments. Users are divided into level one, level two and level three. Level one users are a highly engaged core user group with moderate age, long reading time and a large number of comments; level two users are a user group with slightly lower engagement than level one users; and level three users are a user group with lower engagement.
[0052] The prediction model module builds a demand prediction model based on features, feature heat values F and multiple key factors. The demand prediction model will be intermittently applied to new publications to generate corresponding feature predicted reading volume and feature predicted sales volume. By comparing the difference between feature predicted reading volume and actual reading volume and feature predicted sales volume and actual sales volume, the demand prediction model is adjusted and optimized.
[0053] The display module uses big data analysis algorithms, combined with dimensions such as feature classification, feature heat value F, and feature reading volume, and then filters out some comments from a certain type of user, and generates a detailed visualization report based on these key information. The visualization report includes chart data and the content of the filtered comments, distribution characteristics, and user feedback trend charts.
[0054] The specific process of the demand analysis system is as follows:
[0055] S1. The data collection module collects data related to the types and sales volume of publications in the publishing market and user data, and performs preprocessing to obtain user information data and publication information data.
[0056] S2. Analyze user information data using the user analysis module and divide users according to their age, usage time and number of comments.
[0057] S3, the publication analysis module uses big data and NIP to divide and summarize the content and evaluation of publications and extract feature data and feature types, and calculates the feature heat value F based on data such as the number of readings and comments.
[0058] S4. Establish a demand forecasting model based on the characteristic heat value F and key factors, use the forecasting model to predict new publications and generate predicted reading volume, and verify and improve the forecasting model based on actual market feedback.
[0059] S5. The display module uses a big data analysis algorithm to filter out some comments of a category of users, and generates a visualization report based on feature classification, feature heat value F, feature reading volume, and comments of a category of users, and displays it to the author.
[0060] Specific workflow:
[0061] Through the user analysis module, users are accurately divided, and high-quality comments with reference value are screened out from the comments of first-level users. These comments often contain in-depth insights and unique insights into the publication. The prediction model module is used to build a prediction model, predict the future performance of the publication to obtain prediction data, and correct the prediction model by comparing the prediction data. The prediction data, feature sales volume M, feature reading volume C, feature comment volume D and sentiment value R are displayed in the form of charts, and a trend change chart is drawn. The key comments in the first-level user comments are added, and the produced visualization report is submitted to the publication author.
[0062] Furthermore, the present invention divides users through a user analysis module, and then screens out comments from first-level users, which are more representative and have reference value. Combined with the prediction data given by the prediction model, the prediction data can provide a good reference. A detailed visualization report is produced using this information. The report charts the data and draws a trend change graph, and incorporates the commentary opinions in the first-level user comments, providing a reliable reference for publication authors.
[0063] Some of the data in the above formulas are dimensionless and numerically calculated. Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.
[0064] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. The publishing market demand analysis system based on data mining is characterized by: include: Data collection module, user analysis module, publication analysis module, prediction model module and display module. The data collection module includes a collection unit and a preprocessing unit. The collection unit collects publication types, sales information and user data, and transmits them to the corresponding analysis module. The preprocessing unit is used to clean the data, remove duplicate erroneous items, unify the format and process missing values. The user analysis module receives the user behavior information from the data collection module, splits it, and divides the users into three levels: level one (core users), level two (less engaged users), and level three (less engaged users) based on age, reading time, and number of comments. The publication parsing module includes a feature extraction unit and an analysis unit. The feature extraction unit uses big data algorithms and NLP to extract feature data and types from publications, and uses NLP technology to parse the content, classify features such as themes, plots, etc., and collect evaluation information. The analysis unit performs sentiment analysis to obtain a sentiment value R in the range of 0-1. The feature analysis unit calculates the feature heat value F based on sales volume M, reading volume C, and comment volume D. The prediction model module will build a demand prediction model based on key factors such as feature heat value F, feature classification, author popularity, user market trends, etc. The model is continuously applied to new publications to generate feature predicted reading volume and sales volume. By comparing the difference between the predicted and actual reading volume, the model is adjusted and optimized. The display module uses a big data analysis algorithm to comprehensively classify feature categories, popularity values and reading volumes, as well as feature-predicted reading volumes, to screen a category of user comments, and based on this, generates a detailed and intuitive visualization report, covering chart data, comment content, distribution characteristics, and user feedback trends.
2. The publishing market demand analysis system based on data mining according to claim 1 is characterized in that: The data collection module includes a collection unit and a preprocessing unit. The collection unit is used to collect and classify the types of publications, record the sales information of each publication in different time periods and different sales channels, collect user information and transmit the information to the preprocessing unit. The preprocessing unit is used to preprocess the original data, clean the publication type information, sales information and user information, and remove duplicate, erroneous or invalid data; Unify data in different formats and handle missing values, fill in the missing values using reasonable methods, and then transmit information such as publication type sales to the publication analysis module. Package user registration information, purchase history, browsing behavior, social media interaction and other information into user behavior information packages and transmit them to the user analysis module.
3. The publishing market demand analysis system based on data mining according to claim 1 is characterized in that: The user analysis module is used to receive the user behavior information package transmitted by the data acquisition module, split the user information package, and use an algorithm to classify and label the users according to multiple dimensions such as user age, reading time and number of comments, and divide the users into level one, level two and level three. The level one users are a highly engaged core user group with moderate age, long reading time and large number of comments; the level two users are a user group with slightly lower engagement than the level one users; and the level three users are a user group with lower engagement.
4. The publishing market demand analysis system based on data mining according to claim 1 is characterized in that: The publication parsing module includes a feature extraction unit and an analysis unit. The feature extraction unit uses natural language processing technology to extract feature types and feature sales M, feature reading volume C and feature comment volume D from the publication content and evaluation. The feature analysis unit calculates the feature heat value F based on the publication feature sales M, feature reading volume C and feature comment volume D. The feature extraction unit uses natural language processing technology (NLP) to parse the content of the publication, extracts and classifies feature information, and collects evaluation information from readers. The analysis unit performs sentiment analysis on the evaluation information to determine its positive, negative or neutral attitude, associates the evaluation with the feature data, and obtains a sentiment value R. The sentiment value range is 0 to 1, and the sentiment value of the positive comment is greater than the sentiment value of the negative comment.
5. The publishing market demand analysis system based on data mining according to claim 1 is characterized in that: The characteristic analysis unit calculates the characteristic heat value F according to the formula, and the specific formula is as follows: Among them, M represents the characteristic sales volume, C represents the characteristic reading volume, R represents the sentiment value, D represents the characteristic comment volume, X1 represents the weight coefficient of the characteristic sales volume, X2 represents the weight coefficient of the characteristic reading volume, X3 represents the weight coefficient of the sentiment value, and X4 represents the weight coefficient of the characteristic comment volume.
6. The publishing market demand analysis system based on data mining according to claim 1 is characterized in that: The prediction model module constructs a demand prediction model based on features, feature heat values F and multiple key factors. The demand prediction model is intermittently applied to new publications to generate corresponding feature predicted reading volumes and feature predicted sales volumes. The demand prediction model is adjusted and optimized by comparing the difference between feature predicted reading volumes and actual reading volumes and between feature predicted sales volumes and actual sales volumes.
7. The publishing market demand analysis system based on data mining according to claim 1 is characterized in that: The display module uses a big data analysis algorithm, combined with dimensions such as feature classification, feature heat value F, and feature reading volume, and then filters out some comments of a category of users, and generates a detailed visualization report based on these key information. The visualization report includes chart data and the filtered comment content, distribution characteristics, and user feedback trend chart.
8. The publishing market demand analysis system based on data mining is characterized by: The specific process of the demand analysis system is as follows: S1. The data collection module collects data related to the types and sales volume of publications in the publishing market and user data, and performs preprocessing to obtain user information data and publication information data. S2. Analyze user information data using the user analysis module and divide users according to their age, usage time and number of comments. S3, the publication analysis module uses big data and NIP to divide and summarize the content and evaluation of publications and extract feature data and feature types, and calculates the feature heat value F based on data such as the number of readings and comments. S4. Establish a demand forecasting model based on the characteristic heat value F and key factors, use the forecasting model to predict new publications and generate predicted reading volume, and verify and improve the forecasting model based on actual market feedback. S5. The display module uses a big data analysis algorithm to filter out some comments of a category of users, and generates a visualization report based on feature classification, feature heat value F, feature reading volume, and comments of a category of users, and displays it to the author.