New media information management system based on AI
By designing a new media information management system based on AI, the efficiency and accuracy problems of traditional systems when facing massive data are solved, efficient collection, analysis and classification of new media data are achieved, personalized services are provided and the effectiveness of information dissemination is improved.
Patent Information
- Application Number
- CN202510342937.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
When facing massive data, traditional new media content management systems have problems such as slow content review speed, low personalized recommendation accuracy, difficulty in tracking hot topics in real time, and inadequate filtering of bad information.
A new media information management system based on AI was designed, including original data acquisition module, real-time information acquisition module, original text analysis module, text feature extraction module, user feature classification module, platform type classification module and information classification module. Through the collaborative work of these modules, efficient collection, analysis and classification of new media data can be achieved.
By efficiently collecting and processing a large amount of new media data, the system can timely and accurately understand the text content and classify it in detail, provide personalized services, improve the effectiveness of information dissemination, and improve the efficiency of new media information management.
Smart Images

Figure CN120218858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of new media technologies, and in particular, to a new media information management system based on AI. Background Art
[0002] With the development of the Internet, new media platforms such as social media and video sharing websites have become one of the main sources for people to obtain information. However, traditional content management systems often seem powerless in the face of massive data, and there are the following problems: slow content review speed, low accuracy of personalized recommendations, difficulty in real-time tracking of hot topics, and insufficiently timely and effective filtering of bad information. To solve these problems, it is particularly important to build a more intelligent new media information management system by combining AI technology.
[0003] Chinese Patent Publication No. CN118051631A discloses a method and system for information analysis and management of digital new media based on big data, which relates to the field of digital new media, including obtaining digital media data, and based on data preprocessing according to the digital media data, obtaining digital media standard data, and based on data mining according to the digital media classification database, obtaining digital media data feature information; thus, it can be seen that this invention does not comprehensively classify and analyze the received data based on the data in the existing database and the new media data received in real time, and there is a problem of low management efficiency for new media data. Summary of the Invention
[0004] The purpose of the present invention is to provide a new media information management system based on AI to solve at least one of the problems existing in the prior art.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A new media information management system based on AI, characterized by comprising: An original data acquisition module for acquiring original media data and user historical data; A real-time information acquisition module for acquiring media text data within a monitoring period; An original text analysis module for constructing a mapping relationship between original keywords and the original text field; A text feature extraction module for analyzing the keywords of the media text data within the monitoring period according to the original media data, and analyzing the text features of the media text data according to the keyword analysis result of the media text data and the construction result of the mapping relationship between the original keywords and the original text field; A user feature classification module for analyzing the browsing status and active status of the user according to the user historical data, and analyzing the user features according to the analysis results of the browsing status and active status of the user; A platform type classification module for classifying the types of platform users according to the results of user feature analysis; An information classification module for classifying the popularity of media text data according to the analysis results of the text features of the media text data and the analysis results of the platform user types within the monitoring period.
[0006] Furthermore, the original text analysis module is used to count the proportion μ(z,k) of the original text fields of each original keyword, and establish a mapping relationship between the original keyword and the original text field when μ(z,k) exceeds the mapping ratio threshold K, and set the correlation coefficient of this mapping relationship to f(z,k).
[0007] Furthermore, the text feature extraction module includes a text processing unit, which is used to generate a media text vocabulary according to the original text data, and perform word segmentation and data cleaning operations on the media text data within the monitoring period to obtain real-time media text phrases a[i], i ∈ N + , where a[i] represents the text phrase of the i-th media text data within the monitoring period; The text processing unit constructs a word segmentation vector α[i][j] of the real-time media text phrase a[i], j ∈ N + , where α[i][j] represents the word vector of the j-th text word in the text phrase of the i-th media text data.
[0008] Furthermore, the text feature extraction module further includes a keyword analysis unit, which is used to set the key score of each text word in the real-time media text phrase to score[i][j]; The keyword analysis unit constructs an undirected matrix D(i) of the real-time media text phrase, and calculates the iterative key score Score[i][j] of each text word according to the undirected matrix D(i) of the real-time media text phrase.
[0009] Furthermore, the text feature extraction module further includes a feature analysis unit, which is used to analyze the text features of the media text data according to the keyword analysis results of the media text data within the monitoring period. The text feature analysis results of the media text data include weak association features, general association features, strong association features, and timeliness association features.
[0010] Furthermore, the user feature classification module includes a browsing frequency analysis unit, which is used to compare the user's historical browsing frequency v with each preset browsing frequency, and analyze the active status of the user according to the comparison results. The active status of the user includes low active status, high active status, and normal status.
[0011] Further, the user feature analysis module further includes a historical content analysis unit, which is used to count the number N(k) of the original text fields browsed by the user historically, sort them in descending order, sum up the top three numbers in the sorting result, denote the sum result as NK, and calculate a partial ratio μ according to the sum result NK; The historical content analysis unit compares the partial ratio μ with a preset ratio U, and analyzes the browsing state of the user according to the comparison result. The browsing state of the user includes a divergent state and a convergent state.
[0012] Further, the user feature analysis module further includes a user feature analysis unit, which is used to analyze user features according to the analysis results of the user's browsing state and active state. User features include deep users, functional users, sticky users, and occasional users.
[0013] Further, the platform type classification module counts the number of users with various user features, and denotes the statistical results as n1, n2, n3, and n4, where n1 represents the number of deep users, n2 represents the number of functional users, n3 represents the number of sticky users, and n4 represents the number of occasional users; The platform type classification module is also used to classify the platform user type according to the user feature analysis result. The classification result of the platform user type includes a comprehensive type and a concentrated type.
[0014] Further, the information classification module includes an information classification unit, which is used to classify the media text data by heat according to the text feature analysis result of the media text data within the monitoring period and the analysis result of the platform user type: When the platform user type is the concentrated type, if the text feature of the media text data is a strong correlation feature, the information classification unit classifies the media text data as the first-level heat; if the text feature of the media text data is a general correlation feature or a weak correlation feature, the information classification unit classifies the media text data as the second-level heat; if the text feature of the media text data is a time-effect correlation feature, the information classification unit classifies the media text data as the first-level heat; When the platform user type is the comprehensive type, if the text feature of the media text data is a strong correlation feature, the information classification unit classifies the media text data as the second-level heat; if the text feature of the media text data is a general correlation feature or a weak correlation feature, the information classification unit classifies the media text data as the third-level heat; if the text feature of the media text data is a time-effect correlation feature, the information classification unit classifies the media text data as the time-effect heat.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the original data acquisition module and the real-time information acquisition module, the system can efficiently collect and process a large amount of original media data, user historical data, and the latest media text data within the monitoring period, ensuring the timeliness and accuracy of information; The original text analysis module and the text feature extraction module can more accurately understand the text content through in-depth analysis of keywords and their mapping relationships, and conduct detailed classification according to their characteristics; The user feature classification module and the platform type classification module can more accurately understand user preferences and behavior patterns by analyzing multi-dimensional information such as the browsing status and active status of users, so as to provide more personalized services and support for users; The information classification module classifies information according to the text features of media text data and the analysis results of platform user types, which helps the platform push the most suitable information to the most interested user groups and improve the effectiveness of information dissemination. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic structural diagram of the new media information management system based on AI in this embodiment.
[0018] Figure 2 It is a schematic structural diagram of the text feature extraction module in this embodiment.
[0019] Figure 3 It is a schematic structural diagram of the user feature classification module in this embodiment.
[0020] Figure 4 It is a schematic structural diagram of the information classification module in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to more clearly illustrate the present invention, the present invention will be further described below in conjunction with the preferred embodiments and the drawings. Similar components in the drawings are denoted by the same reference numerals. Those skilled in the art should understand that the specific content described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.
[0022] It should be noted that although terms such as first, second, and third may be used in the embodiments of the present application for description, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, without departing from the scope of the embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.
[0023] Specifically, the AI-based new media information management system described in this embodiment is applied to the statistics and management of new media information on a network service platform; the network service platform described in this embodiment is a text creation platform such as a blog, comment, or forum.
[0024] Please refer to Figure 1 as shown, which is a schematic structural diagram of the AI-based new media information management system described in this embodiment, including: A raw data acquisition module for acquiring raw media data and user historical data; the raw media data includes raw text data, raw keywords, and raw text fields; the raw text data is historical creative text within the network service platform; the raw keywords are keywords accumulated historically on the network service platform; the raw text fields are fields for classifying the raw text data on the network service platform; the user historical data includes the number of users, user historical browsing data, and user historical browsing frequencies.
[0025] Specifically, the raw media data described in this embodiment is multiple sets of data, each set of data includes a raw text data, multiple raw keywords, and a raw text field. The mapping relationship between the raw text field and the raw text data is one-to-many, the mapping relationship between the raw keywords and the raw text field is many-to-many, and the mapping relationship between the raw text data and the raw keywords is one-to-many; this embodiment does not specifically limit the acquisition methods of the raw media data and the user historical browsing data. Those skilled in the art can freely set them as long as the acquisition requirements of the raw media data and the user historical browsing data are met. In this embodiment, the raw media data and the user historical browsing data are obtained through the background server of the network service platform.
[0026] Please continue to refer to Figure 1 as shown, the system further includes a real-time information acquisition module for acquiring media text data within a monitoring period; the media text data includes media text and media text fields.
[0027] It can be understood that the acquisition methods of the media text data, the raw media data, and the user historical data in this embodiment are to obtain them through the server of the network service platform.
[0028] Specifically, this embodiment does not specifically limit the value of the duration of the monitoring period. Those skilled in the art can freely set it as long as the value requirement of the duration of the monitoring period is met. In this embodiment, the value of the duration of the monitoring period is 3 hours; it can be understood that the content of the media text field described in this embodiment is the same as that of the raw text field, but the media text and the raw text data to which they belong are different, so they have different names for distinction.
[0029] Please continue to refer to Figure 1 As shown, the system further includes an original text analysis module, which is connected to the original data acquisition module. The original text analysis module is used to construct a mapping relationship between the original keywords and the original text fields, and set a correlation coefficient to represent the degree of association between the original keywords and the original text fields; The original text analysis module is used to count the proportion μ(z,k) of the original text fields of each original keyword, and construct a mapping relationship between the original keywords and the original text fields according to the statistical results: if μ(z,k) < K, the original text analysis module does not establish a mapping relationship between the original keyword and the original text field; if μ(z,k) ≥ K, the original text analysis module establishes a mapping relationship between the original keyword and the original text field, and sets the correlation coefficient of this mapping relationship to f(z,k), where f(z,k)=[μ(z,k)-K] / K; where μ(z,k) is numerically defined as the ratio of the number of occurrences of the z-th original keyword in the k-th original text field to the total number of occurrences of the z-th original keyword, z ∈ N + , k ∈ N + , and K is a mapping ratio threshold; through numerical comparison, to determine the mapping relationship between the original keywords and the original text fields, and set a correlation coefficient to numerically represent their degree of association on the premise that a mapping relationship is determined to exist.
[0030] Specifically, in this embodiment, the value of the mapping ratio threshold K is not specifically limited, and those skilled in the art can freely set it as long as it meets the value requirements of the mapping ratio threshold K. The optimal value of the mapping ratio threshold K in this embodiment is 0.53.
[0031] Please continue to refer to Figure 1 As shown, the system further includes a text feature extraction module, which is connected to the original text analysis module and the real-time information acquisition module. The text feature extraction module is used to analyze the keywords of the media text data within the monitoring period according to the original media data, and analyze the text features of the media text data according to the keyword analysis results of the media text data.
[0032] Please refer to Figure 2 As shown, the text feature extraction module includes a text processing unit, which is used to preprocess the media text within the monitoring period according to the original text data and the original keywords to obtain real-time media text phrases, and construct word vectors of the real-time media text phrases to numerically represent the real-time media text phrases with word vectors; The text processing unit is used to generate a media text vocabulary based on the original text data, and perform word segmentation and data cleaning operations on the media text data within the monitoring period according to the media text vocabulary, so as to obtain real-time media text phrases a[i], where i ∈ N + , where a[i] represents the text phrase of the i-th media text data within the monitoring period; The text processing unit constructs a word segmentation vector α[i][j] of the real-time media text phrase a[i] according to the word vector file, where j ∈ N + , where α[i][j] represents the word vector of the j-th text word in the text phrase of the i-th media text data; by processing the original text data, real-time media text phrases adapted to the current platform are obtained, and then word vectors are assigned to the real-time media text phrases, providing data support for the subsequent mathematical operation process of keyword extraction.
[0033] Specifically, in this embodiment, the technical means of "generating a media text vocabulary based on the original text data" is not specifically limited, and those skilled in the art can set it freely. In this embodiment, the original text data is input into the Llama2 model for training to output the media text vocabulary. Moreover, in this example, the technical means of "performing word segmentation and data cleaning operations on the media text data within the monitoring period according to the media text vocabulary" is a prior art and can be implemented through the Python programming language, which will not be elaborated in this embodiment; in this embodiment, the word vector file is obtained by acquiring a pre-trained word vector file through GitHub and inputting the pre-trained word vector file into the Llama2 model for training to obtain a word vector file adapted to this embodiment. The vector dimension of the word vector file in this embodiment is 64.
[0034] Please continue to refer to Figure 2 As shown, the text feature extraction module further includes a keyword analysis unit, which is connected to the text processing unit. The keyword analysis unit is used to analyze the keywords of the media text data according to the construction result of the word segmentation vector of the real-time media text phrase and the construction result of the mapping relationship between the original keywords and the original text field; The keyword analysis unit sets the key score of each text word in the real-time media text phrase as score[i][j], and initializes the key score of each text word to 1 / Ni, where Ni represents the number of text words in the i-th media text phrase; by initializing and assigning equal values to each text word, it is convenient for the subsequent operation of iteratively calculating the key score; The keyword analysis unit constructs an undirected matrix D(i) of the real-time media text phrase, and calculates the iterative key score Score[i][j] of each text word according to the undirected matrix D(i) of the real-time media text phrase. It is set that ; where, IN(j) represents the set of text words that appear in the same window as the j-th text word of the i-th media text, D(i)[j][c] represents the element value of the c-th column in the j-th row of the undirected matrix, D(i)[j][d] represents the element value of the d-th column in the j-th row of the undirected matrix, η is the damping coefficient, and 0.6 < η < 1; The keyword analysis unit sorts the iterative key scores Score[i][j] of each text word in the media text in descending order, and takes the top 3 sorted text words as the keywords of the media text; by setting an undirected matrix to numerically represent the relationships between text words in the media text one by one, and using matrix elements to represent the relationships between each text word, the keywords in the media text data are mathematically extracted, realizing the efficient extraction of keywords in the media text data, and providing data support for the subsequent text feature analysis process of the media text data.
[0035] Specifically, the construction process of the undirected matrix D of the real-time media text phrase in this embodiment is to set the number of window words N, and calculate the matrix element D(i)[a][b] between text words that appear in the same window, and set D(i)[a][b] = {α[i][a]·α[i][b]} / {|α[i][a]|×|α[i][b]|}; set the matrix elements between text words that do not appear in the same window to 0; at the same time, in this embodiment, the value of the damping coefficient η is not specifically limited, and those skilled in the art can freely set it as long as the value requirement of the damping coefficient is met. The best value of the damping coefficient η is 0.65.
[0036] Please continue to refer to Figure 2As shown, the text feature extraction module further includes a feature analysis unit. The feature analysis unit is connected to the keyword analysis unit. The feature analysis unit is used to analyze the text features of the media text data according to the keyword analysis results of the media text data within the monitoring period, and use the analysis results of the text features of each media text data to represent the degree to which the media text data matches its media field: When there is a mapping relationship between the keywords of the media text data and the media text field, the feature analysis unit calculates the feature heat β(i) of the media text data, and sets β(i) = {f(1)(k) × Score1 + f(2)(k) × Score2 + f(3)(k) × Score3} / [Score1 + Score2 + Score3]. At this time, if β(i) < B1, the feature analysis unit determines that the text feature of this media text data is a weak association feature; if B1 ≤ β(i) < B2, the feature analysis unit determines that the text feature of this media text data is a general association feature; if β(i) ≥ B1, the feature analysis unit determines that the text feature of this media text data is a strong association feature; When there is no mapping relationship between the keywords of the media text data and the media text field, the feature analysis unit determines that the text feature of this media text data is a timeliness association feature; where f(1)(k) represents the correlation coefficient of the mapping relationship between the first keyword and the media field, f(2)(k) represents the correlation coefficient of the mapping relationship between the second keyword and the media field, f(3)(k) represents the correlation coefficient of the mapping relationship between the third keyword and the media field, Score1 represents the iterative key score of the first keyword, Score2 represents the iterative key score of the second keyword, Score3 represents the iterative key score of the third keyword, B1 is the first preset correlation coefficient, B2 is the second preset correlation coefficient, and B1 < B2; Analyze the text features of the media text data based on the keyword analysis results and mapping relationship of the media text data.
[0037] Specifically, in this embodiment, the values of the first preset correlation coefficient B1 and the second preset correlation coefficient B2 are not specifically limited, and those skilled in the art can freely set them as long as the value requirements of the first preset correlation coefficient B1 and the second preset correlation coefficient B2 are met. In this embodiment, the value of the first preset correlation coefficient B1 is 0.3, and the value of the second preset correlation coefficient B2 is 0.5.
[0038] Please continue to refer to Figure 1 As shown, the system further includes a user feature classification module. The user feature classification module is connected to the original data acquisition module. The user feature classification module is used to analyze the user features according to the user historical data.
[0039] Please refer to Figure 3As shown, the user feature classification module includes a browsing frequency analysis unit. The browsing frequency analysis unit is used to compare the user's historical browsing frequency v with each preset browsing frequency, and analyze the user's active status based on the comparison result: if v < V1, the browsing frequency analysis unit determines that the user's active status is a low-active status; if v ≥ V2, the browsing frequency analysis unit determines that the user's active status is a high-active status; if V1 ≤ v < V2, the browsing frequency analysis unit determines that the user's active status is a normal status; where V1 is the first preset user browsing frequency, V2 is the second preset user browsing frequency, and V1 < V2. By comparing and analyzing the user's browsing frequency, the active status of the user is judged, and an accurate judgment of the user's active status is achieved, and thus an accurate analysis of the user's features is realized.
[0040] Specifically, in this embodiment, the values of the first preset user browsing frequency V1 and the second preset user browsing frequency V2 are not specifically limited, and those skilled in the art can freely set them as long as they meet the value requirements of the first preset user browsing frequency V1 and the second preset user browsing frequency V2. In this embodiment, the value of the first preset user browsing frequency V1 is 10 times per day, and the value of the second preset user browsing frequency V2 is 20 times per day.
[0041] Please continue to refer to Figure 3 As shown, the user feature analysis module further includes a historical content analysis unit. The historical content analysis unit is used to analyze the user's browsing status based on the user's historical browsing data, and use the user's browsing status to represent the concentration degree of the user's browsing content. The historical content analysis unit counts the number N(k) of the original text fields browsed by the user historically, and sorts them in descending order. The historical content analysis unit sums the quantities of the top three in the sorting result, and records the summation result as NK. The historical content analysis unit calculates a partial ratio μ according to the summation result NK, and sets μ = NK / NY, where NY is the quantity of the original text data browsed by the user historically. The historical content analysis unit compares the partial ratio μ with the preset ratio U, and analyzes the user's browsing status based on the comparison result: if μ < U, the historical content analysis unit determines that the user's browsing status is a divergent status; if μ ≥ U, the historical content analysis unit determines that the user's browsing status is a convergent status. By integrating and analyzing the user's historical browsing data, the user's browsing habit (that is, the concentration degree of the user's browsing content) is judged from the data, and is represented by the user's browsing status, and an accurate analysis of the user's browsing status is achieved, and thus an accurate analysis of the user's features is realized.
[0042] Specifically, in this embodiment, the value of the preset ratio U is not specifically limited, and those skilled in the art can freely set it as long as the value requirement of the preset ratio U is met. In this embodiment, the optimal value of the preset ratio U is 60%.
[0043] Please continue to refer to Figure 3 As shown, the user feature analysis module further includes a user feature analysis unit. The user feature analysis unit is connected to the browsing frequency analysis unit and the historical content analysis unit. The user feature analysis unit is used to analyze user features based on the analysis results of the user's browsing state and active state: If the user's active state is a high-active state or a normal state and the user's browsing state is a convergent state, the user feature analysis unit determines that the user feature is a deep user; If the user's active state is a low-active state and the user's browsing state is a convergent state, the user feature analysis unit determines that the user feature is a functional user; If the user's active state is a high-active state and the user's browsing state is a divergent state, the user feature analysis unit determines that the user feature is a sticky user; If the user's active state is a normal state or a low-active state and the user's browsing state is a divergent state, the user feature analysis unit determines that the user feature is an occasional user; By comprehensively evaluating the user features under the current platform by combining the user's active state and the user's browsing state, and using deep users, functional users, sticky users, and occasional users to represent various types of users, the concise and accurate analysis and judgment of user features are realized, and the accurate analysis of user features is realized.
[0044] Please continue to refer to Figure 1 As shown, the system further includes a platform type classification module. The platform type classification module is connected to the user feature analysis module. The platform type classification module is used to classify the platform user types according to the user feature analysis results; The platform type classification module counts the number of users with various user features, and records the statistical results as n1, n2, n3, and n4. Among them, n1 represents the number of deep users, n2 represents the number of functional users, n3 represents the number of sticky users, and n4 represents the number of occasional users; The platform type classification module is also used to classify the platform user types according to the user characteristic analysis results: if n1 / (n1 + n2 + n3 + n4) < MA, the platform type classification module determines that the platform user type is a comprehensive type; if n1 / (n1 + n2 + n3 + n4) ≥ MA, the platform type classification module determines that the platform user type is a concentrated type; where MA is a preset user proportion; by the platform type analysis module, the customer types of the platform are judged based on the user characteristic analysis results of all users within the platform, and this is used as one of the data for predicting the popularity level of media text information, improving the accuracy of the popularity classification of media text data, and further improving the management efficiency of new media information.
[0045] Specifically, in this embodiment, no specific limitation is imposed on the value of the preset user proportion MA. Those skilled in the art can freely set it as long as it meets the value requirements of the preset user proportion MA. In this embodiment, the value of the preset user proportion MA is 0.4.
[0046] Please continue to refer to Figure 1 As shown, the system further includes an information classification module. The information classification module is connected to the platform type classification module and the text feature extraction module. The information classification module is used to classify the popularity of media text data according to the text feature analysis results of media text data within the monitoring period and the analysis results of platform user types.
[0047] Please refer to Figure 4 As shown, the information classification module includes an information classification unit. The information classification unit is used to classify the popularity of media text data according to the text feature analysis results of media text data within the monitoring period and the analysis results of platform user types: When the platform user type is a concentrated type, if the text feature of the media text data is a strong correlation feature, the information classification unit classifies the media text data as a first-level popularity; if the text feature of the media text data is a general correlation feature or a weak correlation feature, the information classification unit classifies the media text data as a second-level popularity; if the text feature of the media text data is a timeliness correlation feature, the information classification unit classifies the media text data as a first-level popularity; When the platform user type is of the comprehensive type, if the text feature of the media text data is a strong correlation feature, the information classification unit classifies the media text data as the secondary popularity level; if the text feature of the media text data is a general correlation feature or a weak correlation feature, the information classification unit classifies the media text data as the tertiary popularity level; if the text feature of the media text data is a timeliness correlation feature, the information classification unit classifies the media text data as the timeliness popularity level; by comprehensively considering the platform user type and the text feature of the media text data through the information classification unit to classify the media text data by popularity, it is convenient for the subsequent platform to push and manage the media text data, improves the accuracy of the popularity classification of the media text data, and further improves the management efficiency of the new media information.
[0048] Specifically, the primary popularity level, secondary popularity level, and tertiary popularity level described in this embodiment represent the decreasing order of predicting the information popularity within the platform, and the timeliness popularity level represents the high-popularity information within a short period of time.
[0049] Please continue to refer to Figure 4 As shown, the information classification unit further includes a storage unit, and the storage unit is used to store the media text data according to the popularity classification result of the media text data.
[0050] Specifically, this embodiment does not specifically limit the process of storing the media text data, which is a prior art and will not be elaborated in this embodiment. The media text data is stored in the platform database.
[0051] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A new media information management system based on AI, characterized in that: include: The original data collection module is used to collect original media data and user historical data; Real-time information collection module, used to collect media text data within the monitoring period; The original text analysis module is used to construct the mapping relationship between the original keywords and the original text fields; A text feature extraction module is used to analyze the keywords of the media text data within the monitoring period according to the original media data, and analyze the text features of the media text data according to the keyword analysis results of the media text data and the mapping relationship construction results between the original keywords and the original text fields; A user feature classification module is used to analyze the browsing status and active status of users based on user history data, and to analyze user features based on the results of the browsing status and active status analysis; The platform type classification module is used to classify the platform user types according to the user feature analysis results; The information classification module is used to classify the popularity of media text data based on the text feature analysis results of the media text data and the analysis results of the platform user types within the monitoring period.
2. The AI-based new media information management system according to claim 1, characterized in that: The original text analysis module is used to count the proportion μ(z, k) of the original text field of each original keyword, and when μ(z, k) exceeds the mapping proportion threshold K, a mapping relationship between the original keyword and the original text field is established, and the correlation coefficient of this mapping relationship is set to f(z, k).
3. The AI-based new media information management system according to claim 2 is characterized in that: The text feature extraction module includes a text processing unit, which is used to generate a media text vocabulary according to the original text data, and perform word segmentation and data cleaning operations on the media text data within the monitoring period according to the media text vocabulary to obtain real-time media text phrases a[i], i∈N + , a[i] represents the text phrase of the i-th media text data in the monitoring period; The text processing unit constructs the word segmentation vector α[i][j], j∈N of the real-time media text phrase a[i] according to the word vector file. + ,α[i][j] represents the word vector of the jth text word in the text phrase of the i-th media text data.
4. The AI-based new media information management system according to claim 3 is characterized in that: The text feature extraction module further includes a keyword analysis unit, which is used to set the key score of each text word in the real-time media text phrase to score[i][j]; The keyword analysis unit constructs an undirected matrix D(i) of the real-time media text phrase, and calculates an iterative keyword score Score[i][j] of each text word according to the undirected matrix D(i) of the real-time media text phrase.
5. The AI-based new media information management system according to claim 4 is characterized in that: The text feature extraction module also includes a feature analysis unit, which is used to analyze the text features of the media text data according to the keyword analysis results of the media text data within a monitoring period. The text feature analysis results of the media text data include weak correlation features, general correlation features, strong correlation features and time-related correlation features.
6. The AI-based new media information management system according to claim 5, characterized in that: The user feature classification module includes a browsing frequency analysis unit, which is used to compare the user's historical browsing frequency v with each preset browsing frequency, and analyze the user's active state according to the comparison result. The user's active state includes a low active state, a high active state and a normal state.
7. The AI-based new media information management system according to claim 6, characterized in that: The user feature analysis module further includes a historical content analysis unit, which is used to count the number N(k) of original text fields historically browsed by the user and sort them in descending order, and the historical content analysis unit sums the numbers of the first three digits of the sorting results and records the summation result as NK, and also calculates the partial proportion μ according to the summation result NK; The historical content analysis unit compares the partial ratio μ with the preset ratio U, and analyzes the browsing state of the user according to the comparison result, where the browsing state of the user includes a divergent state and a convergent state.
8. The AI-based new media information management system according to claim 7, characterized in that: The user feature analysis module also includes a user feature analysis unit, which is used to analyze user features according to the user's browsing status analysis results and active status analysis results. User features include deep users, functional users, sticky users and occasional users.
9. The AI-based new media information management system according to claim 8, characterized in that: The platform type classification module counts the number of users with each type of user characteristics, and records the statistical results as n1, n2, n3 and n4, where n1 represents the number of deep users, n2 represents the number of functional users, n3 represents the number of sticky users, and n4 represents the number of occasional users; The platform type classification module is also used to classify platform user types according to the user feature analysis results, and the classification results of platform user types include comprehensive types and concentrated types.
10. The AI-based new media information management system according to claim 9, characterized in that: The information classification module includes an information classification unit, which is used to perform heat classification on the media text data according to the text feature analysis results of the media text data and the analysis results of the platform user type within the monitoring period: When the platform user type is a concentrated type, if the text feature of the media text data is a strong correlation feature, the information classification unit classifies the media text data as the first-level heat; if the text feature of the media text data is a general correlation feature or a weak correlation feature, the information classification unit classifies the media text data as the second-level heat; if the text feature of the media text data is a time-sensitive correlation feature, the information classification unit classifies the media text data as the first-level heat; When the platform user type is a comprehensive type, if the text feature of the media text data is a strong correlation feature, the information classification unit classifies the media text data as secondary heat; if the text feature of the media text data is a general correlation feature or a weak correlation feature, the information classification unit classifies the media text data as tertiary heat; if the text feature of the media text data is a time-related correlation feature, the information classification unit classifies the media text data as time-related heat.
Citation Information
Patent Citations
Information analysis and management method and system for digital new media based on big data
CN118051631A