A risk prediction method for news network data based on large models
Through a large-scale model-based method, a multi-level analysis and screening model is established, which solves the problem of insufficient efficiency and accuracy of news network data processing in the existing technology, and realizes rapid risk prediction processing and efficient monitoring and analysis.
Patent Information
- Application Number
- CN202411854177.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The existing technology is difficult to effectively monitor and process large amounts of news network data in real time, resulting in the inability to guarantee processing efficiency and accuracy, and the traditional manual review speed is slow, which cannot meet the requirements of Internet data processing.
A large-scale model-based news network data risk prediction processing method is adopted, and a multi-level analysis and screening sub-model is established to realize the risk prediction processing of data by extracting text data and associated communication characteristics of news network data.
It realizes rapid risk prediction and processing of news network data, improves processing efficiency and accuracy, and can effectively monitor and analyze large amounts of data to meet the needs of Internet data processing.
Smart Images

Figure CN119313170B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of comprehensive analysis and processing of news network data, and in particular to a news network data risk prediction and processing method based on a large model. Background Art
[0002] As an important part of Internet data, news network data has the characteristics of rapid dissemination, short exposure time, and netizens' opinions compared with traditional media, and its value is also increasing. If we want to monitor the development and changes of online news in real time, the speed of traditional manual review is too slow, and the processing efficiency and accuracy of large amounts of data cannot be guaranteed. The singleness of its review content objectives also gradually fails to meet the requirements of Internet data processing. Summary of the invention
[0003] In view of the shortcomings of the prior art, the present invention provides a news network data risk prediction processing method based on a large model. By extracting the text data and related communication features of the news network data and establishing a multi-level analysis and screening sub-model based on the large model, the final risk prediction processing results can be quickly output by accurately analyzing and classifying text types.
[0004] To achieve the above-mentioned purpose, the present invention provides a news network data risk prediction processing method based on a large model, which is characterized by comprising:
[0005] S1, using news network data to perform visualization processing to obtain text features of news network data;
[0006] S2, using the text features of the news network data to establish a data correlation analysis and screening model based on a large model;
[0007] S3. Obtaining a news network data risk prediction processing result according to the data correlation analysis screening model.
[0008] Preferably, the text features of the news network data obtained by visualizing the news network data include:
[0009] Acquire text data of news network data and perform text noise reduction processing to obtain noise-reduced text data of news network data;
[0010] Using the noise-reduced text data of the news network data to perform word segmentation processing to obtain a text word segmentation set of the news network data;
[0011] Using the noise-reduced text data of the news network data to perform part-of-speech tagging to obtain text part-of-speech tags for the news network data;
[0012] Using the denoised text data, text segmentation set and text part-of-speech tagging of the news network data as text features of the news network data;
[0013] The text denoising process is to remove meaningless content from text data.
[0014] Furthermore, using the text features of the news network data to establish a data correlation analysis and screening model based on a large model includes:
[0015] Using the text features of the news network data to establish a text data word vector classification model based on the big model;
[0016] Acquire relevant feature data of the news network data according to the text features of the news network data;
[0017] Establishing a correlation feature analysis model based on a large model using the correlation feature data of the news network data;
[0018] The text data word vector classification model and the association feature analysis model are used as a data association analysis screening model.
[0019] Furthermore, using the text features of the news network data to establish a text data word vector classification model based on a large model includes:
[0020] Dividing the denoised text data into a first data set using the text features of the news network data;
[0021] Dividing the text word set into a second data set using the text features of the news network data;
[0022] Using the text features of the news network data and corresponding text part-of-speech tags to divide the data into a third data set;
[0023] The text features of the news network data are used to correspond to the relative positions of the text segmentation set in the noise reduction text data to divide it into a fourth data set;
[0024] The first data set, the second data set and the third data set are used as input, and the fourth data set is used as output, and training is performed based on the large model to obtain a text data word vector classification model.
[0025] Furthermore, obtaining the associated feature data of the news network data according to the text features of the news network data includes:
[0026] According to the text features of the news network data, the corresponding local port IP address and server port IP address are respectively obtained;
[0027] Acquire the corresponding news network website IP address according to the text features of the news network data;
[0028] Obtaining a classification label of news network data corresponding to the news network data;
[0029] The local port IP address, the server port IP address, the news network website IP address and the classification label of the news network data are used as the associated characteristic data of the news network data.
[0030] Furthermore, using the associated feature data of the news network data to establish an associated feature analysis model based on the big model includes:
[0031] Using the local port IP address corresponding to the associated characteristic data of the news network data as the fifth data set;
[0032] Using the server port IP address corresponding to the associated characteristic data of the news network data as the sixth data set;
[0033] Using the associated characteristic data of the news network data corresponding to the news website IP address as the seventh data set;
[0034] Using the associated feature data of the news network data to correspond to the classification labels of the news network data as an eighth data set;
[0035] Using the fifth data set and the sixth data set as input, and the corresponding relationship between the fifth data set and the sixth data set as output, training is performed based on the large model to obtain a port IP address association model;
[0036] Using the sixth data set and the seventh data set as input, and the corresponding relationship between the sixth data set and the seventh data set as output, training is performed based on the large model to obtain a port website association model;
[0037] Using the seventh data set and the eighth data set as input, and the corresponding relationship between the seventh data set and the eighth data set as output, training is performed based on the large model to obtain a news website classification association model;
[0038] The port IP address association model, the port website association model and the news website classification association model are used as association feature analysis models.
[0039] Furthermore, the news network data risk prediction processing results obtained according to the data correlation analysis screening model include:
[0040] S3-1, using the data correlation analysis and screening model to establish a comprehensive data matrix of news network data;
[0041] S3-2. Obtaining a news network data risk prediction processing result based on the comprehensive data matrix of the news network data.
[0042] Furthermore, the comprehensive data matrix of news network data established by using the data correlation analysis and screening model includes:
[0043] S3-1-1, using the data association analysis screening model to correspond to the text data word vector classification model to obtain the model output result of the text data word vector classification model as the main feature of the basic news network data;
[0044] S3-1-2, using the data correlation analysis screening model to correspond to the correlation feature analysis model to obtain the model output result of the correlation feature analysis model as the secondary feature of the basic news network data;
[0045] S3-1-3. Establish a comprehensive data matrix of news network data using the main features of the basic news network data and the secondary features of the basic news network data.
[0046] Furthermore, using the main features of the basic news network data and the secondary features of the basic news network data to establish a comprehensive data matrix of the news network data includes:
[0047] S3-1-3-1, obtaining historical normal text features and historical abnormal text features respectively according to the text features of the news network data corresponding to the main features of the basic news network data;
[0048] S3-1-3-2, obtaining the central word of the historical normal text feature as the historical normal text reference word;
[0049] S3-1-3-3, obtaining the central word of the historical abnormal text feature as the historical abnormal text reference word;
[0050] S3-1-3-4, using the historical normal text feature corresponding to the text word set to establish a historical normal text vector according to the historical normal text benchmark word;
[0051] S3-1-3-5, using the historical abnormal text feature corresponding to the text word set to establish a historical abnormal text vector according to the historical abnormal text benchmark word;
[0052] S3-1-3-6, according to the historical normal text features, respectively obtain the corresponding historical normal basic news network data secondary features as historical normal association labels;
[0053] S3-1-3-7, according to the historical abnormal text features, respectively obtain the corresponding historical abnormal basic news network data secondary features as historical abnormality association labels;
[0054] S3-1-3-8. Use the historical normal text vectors, historical abnormal text vectors, historical normal associated labels and historical abnormal associated labels as a comprehensive data matrix of news network data.
[0055] Furthermore, the news network data risk prediction processing results obtained according to the comprehensive data matrix of the news network data include:
[0056] S3-2-1, determine whether the text feature of the news network data has a historical normal text benchmark word corresponding to the comprehensive data matrix of the news network data, if so, the text benchmark word screening state is normal, and directly execute S3-2-3, otherwise, execute S3-2-2;
[0057] S3-2-2, determine whether the text features of the news network data have historical abnormal text benchmark words corresponding to the comprehensive data matrix of the news network data. If so, the text benchmark word screening status is abnormal, and execute S3-2-3; otherwise, delete the current historical normal text benchmark words and historical abnormal text benchmark words, and return to S3-1-3-2;
[0058] S3-2-3, obtaining the similarity between the main features of the basic news network data and the historical normal text vector corresponding to the comprehensive data matrix of the news network data as the first text similarity;
[0059] S3-2-4, obtaining the similarity between the main features of the basic news network data and the historical abnormal text vector corresponding to the comprehensive data matrix of the news network data as the second text similarity;
[0060] S3-2-5, determining whether the first text similarity is greater than the second text similarity, if so, the text similarity screening state is normal, otherwise, the text similarity screening state is abnormal;
[0061] S3-2-6, judging whether the secondary features of the basic news network data are consistent with the historical normal associated labels corresponding to the news network data risk prediction processing results, if so, the associated feature data screening status is normal, otherwise, executing S3-2-7;
[0062] S3-2-7, determine whether the secondary features of the basic news network data are consistent with the historical abnormal association labels corresponding to the news network data risk prediction processing results. If so, the associated feature data screening status is abnormal. Otherwise, delete the current historical normal association labels and historical abnormal association labels, and return to S3-1-3-6;
[0063] S3-2-8. Determine whether the text benchmark word screening status, text similarity screening status and associated feature data screening status are all consistent. If so, output the text benchmark word screening status, text similarity screening status and associated feature data screening status as the news network data risk prediction processing result. Otherwise, return to S1.
[0064] Compared with the closest prior art, the present invention has the following beneficial effects:
[0065] For a large amount of disordered news network data, the first focus is on text content extraction, and based on this, text features and related classification and screening models are established. At the same time, it is noted that news network data are aggregated on network servers, and the source of their information data and the relevant information of the publishing users also need to be taken into consideration. From this, the peripheral features of the news network data other than text data are derived, and a multi-level analysis and screening model is also established to improve processing efficiency. In the process of model establishment, a large model commonly used in the industry is introduced to ensure the stability and reliability of the model establishment. That is, the focus of this solution is on the extraction, analysis, and processing of text features and the extraction, analysis, and processing of related features. Through the mutual correlation combined with historical manual review data, the final risk processing results are quickly output. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a flow chart of a news network data risk prediction processing method based on a large model provided by the present invention. DETAILED DESCRIPTION
[0067] The specific implementation modes of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0068] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0069] Embodiment 1:
[0070] The present invention provides a news network data risk prediction processing method based on a large model, such as Figure 1 As shown, including:
[0071] S1, using news network data to perform visualization processing to obtain text features of news network data;
[0072] S2, using the text features of the news network data to establish a data correlation analysis and screening model based on a large model;
[0073] S3. Obtaining a news network data risk prediction processing result according to the data correlation analysis screening model.
[0074] S1 specifically includes:
[0075] S1-1, obtaining text data of news network data and performing text noise reduction processing to obtain noise-reduced text data of news network data;
[0076] S1-2, performing word segmentation processing on the noise-reduced text data of the news network data to obtain a text word segmentation set of the news network data;
[0077] S1-3, performing part-of-speech tagging processing on the denoised text data of the news network data to obtain text part-of-speech tags for the news network data;
[0078] S1-4, using the denoised text data, text segmentation set and text part-of-speech tagging of the news network data as text features of the news network data;
[0079] The text noise reduction process is to remove meaningless content in the text data.
[0080] S2 specifically includes:
[0081] S2-1, using the text features of the news network data to establish a text data word vector classification model based on a large model;
[0082] S2-2, acquiring relevant feature data of the news network data according to the text features of the news network data;
[0083] S2-3, using the associated feature data of the news network data to establish an associated feature analysis model based on the large model;
[0084] S2-4. Utilize the text data word vector classification model and the association feature analysis model as a data association analysis screening model.
[0085] S2-1 specifically includes:
[0086] S2-1-1, dividing the noise-reduced text data into a first data set using the text features of the news network data;
[0087] S2-1-2, using the text features of the news network data to divide the corresponding text word set into a second data set;
[0088] S2-1-3, using the text features of the news network data to correspond to the text part-of-speech tags to divide it into a third data set;
[0089] S2-1-4, using the text features of the news network data to correspond to the text segmentation set at the relative position of the noise reduction text data to divide it into a fourth data set;
[0090] S2-1-5. Use the first data set, the second data set and the third data set as input and the fourth data set as output, and perform training based on the large model to obtain a text data word vector classification model.
[0091] S2-2 specifically includes:
[0092] S2-2-1, respectively obtaining the corresponding local port IP address and server port IP address according to the text features of the news network data;
[0093] S2-2-2, obtaining the corresponding news network website IP address according to the text features of the news network data;
[0094] S2-2-3, obtaining the classification label of the news network data corresponding to the news network data;
[0095] S2-2-4. Utilize the local port IP address, server port IP address, news network website IP address and classification label of news network data as associated characteristic data of news network data.
[0096] S2-3 specifically includes:
[0097] S2-3-1, using the local port IP address corresponding to the associated characteristic data of the news network data as the fifth data set;
[0098] S2-3-2, using the server port IP address corresponding to the associated characteristic data of the news network data as the sixth data set;
[0099] S2-3-3, using the news website IP address corresponding to the associated characteristic data of the news network data as the seventh data set;
[0100] S2-3-4, using the associated feature data of the news network data to correspond to the classification labels of the news network data as an eighth data set;
[0101] S2-3-5, using the fifth data set and the sixth data set as input, and the corresponding relationship between the fifth data set and the sixth data set as output, training is performed based on the large model to obtain a port IP address association model;
[0102] S2-3-6, using the sixth data set and the seventh data set as input, and the corresponding relationship between the sixth data set and the seventh data set as output, training is performed based on the large model to obtain a port website association model;
[0103] S2-3-7, using the seventh data set and the eighth data set as input, and the corresponding relationship between the seventh data set and the eighth data set as output, training is performed based on the large model to obtain a news website classification association model;
[0104] S2-3-8. Use the port IP address association model, the port website association model and the news website classification association model as an association feature analysis model.
[0105] S3 specifically includes:
[0106] S3-1, using the data correlation analysis and screening model to establish a comprehensive data matrix of news network data;
[0107] S3-2. Obtaining a news network data risk prediction processing result based on the comprehensive data matrix of the news network data.
[0108] S3-1 specifically includes:
[0109] S3-1-1, using the data association analysis screening model to correspond to the text data word vector classification model to obtain the model output result of the text data word vector classification model as the main feature of the basic news network data;
[0110] S3-1-2, using the data correlation analysis screening model to correspond to the correlation feature analysis model to obtain the model output result of the correlation feature analysis model as the secondary feature of the basic news network data;
[0111] S3-1-3. Establish a comprehensive data matrix of news network data using the main features of the basic news network data and the secondary features of the basic news network data.
[0112] S3-1-3 specifically includes:
[0113] S3-1-3-1, obtaining historical normal text features and historical abnormal text features respectively according to the text features of the news network data corresponding to the main features of the basic news network data;
[0114] S3-1-3-2, obtaining the central word of the historical normal text feature as the historical normal text reference word;
[0115] S3-1-3-3, obtaining the central word of the historical abnormal text feature as the historical abnormal text reference word;
[0116] S3-1-3-4, using the historical normal text feature corresponding to the text word set to establish a historical normal text vector according to the historical normal text benchmark word;
[0117] S3-1-3-5, using the historical abnormal text feature corresponding to the text word set to establish a historical abnormal text vector according to the historical abnormal text benchmark word;
[0118] S3-1-3-6, according to the historical normal text features, respectively obtain the corresponding historical normal basic news network data secondary features as historical normal association labels;
[0119] S3-1-3-7, according to the historical abnormal text features, respectively obtain the corresponding historical abnormal basic news network data secondary features as historical abnormality association labels;
[0120] S3-1-3-8. Use the historical normal text vectors, historical abnormal text vectors, historical normal associated labels and historical abnormal associated labels as a comprehensive data matrix of news network data.
[0121] S3-2 specifically includes:
[0122] S3-2-1, determine whether the text feature of the news network data has a historical normal text benchmark word corresponding to the comprehensive data matrix of the news network data, if so, the text benchmark word screening state is normal, and directly execute S3-2-3, otherwise, execute S3-2-2;
[0123] S3-2-2, determine whether the text features of the news network data have historical abnormal text benchmark words corresponding to the comprehensive data matrix of the news network data. If so, the text benchmark word screening status is abnormal, and execute S3-2-3; otherwise, delete the current historical normal text benchmark words and historical abnormal text benchmark words, and return to S3-1-3-2;
[0124] S3-2-3, obtaining the similarity between the main features of the basic news network data and the historical normal text vector corresponding to the comprehensive data matrix of the news network data as the first text similarity;
[0125] S3-2-4, obtaining the similarity between the main features of the basic news network data and the historical abnormal text vector corresponding to the comprehensive data matrix of the news network data as the second text similarity;
[0126] S3-2-5, determining whether the first text similarity is greater than the second text similarity, if so, the text similarity screening state is normal, otherwise, the text similarity screening state is abnormal;
[0127] S3-2-6, judging whether the secondary features of the basic news network data are consistent with the historical normal associated labels corresponding to the news network data risk prediction processing results, if so, the associated feature data screening status is normal, otherwise, executing S3-2-7;
[0128] S3-2-7, determine whether the secondary features of the basic news network data are consistent with the historical abnormal association labels corresponding to the news network data risk prediction processing results. If so, the associated feature data screening status is abnormal. Otherwise, delete the current historical normal association labels and historical abnormal association labels, and return to S3-1-3-6;
[0129] S3-2-8. Determine whether the text benchmark word screening status, text similarity screening status and associated feature data screening status are all consistent. If so, output the text benchmark word screening status, text similarity screening status and associated feature data screening status as the news network data risk prediction processing result. Otherwise, return to S1.
[0130] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0131] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0132] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A news network data risk prediction processing method based on a large model, characterized in that: include: S1, using news network data to perform visualization processing to obtain text features of news network data; S1-1, obtaining text data of news network data and performing text noise reduction processing to obtain noise-reduced text data of news network data; S1-2, performing word segmentation processing on the noise-reduced text data of the news network data to obtain a text word segmentation set of the news network data; S1-3, performing part-of-speech tagging processing on the denoised text data of the news network data to obtain text part-of-speech tags for the news network data; S1-4, using the denoised text data, text segmentation set and text part-of-speech tagging of the news network data as text features of the news network data; Wherein, the text denoising process is to remove meaningless content in the text data; S2, using the text features of the news network data to establish a data correlation analysis and screening model based on a large model; S2-1, using the text features of the news network data to establish a text data word vector classification model based on a large model; S2-1-1, dividing the noise-reduced text data into a first data set using the text features of the news network data; S2-1-2, using the text features of the news network data to divide the corresponding text word set into a second data set; S2-1-3, using the text features of the news network data to correspond to the text part-of-speech tags to divide it into a third data set; S2-1-4, using the text features of the news network data to correspond to the text segmentation set at the relative position of the noise reduction text data to divide it into a fourth data set; S2-1-5, using the first data set, the second data set and the third data set as input, and the fourth data set as output, training based on the large model to obtain a text data word vector classification model; S2-2, acquiring relevant feature data of the news network data according to the text features of the news network data; S2-2-1, respectively obtaining the corresponding local port IP address and server port IP address according to the text features of the news network data; S2-2-2, obtaining the corresponding news network website IP address according to the text features of the news network data; S2-2-3, obtaining the classification label of the news network data corresponding to the news network data; S2-2-4, using the local port IP address, server port IP address, news network website IP address and classification label of news network data as associated feature data of news network data; S2-3, using the associated feature data of the news network data to establish an associated feature analysis model based on the large model; S2-3-1, using the local port IP address corresponding to the associated characteristic data of the news network data as the fifth data set; S2-3-2, using the server port IP address corresponding to the associated characteristic data of the news network data as the sixth data set; S2-3-3, using the news website IP address corresponding to the associated characteristic data of the news network data as the seventh data set; S2-3-4, using the associated feature data of the news network data to correspond to the classification labels of the news network data as an eighth data set; S2-3-5, using the fifth data set and the sixth data set as input, and the corresponding relationship between the fifth data set and the sixth data set as output, training is performed based on the large model to obtain a port IP address association model; S2-3-6, using the sixth data set and the seventh data set as input, and the corresponding relationship between the sixth data set and the seventh data set as output, training is performed based on the large model to obtain a port website association model; S2-3-7, using the seventh data set and the eighth data set as input, and the corresponding relationship between the seventh data set and the eighth data set as output, training is performed based on the large model to obtain a news website classification association model; S2-3-8, using the port IP address association model, the port website association model and the news website classification association model as an association feature analysis model; S2-4, using the text data word vector classification model and the association feature analysis model as a data association analysis screening model; S3, obtaining a news network data risk prediction processing result according to the data correlation analysis screening model; S3-1, using the data correlation analysis and screening model to establish a comprehensive data matrix of news network data; S3-1-1, using the data association analysis screening model to correspond to the text data word vector classification model to obtain the model output result of the text data word vector classification model as the main feature of the basic news network data; S3-1-2, using the data correlation analysis screening model to correspond to the correlation feature analysis model to obtain the model output result of the correlation feature analysis model as the secondary feature of the basic news network data; S3-1-3, using the main features of the basic news network data and the secondary features of the basic news network data to establish a comprehensive data matrix of the news network data; S3-1-3-1, obtaining historical normal text features and historical abnormal text features respectively according to the text features of the news network data corresponding to the main features of the basic news network data; S3-1-3-2, obtaining the central word of the historical normal text feature as the historical normal text reference word; S3-1-3-3, obtaining the central word of the historical abnormal text feature as the historical abnormal text benchmark word; S3-1-3-4, using the historical normal text feature corresponding to the text word set to establish a historical normal text vector according to the historical normal text benchmark word; S3-1-3-5, using the historical abnormal text feature corresponding to the text word set to establish a historical abnormal text vector according to the historical abnormal text benchmark word; S3-1-3-6, according to the historical normal text features, respectively obtain the corresponding historical normal basic news network data secondary features as historical normal association labels; S3-1-3-7, according to the historical abnormal text features, respectively obtain the corresponding historical abnormal basic news network data secondary features as historical abnormality association labels; S3-1-3-8, using the historical normal text vectors, historical abnormal text vectors, historical normal associated labels and historical abnormal associated labels as a comprehensive data matrix of news network data; S3-2, obtaining a news network data risk prediction processing result according to the comprehensive data matrix of the news network data; S3-2-1, determine whether the text feature of the news network data has a historical normal text benchmark word corresponding to the comprehensive data matrix of the news network data, if so, the text benchmark word screening state is normal, and directly execute S3-2-3, otherwise, execute S3-2-2; S3-2-2, determine whether the text features of the news network data have historical abnormal text benchmark words corresponding to the comprehensive data matrix of the news network data. If so, the text benchmark word screening status is abnormal, and execute S3-2-3; otherwise, delete the current historical normal text benchmark words and historical abnormal text benchmark words, and return to S3-1-3-2; S3-2-3, obtaining the similarity between the main features of the basic news network data and the historical normal text vector corresponding to the comprehensive data matrix of the news network data as the first text similarity; S3-2-4, obtaining the similarity between the main features of the basic news network data and the historical abnormal text vector corresponding to the comprehensive data matrix of the news network data as the second text similarity; S3-2-5, determining whether the first text similarity is greater than the second text similarity, if so, the text similarity screening state is normal, otherwise, the text similarity screening state is abnormal; S3-2-6, judging whether the secondary features of the basic news network data are consistent with the historical normal associated labels corresponding to the news network data risk prediction processing results, if so, the associated feature data screening status is normal, otherwise, executing S3-2-7; S3-2-7, determine whether the secondary features of the basic news network data are consistent with the historical abnormal association labels corresponding to the news network data risk prediction processing results. If so, the associated feature data screening status is abnormal. Otherwise, delete the current historical normal association labels and historical abnormal association labels, and return to S3-1-3-6; S3-2-8. Determine whether the text benchmark word screening status, text similarity screening status and associated feature data screening status are all consistent. If so, output the text benchmark word screening status, text similarity screening status and associated feature data screening status as the news network data risk prediction processing result. Otherwise, return to S1.
Citation Information
Patent Citations
Risk prediction method and device, electronic equipment and storage medium
CN111753520A
Internet news analysis system and method based on big data
CN118093979A
News data processing method and device, computer equipment, storage medium and computer program product
CN118586391A