Big data information analysis method, device and system based on cloud computing, and storage medium
By conducting big data information analysis on network data on cloud computing platforms, and using cluster analysis and sentiment scoring technology to separate noise data, the problem of noise data affecting the accuracy of analysis in the existing technology is solved, and more efficient and reliable data analysis is achieved.
Patent Information
- Application Number
- CN202510490848.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing network data analysis methods are affected by a large number of noise data such as marketing information and rumors information synthesized by AI, resulting in a decrease in analysis accuracy and unreliable results.
The big data information analysis method based on cloud computing is adopted, and the evaluation articles of target keywords are obtained, and the text features are extracted for cluster analysis are obtained, templated features and content scores are obtained, and a two-dimensional evaluation point cloud map is constructed to separate the influence of batch generated content.
Effectively remove the impact of noise data, improve the accuracy and credibility of data analysis, and is especially suitable for scenarios where rapid response to massive text data, such as product word-of-mouth analysis.
Smart Images

Figure CN120011570A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network information data processing, and in particular, relates to a big data information analysis method, device, storage medium and system based on cloud computing. Background Art
[0002] At present, with the increasing popularity of Internet technology, the amount of data in the network is becoming increasingly large, and big data information analysis and processing methods have emerged.
[0003] Big data information analysis refers to the process of extracting valuable information, patterns or knowledge from massive, heterogeneous and dynamic Internet data and analyzing them. It combines big data technology, data mining algorithms, machine learning, natural language processing (NLP) and network analysis methods to solve challenges such as large data scale, complex structure and strong real-time performance.
[0004] At present, the analysis of network data is mainly based on machine learning, deep learning and natural language processing, and the output accuracy of these methods depends on the "purity" of the data input. However, there is a large amount of noise data in the network, such as marketing information and rumor information produced in batches by AI synthesis, which affects the accuracy of analysis of various data analysis systems and makes the results unreliable. Therefore, further improvements need to be made to the existing information analysis methods. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a big data information analysis method based on cloud computing, which aims to solve the problem that there is a large amount of noise data in the current network, such as marketing information and rumor information produced in batches by AI synthesis, which affects the analysis accuracy of various data analysis systems and causes unreliable analysis results.
[0006] The embodiment of the present application is implemented by providing a big data information analysis method based on cloud computing, the method comprising: Obtaining target keywords of the content to be analyzed, and obtaining a number of evaluation articles that evaluate the content of the target keywords based on network information; Acquire the writing features of each of the evaluation articles, where the writing features are used to characterize the structural logic features and / or content expression features of the evaluation articles; Performing cluster analysis on all the writing features to cluster the writing features with high similarity into one class, obtaining a typical templated feature representing the class from each class, and obtaining a number of templated features; performing similarity calculation between the writing features of each evaluation article and the templated features, and obtaining a templated score for each evaluation article; Conduct content analysis on the evaluation articles based on the sentiment analysis model to obtain content scores of each evaluation article on the target keyword; The templated score and content score are used as the horizontal and vertical coordinates respectively, and each evaluation article is used as a data point to construct a two-dimensional evaluation point cloud map.
[0007] Another object of the embodiment of the present application is to provide a big data information analysis device based on cloud computing, the big data information analysis device based on cloud computing comprising: An evaluation article acquisition module is used to acquire target keywords of the content to be analyzed, and acquire a number of evaluation articles that evaluate the content of the target keywords based on network information; A writing feature acquisition module, used to acquire writing features of each of the evaluation articles, wherein the writing features are used to characterize the structural logic features and / or content expression features of the evaluation articles; The template score acquisition module is used to perform cluster analysis on all the writing features to cluster the writing features with high similarity into one class, obtain a typical template feature representing the class from each class, and obtain a number of template features; perform similarity calculation between the writing features of each evaluation article and the template feature to obtain a template score for each evaluation article; A content score acquisition module is used to perform content analysis on the evaluation articles based on the sentiment analysis model to obtain a content score of each evaluation article for evaluating the target keyword; The two-dimensional evaluation point cloud map acquisition module is used to construct a two-dimensional evaluation point cloud map using the templated score and content score as the horizontal axis and the vertical axis respectively, and each evaluation article as a data point.
[0008] Another object of an embodiment of the present application is to provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor executes the steps of the cloud computing-based big data information analysis method as described above.
[0009] Another object of an embodiment of the present application is to provide a big data information analysis system based on cloud computing, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the big data information analysis method based on cloud computing as described above.
[0010] The embodiment of the present application provides a big data information analysis method based on cloud computing. Its outstanding advantage is that the present application utilizes the high computing power of the cloud platform to deeply integrate big data analysis technology with information extraction technology, and effectively separates the evaluation of the target to be evaluated from the batch-generated content and the non-patterned generated content in the form of a two-dimensional image, thereby obtaining more real and objective evaluation data and removing the influence of noise. It is particularly suitable for scenarios such as product word-of-mouth analysis that require rapid response to massive text data, and is accurate and efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 An application environment diagram of a big data information analysis method based on cloud computing provided in an embodiment of the present application; Figure 2 A flowchart of a big data information analysis method based on cloud computing provided in an embodiment of the present application; Figure 3 A schematic diagram of a two-dimensional evaluation point cloud map provided in an embodiment of the present application; Figure 4 A schematic diagram of an input box provided in an embodiment of the present application; Figure 5 A schematic diagram of a three-dimensional evaluation point cloud map provided in an embodiment of the present application; Figure 6 A structural block diagram of a big data information analysis device based on cloud computing provided in an embodiment of the present application; Figure 7 FIG. 4 is a block diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0012] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0013] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first unit or module from another unit or module. For example, a first script may be referred to as a second script, and similarly, a second script may be referred to as a first script without departing from the scope of this application.
[0014] Figure 1 The application environment diagram of the big data information analysis method based on cloud computing provided in the embodiment of the present application is as follows: Figure 1 As shown, in the application environment, a terminal 110 and a computer device 120 are included.
[0015] The computer device 120 may be an independent physical server or terminal, or a server cluster composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, cloud storage, and CDN.
[0016] The terminal 110 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, etc., but is not limited thereto. The terminal 110 and the computer device 120 may be connected via a network, and this application does not limit this.
[0017] like Figure 2 As shown, in one embodiment, a big data information analysis method based on cloud computing is proposed. This embodiment mainly applies this method to the above Figure 1 A method for analyzing big data information based on cloud computing may specifically include the following steps: Step S10, obtaining a target keyword of the content to be analyzed, and obtaining a number of evaluation articles that evaluate the content of the target keyword based on network information.
[0018] In this embodiment, based on big data network retrieval, crawlers can be used to confirm and obtain articles that are relevant to the target keyword. For example, the content to be analyzed is a film or TV work, and the target keyword can be the full name or alias of the work. The system automatically captures content articles related to the target keyword on the network and scores or evaluates it.
[0019] Step S20, obtaining the writing features of each of the evaluation articles, wherein the writing features are used to characterize the structural logic features and / or content expression features of the evaluation articles.
[0020] In this embodiment, because the content produced by the batch generation engine is limited by factors such as cost, it often has a patterned formula or template. The number of these non-real evaluation articles or information released may be hundreds or thousands or more, and their content is highly homogenized. Therefore, a natural language processing model can be used to perform feature analysis on the evaluation articles, quantify and obtain the writing features of these contents in terms of overall structural logic and / or content expression, so as to facilitate analysis and processing in subsequent steps.
[0021] Step S30, clustering analysis is performed on all the writing features to cluster the writing features with high similarity into one class, and a typical templated feature representing the class is obtained from each class to obtain a plurality of templated features; similarity is calculated between the writing features of each evaluation article and the templated features to obtain a templated score for each evaluation article.
[0022] This step is used to screen and extract articles that are batch-produced by similar templates or generated by generative models from a large number of evaluation articles. A typical text feature that can characterize the characteristics of each class can be obtained from each class, and the feature can be set as the template feature of this class. The template feature can be quantified in the form of a digital matrix, vector, etc. Among them, the template feature is the center of the cluster, and multiple template features are obtained for multiple classes.
[0023] This step also calculates the similarity between each text feature and the template feature. If the text feature has a high similarity with any template feature, the text feature is assigned a high template score value. The template score is used to measure the possibility that the content is batch generated.
[0024] As a preferred embodiment of the present application, the traditional K-means clustering method generally presets or predicts the approximate number of clusters, but the effect of this method applied to the present application is not good. Therefore, this embodiment adopts a hierarchical clustering method or a DBSCAN clustering method, which is more suitable for situations where there is no preset number of categories, suitable for processing data with uneven density and more noise, and can automatically determine the number of clusters. Through clustering, similar text structures, logical relationships, language styles, etc. are identified, and then several text templates are obtained.
[0025] Step S40, performing content analysis on the evaluation articles based on the sentiment analysis model to obtain a content score of each evaluation article for evaluating the target keyword.
[0026] In this embodiment, SnowNLP, BERT or other mainstream sentiment analysis models can be called to calculate the sentiment evaluation score of each evaluation article on the target content. For example, if the evaluation article likes a product very much, the content score is set to 1. If the evaluation article dislikes a product very much, the score is set to -1. The content score is limited to the value range [−1, 1] in a uniform manner to facilitate unified display.
[0027] Step S50, using the templated score and content score as the horizontal axis and the vertical axis respectively, and each evaluation article as a data point, to construct a two-dimensional evaluation point cloud map.
[0028] In this embodiment, the entire solution can be set up in the cloud, using cloud computing to store and host interactive charts, and using the visualization library to dynamically render Figure 3 In the large-scale point cloud shown in the figure, the text templates formed by different cluster centers can be represented by different identifiers.
[0029] The outstanding advantage of this application is that it can accurately eliminate a large amount of content that is generated in a pattern and is highly repetitive, highly patchwork, and empty from the evaluation system, and use the point cloud map formed by big data to intuitively display the evaluation articles with different levels of originality scores for the target keywords. The evaluation is clear and intuitive, which significantly improves the authenticity and objectivity of the evaluation system and facilitates users to obtain the public's true evaluation of the content to be analyzed.
[0030] In a preferred embodiment, the method for acquiring the structural logical features is: Based on the grammatical attribute types of the text, construct the functional attribute labels of several text segments: ; Based on the functional attribute tags, the evaluation article is segmented to obtain several content segments: ; Based on the functional attribute label corresponding to each content segment, a structural logic feature vector is obtained: , , in, Indicates Function attribute tags, Indicates content segments, Indicates the evaluation article. represents the structural logical feature vector, Indicates The functional attribute label corresponding to each content segment.
[0031] In the embodiments of the present application, there may be multiple ways to analyze the characteristics of the text. The present application weighs and selects the functional attribute labels of the content of the article as the core metric, and a sequence annotation model can be used to identify the functional attribute labels of the text segment. The functional attribute label refers to the grammatical function of the text segment, for example, "introduction", "argument", "conclusion", "quote", "example", etc. It can be understood that a natural paragraph may be divided into multiple content segments, and multiple natural paragraphs may belong to the same content segment. After segmentation, each content segment corresponds to its own structural label.
[0032] In a preferred embodiment, the content representation feature at least includes an information density feature, and the method for obtaining the information density feature is: Get the content segment, divide the content segment into several sentences, and obtain the semantic vector of each sentence: Get the mean vector of the semantic vectors of all sentences in the content segment , and then get the square of the Euclidean distance between the semantic vector of each sentence and the mean vector, and get the average value SemVar(P) of all sentences; Based on the average value SemVar(P), obtain the information density value Den(P) of the content segment; in, Indicates Sentences The semantic vector of represents the semantic vector generation function, Indicates the content segment to be segmented. Indicates the number of sentences formed by content segment segmentation, Indicates that there is dimensional vector space.
[0033] In the embodiments of the present application, the content expression features have various manifestations, such as semantic features, language style features, etc. In this embodiment, the researchers took into account that the content produced in batches is often relatively vague, there is a large amount of homogeneous content in the context, and the content information density is low. Therefore, while judging whether the overall structure is patterned, the content density judgment is further introduced to analyze the content of the segmented text.
[0034] First, based on the semantic vectorization model, a semantic vector is obtained. The semantic vectorization model can use a pre-trained Transformer encoding model, which will not be described here. If the information transmission of sentences in the content segment is compact, the information density is high; if the sentence has a high amount of redundant description and a large number of similar discussions, the density is low.
[0035] In this embodiment, each sentence is mapped into a high-dimensional space For a point in the content segment, sentences with similar semantics are closer in space and have higher cosine similarity. The mean vector of the semantic vectors of all sentences in the content segment satisfy: , is the semantic center vector of the text segment. Get the mean vector of the semantic vectors of all sentences in the content segment The method is: ; The calculation method of the average value SemVar(P) of the squared Euclidean distance between the semantic vector of each sentence and the mean vector is: Based on the average value SemVar(P), the information density value of the entire content segment can be obtained. The information density value can be mapped to the [0,1] interval through normalization. The larger the value, the higher the information density. The calculation method of the Den(P) value of the content segment is: Through the above method, we can judge whether the article is templated in terms of overall structure and content. It is understandable that the information density is similar when the language expression is replaced in different articles. The content expression feature can contain more dimensions, and only the judgment method based on the information density dimension is provided here.
[0036] In a preferred embodiment, when the writing feature represents the structural logic feature and the content expression feature of the evaluation article, the method for obtaining the writing feature is: Among them, F represents the writing characteristics.
[0037] In the embodiments of the present application, Indicates The label corresponding to the content segment, Indicates The information density value corresponding to each content paragraph. The content generated in batches by similar models and templates often has similar argumentation methods. Therefore, the article can be quantitatively evaluated from both the content and structure levels. There are m elements in the text feature F, each of which contains the label of the paragraph and the information density of the paragraph. When there are more dimensions in the content expression feature, each element can be expanded. For example, it also contains semantic features and language style features. At this time, The elements are: .
[0038] In a preferred embodiment, a one-hot encoding vector can be generated for each content segment (such as "introduction", "argument", etc.) according to the functional label ,For example ∈ , where K is the total number of label categories. For example, if the label set is {introduction, argument, conclusion}, then the vector corresponding to the "argument" segment is [0,1,0]. This facilitates feature matching and clustering. Preferably, the position information of the content segment in the article is also captured during clustering, such as the normalized position of the paragraph order, to synthesize the final structural logic feature vector and improve matching accuracy.
[0039] In a preferred embodiment, the method of performing cluster analysis on all the text features to cluster text features with high similarity into one class includes: Obtain the preset clustering threshold and the similarity value of the similarity matrix of each class; If and only if the similarity value is higher than the clustering threshold, the clustering result is output and a typical template feature characterizing the class is obtained.
[0040] In the embodiment of the present application, a limit value needs to be set for generating clustering results, that is, only text features with similarity higher than a threshold are clustered into one class to avoid features with low similarity being mistakenly clustered into one class, thereby reducing the probability of misjudgment.
[0041] In a preferred embodiment, the method for obtaining the templated feature further includes: Get the input false information source address; Based on the false information source address, obtain the false information source content; The text feature of the content of the false information source is obtained, and the feature is set as a templated feature.
[0042] In the embodiments of the present application, Figure 4 As shown, users can also manually input the web address of the templated article into the system based on their own search or experience, and enter the web link of the false information source address into the text box. The number of link content entered can be increased or decreased by clicking the "+" or "-" sign.
[0043] In a preferred embodiment, the method further comprises: Obtaining the publishing timestamp information of the evaluation article; A spatial rectangular evaluation coordinate system is constructed, and the templated score is used as the x-axis coordinate, the content score is used as the z-axis coordinate, the publishing timestamp information is used as the y-axis coordinate, and each evaluation article is used as a data point to obtain a three-dimensional evaluation point cloud map.
[0044] In the embodiments of the present application, Figure 5 As shown in Figure 1, the time dimension is further introduced, so that users can clearly know the changes in the reputation of the target to be evaluated over time. The information displayed by the three-dimensional point cloud map is more comprehensive, objective and specific.
[0045] like Figure 6 As shown, in one embodiment, a big data information analysis device based on cloud computing is provided. The big data information analysis device based on cloud computing can be integrated into the above-mentioned computer device 120, and specifically may include: An evaluation article acquisition module 510 is used to acquire a target keyword of the content to be analyzed, and acquire a number of evaluation articles that evaluate the content of the target keyword based on network information; A text feature acquisition module 520 is used to acquire text features of each of the evaluation articles, wherein the text features are used to characterize the structural logic features and / or content expression features of the evaluation article; The template score acquisition module 530 is used to perform cluster analysis on all the text features to cluster text features with high similarity into one class, obtain a typical template feature representing the class from each class, and obtain a plurality of template features; perform similarity calculation between the text features of each evaluation article and the template feature to obtain a template score for each evaluation article; A content score acquisition module 540 is used to perform content analysis on the evaluation articles based on the sentiment analysis model to obtain a content score of each evaluation article for evaluating the target keyword; The two-dimensional evaluation point cloud map acquisition module 550 is used to construct a two-dimensional evaluation point cloud map using the templated score and content score as the horizontal axis and the vertical axis respectively and each evaluation article as a data point.
[0046] In the embodiments of the present application, for the explanation and description of the above-mentioned big data information analysis device based on cloud computing, please refer to the explanation and description of the above-mentioned corresponding method. For the description of the above-mentioned big data information analysis method based on cloud computing, please refer to the above text and will not be repeated here.
[0047] The outstanding advantage of the embodiment is that it utilizes the high computing power of the cloud platform to deeply integrate big data analysis technology and information extraction technology, and effectively separates the evaluation of the target to be evaluated from the batch-generated content and the non-patterned generated content in the form of a two-dimensional image, thereby obtaining more real and objective evaluation data and removing the impact of noise. It is especially suitable for scenarios such as product word-of-mouth analysis that require rapid response to massive text data, and is accurate and efficient.
[0048] Figure 7 The internal structure diagram of a computer device in one embodiment is shown. The computer device may specifically be Figure 1 The computer device 120 in FIG. Figure 7 As shown, the computer device includes a processor, a memory, a network interface, an input device and a display screen connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement a big data information analysis method based on cloud computing. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute a big data information analysis method based on cloud computing. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0049] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0050] In one embodiment, the cloud computing-based big data information analysis device provided by the present application can be implemented in the form of a computer program, which can be used in Figure 7 The computer device shown in the figure is run on the computer device. The memory of the computer device can store various program modules that constitute the cloud computing-based big data information analysis device, such as: Figure 6 The evaluation article acquisition module 510, the text feature acquisition module 520 and the template score acquisition module 530 are shown. The computer program composed of various program modules enables the processor to execute the steps of the cloud computing-based big data information analysis method of each embodiment of the present application described in this specification.
[0051] For example, Figure 7 The computer device shown can be Figure 6 The evaluation article acquisition module 510 in the cloud computing-based big data information analysis device shown executes step S10. The computer device can execute step S20 through the text feature acquisition module 520. And so on.
[0052] In one embodiment, a big data information analysis system based on cloud computing is proposed, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the big data information analysis method based on cloud computing as described above.
[0053] In the embodiments of the present application, please refer to the above for the description of the above-mentioned big data information analysis method based on cloud computing, which will not be repeated here.
[0054] The outstanding advantage of the embodiment of the present application is that it utilizes the high computing power of the cloud platform to deeply integrate big data analysis technology and information extraction technology, and effectively separates the evaluation of the target to be evaluated by batch-generated content and non-patterned generated content in the form of two-dimensional images, thereby obtaining more real and objective evaluation data and removing the impact of noise. It is especially suitable for scenarios such as product word-of-mouth analysis that require rapid response to massive text data, and is accurate and efficient.
[0055] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the steps of the cloud computing-based big data information analysis method as described above.
[0056] In the embodiments of the present application, please refer to the above for the description of the above-mentioned big data information analysis method based on cloud computing, which will not be repeated here.
[0057] In the embodiment of the present application, the program run by the method stored in the storage medium of the embodiment of the present application has the outstanding advantage in that it utilizes the high computing power of the cloud platform to deeply integrate big data analysis technology and information extraction technology, and effectively separates the evaluation of the target to be evaluated from the batch-generated content and the non-patterned generated content in the form of a two-dimensional image, thereby obtaining more real and objective evaluation data and removing the impact of noise. It is particularly suitable for scenarios such as product word-of-mouth analysis that require rapid response to massive text data, and is accurate and efficient.
[0058] It should be understood that, although each step in the flow chart of each embodiment of the present application is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is clear explanation in this article, the execution of these steps does not have strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0059] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0060] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0061] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A big data information analysis method based on cloud computing, characterized in that: The method comprises: Obtaining target keywords of the content to be analyzed, and obtaining a number of evaluation articles that evaluate the content of the target keywords based on network information; Acquire the writing features of each of the evaluation articles, where the writing features are used to characterize the structural logic features and / or content expression features of the evaluation articles; Performing cluster analysis on all the writing features to cluster the writing features with high similarity into one class, obtaining a typical templated feature representing the class from each class, and obtaining a number of templated features; performing similarity calculation between the writing features of each evaluation article and the templated features, and obtaining a templated score for each evaluation article; Conduct content analysis on the evaluation articles based on the sentiment analysis model to obtain content scores of each evaluation article on the target keyword; The templated score and content score are used as the horizontal and vertical coordinates respectively, and each evaluation article is used as a data point to construct a two-dimensional evaluation point cloud map.
2. The big data information analysis method based on cloud computing according to claim 1 is characterized in that: The method for obtaining the structural logical features is: Based on the grammatical attribute types of the text, construct the functional attribute labels of several text segments: ; Based on the functional attribute tags, the evaluation article is segmented to obtain several content segments: ; Based on the functional attribute label corresponding to each content segment, a structural logic feature vector is obtained: , , in, represents the nth functional attribute label, represents the mth content segment, Indicates the evaluation article. represents the structural logical feature vector, Indicates the functional attribute label corresponding to the mth content segment, Represents a feature attribute label.
3. The big data information analysis method based on cloud computing according to claim 2 is characterized in that: The content description feature at least includes an information density feature, and the method for obtaining the information density feature is: Get the content segment, divide the content segment into several sentences, and obtain the semantic vector of each sentence: Get the mean vector of the semantic vectors of all sentences in the content segment , and then get the square of the Euclidean distance between the semantic vector of each sentence and the mean vector, and get the average value SemVar(P) of all sentences; Based on the average value SemVar(P), obtain the information density value Den(P) of the content segment; in, Represents the i-th sentence The semantic vector of represents the semantic vector generation function, Indicates the content segment to be segmented. Indicates the number of sentences formed by content segment segmentation, Indicates that there is dimensional vector space.
4. The big data information analysis method based on cloud computing according to claim 3 is characterized in that: When the writing feature represents the structural logic feature and content expression feature of the evaluation article, the method for obtaining the writing feature is: in, Indicates the characteristics of the text.
5. The big data information analysis method based on cloud computing according to claim 1 is characterized in that: The method of clustering all the text features to cluster the text features with high similarity into one class includes: Obtain the preset clustering threshold and the similarity value of the similarity matrix of each class; If and only if the similarity value is higher than the clustering threshold, the clustering result is output and a typical template feature characterizing the class is obtained.
6. The big data information analysis method based on cloud computing according to claim 1 is characterized in that: The method for obtaining the templated feature also includes: Get the input false information source address; Based on the false information source address, obtain the false information source content; The text feature of the content of the false information source is obtained, and the feature is set as a templated feature.
7. The big data information analysis method based on cloud computing according to claim 1 is characterized in that: The method further comprises: Obtaining the publishing timestamp information of the evaluation article; A spatial rectangular evaluation coordinate system is constructed, and the templated score is used as the x-axis coordinate, the content score is used as the z-axis coordinate, the publishing timestamp information is used as the y-axis coordinate, and each evaluation article is used as a data point to obtain a three-dimensional evaluation point cloud map.
8. A big data information analysis device based on cloud computing, characterized in that: The big data information analysis device based on cloud computing includes: An evaluation article acquisition module is used to acquire target keywords of the content to be analyzed, and acquire a number of evaluation articles that evaluate the content of the target keywords based on network information; A writing feature acquisition module, used to acquire writing features of each of the evaluation articles, wherein the writing features are used to characterize the structural logic features and / or content expression features of the evaluation articles; The template score acquisition module is used to perform cluster analysis on all the writing features to cluster the writing features with high similarity into one class, obtain a typical template feature representing the class from each class, and obtain a number of template features; perform similarity calculation between the writing features of each evaluation article and the template feature to obtain a template score for each evaluation article; A content score acquisition module is used to perform content analysis on the evaluation articles based on the sentiment analysis model to obtain a content score of each evaluation article for evaluating the target keyword; The two-dimensional evaluation point cloud map acquisition module is used to construct a two-dimensional evaluation point cloud map using the templated score and content score as the horizontal axis and the vertical axis respectively, and each evaluation article as a data point.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the cloud computing-based big data information analysis method described in any one of claims 1 to 7.
10. A big data information analysis system based on cloud computing, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the cloud computing-based big data information analysis method as described in any one of claims 1 to 7.