Data analysis method and system, electronic equipment and storage medium
By determining the target proxy address from the preset proxy pool and performing word segmentation and word meaning network analysis on the data, the problems of low cloud computing data acquisition efficiency and poor analysis results are solved, and more efficient and stable data analysis is achieved.
Patent Information
- Application Number
- CN202510096243.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
Cloud computing data analysis methods are prone to the problem of IP addresses being blocked when collecting data. Data collection efficiency is low and the stability is insufficient, making it difficult to conduct in-depth mining analysis, and the analysis effect is poor.
By determining the target proxy address from the preset proxy pool, obtaining the data to be analyzed from the target cloud server cluster, and performing first data preprocessing on the data, including word segmentation processing and word meaning network analysis processing, and finally performing predictive analysis to improve the accuracy and reliability of the data analysis.
It effectively improves data acquisition efficiency and stability, enhances the depth and breadth of data analysis, and thus improves the accuracy and reliability of data analysis.
Smart Images

Figure CN119988908A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data service technology, and in particular to a data analysis method, system, electronic device and storage medium. Background Art
[0002] Cloud servers use cloud computing technology to merge the resources of multiple servers with no upper limit to form a cluster, and then distribute the resources in the cluster to different users. In related technologies, cloud computing data analysis methods are prone to IP address blocking during data collection, low data collection efficiency, and insufficient stability. At the same time, it is difficult to conduct in-depth mining and analysis of relevant data during data analysis, and the analysis effect is poor.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to propose a data analysis method, system, electronic device and storage medium, which can effectively improve the efficiency and stability of data collection, and improve the depth and breadth of data analysis, thereby improving the accuracy and reliability of data analysis.
[0005] To achieve the above object, an embodiment of the present application provides a data analysis method, which includes the following steps:
[0006] Determine a target proxy address from a preset proxy pool, so as to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address;
[0007] Performing a first data preprocessing on the data to be analyzed to obtain first preprocessed data; wherein the first data preprocessing includes word segmentation processing and word meaning network analysis processing;
[0008] Performing predictive analysis on the first preprocessed data to obtain data analysis results.
[0009] In some embodiments, determining a target proxy address from a preset proxy pool to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address includes:
[0010] Obtaining a preset proxy address to verify the validity of the preset proxy address and obtain a first verification result;
[0011] When it is determined that the first verification result corresponding to the preset proxy address is verification passed, performing a risk assessment on the preset proxy address;
[0012] When it is determined that the risk level of the preset proxy address is lower than the preset risk threshold, the preset proxy address is stored in a preset database to construct the preset proxy pool;
[0013] Allocating the target proxy address from the preset proxy pool according to a preset allocation strategy;
[0014] Data is collected through the target proxy address to obtain the data to be analyzed.
[0015] In some embodiments, after performing the step of determining that the risk level of the preset proxy address is lower than a preset risk threshold, storing the preset proxy address in a preset database to construct the preset proxy pool, the method further includes:
[0016] When it is determined that the update period is satisfied, performing validity verification on each of the preset proxy addresses in the preset proxy pool to obtain a second verification result;
[0017] The preset proxy pool is updated according to the second verification result.
[0018] In some embodiments, the first pre-processed data includes a word cloud image and a word meaning network image;
[0019] The step of performing first data preprocessing on the data to be analyzed to obtain first preprocessed data includes:
[0020] Performing word segmentation processing on the data to be analyzed by using a preset word segmentation function to obtain a word segmentation result list; wherein the word segmentation result list includes a plurality of preset word segmentations;
[0021] Traversing the word segmentation result list to perform a preset condition check on each of the preset word segments to obtain a condition check result; wherein the preset condition check includes determining whether the preset word segment is not in a preset stop word character table, and determining whether the word length of the preset word segment is greater than a preset length;
[0022] When it is determined that the conditional test result is that the preset word segmentation is not in the preset stop word character table, and the word length of the preset word segmentation is greater than the preset length, the preset word segmentation is added to the test result list;
[0023] Generate a word cloud image by using a preset word cloud function according to the inspection result list;
[0024] Performing text filtering on the word segmentation result list according to the inspection result list to obtain target text content;
[0025] Performing word frequency calculation on the target text content to determine the target keyword through the word frequency data obtained by calculation;
[0026] When it is determined that the number of the target keywords in the first text in the target text content is greater than a preset threshold, adjusting the association matrix corresponding to the first text according to a preset change value;
[0027] A word meaning network image is constructed based on the association matrix.
[0028] In some embodiments, before performing the first data preprocessing on the data to be analyzed to obtain the first preprocessed data, the method further includes:
[0029] A second data preprocessing is performed on the data to be analyzed to obtain second preprocessed data; wherein the second data preprocessing includes data deduplication processing and outlier processing.
[0030] In some embodiments, performing predictive analysis on the first preprocessed data to obtain data analysis results includes:
[0031] Aggregating the first pre-processed data corresponding to each of the target proxy addresses to obtain preset aggregated data;
[0032] The preset aggregated data is input into a preset prediction model for prediction analysis to obtain the data analysis result; wherein the weight data of the preset prediction model is determined by an optimal weighting algorithm.
[0033] In some embodiments, after inputting the preset aggregated data into a preset prediction model for prediction analysis to obtain the data analysis result, the method further includes:
[0034] The preset aggregated data and the data analysis results are visualized to obtain visualized data information.
[0035] To achieve the above object, another aspect of the embodiment of the present application provides a data analysis system, the system comprising:
[0036] The first module is used to determine a target proxy address from a preset proxy pool, so as to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address;
[0037] The second module is used to perform a first data preprocessing on the data to be analyzed to obtain first preprocessed data; wherein the first data preprocessing includes word segmentation processing and word meaning network analysis processing;
[0038] The third module is used to perform predictive analysis on the first preprocessed data to obtain data analysis results.
[0039] To achieve the above object, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising:
[0040] at least one processor;
[0041] at least one memory for storing at least one program;
[0042] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0043] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0044] The embodiments of the present application include at least the following beneficial effects: The present application provides a data analysis method, system, electronic device and storage medium, which obtains the data to be analyzed from the target cloud server cluster according to the target proxy address by determining the target proxy address from the preset proxy pool. Next, the embodiment of the present invention performs a first data preprocessing on the data to be analyzed to obtain the first preprocessed data, including word segmentation processing and word meaning network analysis processing. Then, the embodiment of the present invention performs a predictive analysis on the first preprocessed data to obtain a data analysis result, and realizes a reliable and accurate analysis of the data. It is easy to understand that the embodiment of the present invention can effectively improve the efficiency and stability of data collection by determining the target proxy address from the preset proxy pool, and at the same time, by performing word segmentation processing and word meaning network analysis processing on the data to be analyzed, it can perform in-depth analysis and mining on the data to be analyzed, effectively improve the integrity and sufficiency of data analysis, and thus effectively improve the accuracy and reliability of data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of the steps of the data analysis method provided by an embodiment of the present invention;
[0046] Figure 2 It is a flowchart of the steps of determining a target proxy address from a preset proxy pool to obtain data to be analyzed from a target cloud server cluster according to the target proxy address provided by an embodiment of the present invention;
[0047] Figure 3 is a flowchart of the steps of updating a preset proxy pool provided by an embodiment of the present invention;
[0048] Figure 4 It is a schematic diagram of the architecture of the IP proxy pool provided by an embodiment of the present invention;
[0049] Figure 5 is a flowchart of steps for performing text analysis on data to be analyzed to obtain text analysis data provided by an embodiment of the present invention;
[0050] Figure 6 is a flowchart of the steps of word frequency analysis provided by an embodiment of the present invention;
[0051] Figure 7 is a flowchart of the steps of word meaning network analysis provided by an embodiment of the present invention;
[0052] Figure 8 is a flowchart of the steps of performing data preprocessing on the data to be analyzed provided by an embodiment of the present invention;
[0053] Fig. 9 It is a flowchart of the steps of performing predictive analysis on text analysis data to obtain data analysis results provided by an embodiment of the present invention;
[0054] Fig.10 is a flowchart of the steps of data visualization provided by an embodiment of the present invention;
[0055] Fig.11 is an overall flow chart of data analysis provided by an embodiment of the present invention;
[0056] Fig.12 is a schematic diagram of the structure of a data analysis system provided by an embodiment of the present invention;
[0057] Fig.13 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0059] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0060] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0062] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0063] Cloud computing technology: a new way of computing based on the Internet, providing computing services through the network, including servers, storage, databases, networks, software, analysis and intelligence resources.
[0064] Data analysis: refers to the use of appropriate statistical analysis methods to analyze large amounts of collected data, summarize, understand and digest them in order to maximize the development of data functions and play the role of data.
[0065] Communication data: refers to connecting data terminals with computers through transmission channels, enabling data terminals in different locations to share software, hardware and information resources, and complete user information resource sharing.
[0066] Cloud servers use cloud computing technology to merge the resources of multiple servers with no upper limit to form a cluster, and then distribute the resources in the cluster to different users. In the related art, the analysis method of cloud computing data is prone to the problem of IP address being blocked when collecting data, and the data collection efficiency is low and the stability is insufficient. For example, when using the IP address of a fixed server to collect data, the IP address is prone to being blocked. A single IP address is used to request data, and the data collection efficiency is low and the stability is insufficient. At the same time, when performing data analysis, it is difficult to conduct in-depth mining and analysis of related data, and the analysis effect is not good. For example, although the collected data is processed and integrated through conventional data processing, there is a lack of deep semantic mining and understanding of related text data (such as policy texts, etc.). At the same time, related text analysis tools often require professional technicians to operate, which is not user-friendly, limiting their application in the actual decision-making process. In addition, in terms of the effect of data analysis, it is difficult to respond quickly and provide effective decision-making effect prediction, which to a certain extent limits the application scope and effect of the system.
[0067] In view of this, a data analysis method, system, electronic device and storage medium are provided in an embodiment of the present application. The scheme determines the target proxy address from a preset proxy pool to obtain the data to be analyzed from the target cloud server cluster according to the target proxy address. Next, the embodiment of the present invention performs a first data preprocessing on the data to be analyzed to obtain first preprocessed data, including word segmentation processing and word meaning network analysis processing. Then, the embodiment of the present invention performs predictive analysis on the first preprocessed data to obtain data analysis results, realizes reliable and accurate analysis of the data, can effectively improve the efficiency and stability of data collection, and improves the depth and breadth of data analysis, thereby improving the accuracy and reliability of data analysis.
[0068] The data analysis method provided in the embodiment of the present application relates to the field of data service technology. The data analysis method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or it can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and cloud servers for basic cloud computing services such as big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the data analysis method, etc., but is not limited to the above forms.
[0069] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0070] Figure 1 is an optional flow chart of the data analysis method provided in the embodiment of the present application. Figure 1The method may include but is not limited to steps S110 to S130.
[0071] Step S110: determining a target proxy address from a preset proxy pool, so as to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address.
[0072] Step S120: Performing a first data preprocessing on the data to be analyzed to obtain first preprocessed data, wherein the first data preprocessing includes word segmentation processing and word meaning network analysis processing.
[0073] Step S130: performing predictive analysis on the first preprocessed data to obtain data analysis results.
[0074] During the working process of this specific embodiment, the embodiment of the present invention first determines the target proxy address from the preset proxy pool to obtain the data to be analyzed from the target cloud server according to the target proxy address. Specifically, the preset proxy pool in the embodiment of the present invention refers to an address set storing several proxy addresses. For example, the preset proxy pool in the embodiment of the present invention includes several IP addresses. Accordingly, the embodiment of the present invention determines the corresponding IP address from the preset proxy pool according to the preset allocation strategy, and uses the IP address as the target proxy address to obtain the data to be analyzed from the target cloud server cluster through the IP address (target proxy address). Among them, the data to be analyzed in the embodiment of the present invention refers to the data that needs to be analyzed, including the basic information of the user, business data and real-time policy information. Then, the embodiment of the present invention performs a first data preprocessing on the data to be analyzed to obtain the first preprocessed data. Specifically, the first data preprocessing in the embodiment of the present invention includes word segmentation processing and word meaning network analysis processing. Among them, the embodiment of the present invention performs text analysis preprocessing on the acquired data information (data to be analyzed) by combining word segmentation processing and word meaning network analysis processing, so as to conduct in-depth mining and understanding of the data to be analyzed, obtain the first preprocessed data, and facilitate subsequent data analysis and prediction. Finally, the embodiment of the present invention obtains data analysis results by performing predictive analysis on the first preprocessed data. Specifically, the embodiment of the present invention inputs the first preprocessed data after deep mining and analysis into the pre-built prediction model to predict the trend of the data and obtain data analysis results. It should be noted that the embodiment of the present invention can dynamically change the proxy address by determining the target proxy address from the preset proxy pool, alleviate the access pressure of a single proxy IP, and improve the stability of IP address access and data collection efficiency. At the same time, it can alleviate the data acquisition abnormality problem caused by the failure of a single proxy IP to be used normally, and improve the reliability and stability of data collection. In addition, the embodiment of the present invention can deeply mine the semantic information of the collected data text by performing word segmentation processing and word meaning network analysis processing on the collected data to be analyzed, and then perform predictive analysis on the preprocessed data, which can effectively improve the depth and breadth of analysis of the data to be analyzed, thereby improving the accuracy and reliability of data analysis.
[0075] Reference Figure 2 In some embodiments of the present invention, determining a target proxy address from a preset proxy pool to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address includes but is not limited to the following steps:
[0076] Step S210: obtaining a preset proxy address to verify the validity of the preset proxy address and obtain a first verification result.
[0077] Step S220: When it is determined that the first verification result corresponding to the preset proxy address is verification passed, a risk assessment is performed on the preset proxy address.
[0078] Step S230: When it is determined that the risk level of the preset proxy address is lower than the preset risk threshold, the preset proxy address is stored in a preset database to construct a preset proxy pool.
[0079] Step S240: obtaining a target proxy address from a preset proxy pool according to a preset allocation strategy.
[0080] Step S250: Collect data through the target proxy address to obtain data to be analyzed.
[0081] In this specific embodiment, the embodiment of the present invention first obtains a preset proxy address to verify the validity of the preset proxy address and obtain a first verification result. Specifically, the preset proxy address in the embodiment of the present invention refers to a proxy IP in the corresponding network area, such as a free or paid proxy IP in the intranet. Accordingly, after collecting the relevant proxy IPs, the embodiment of the present invention verifies the validity of these proxy IPs, including verifying the connectivity, anonymity, speed, stability and risk level of the proxy IPs. For example, the verification formula of the risk level in the embodiment of the present invention is shown in the following formula (1):
[0082]
[0083] Among them, SQ represents the risk assessment coefficient, bi represents the network risk level of the obtained IP address, ci represents the number of times the IP address is attacked, di represents the number of times the IP address executes a command, i represents the number of the IP address i = 1, 2, 3, 4, .... n, SQ>1 indicates high risk, otherwise low risk.
[0084] Next, when it is determined that the first verification result corresponding to the preset proxy address is a verification pass, the embodiment of the present invention performs a risk assessment on the preset proxy address. When it is determined that the risk level of the preset proxy address is lower than the preset risk threshold, the preset proxy address is stored in the preset database to construct a preset proxy pool. Specifically, in the embodiment of the present invention, when the validity verification result of the preset proxy address is a verification pass, that is, the preset proxy address is a valid proxy address, the embodiment of the present invention further determines whether it meets the risk condition. Correspondingly, when it is determined that the risk level corresponding to the valid proxy address is also lower than the preset risk threshold, the embodiment of the present invention generates a corresponding linkage maintenance instruction and stores it in the preset database to construct an IP proxy pool, that is, a preset proxy pool. Finally, the embodiment of the present invention allocates a target proxy address from the preset proxy pool according to a preset allocation strategy, and then collects data through the target proxy address to obtain data to be analyzed. Specifically, after the preset proxy pool is constructed, the embodiment of the present invention allocates a target proxy address from the preset proxy pool according to the preset allocation strategy according to the corresponding data collection requirements, and uses the allocated proxy address to collect data to obtain data to be analyzed. For example, the embodiment of the present invention allocates proxy IPs randomly or according to policies from the Redis proxy pool according to demand, and collects data through the proxy IPs.
[0085] Reference Figure 3 In some embodiments of the present invention, after performing the step of storing the preset proxy address in a preset database to construct a preset proxy pool when determining that the risk level of the preset proxy address is lower than the preset risk threshold, the data analysis method provided by the embodiment of the present invention further includes but is not limited to the following steps:
[0086] Step S310: When it is determined that the update period is satisfied, the validity of each preset proxy address in the preset proxy pool is verified to obtain a second verification result.
[0087] Step S320: updating the preset proxy pool according to the second verification result.
[0088] In this specific embodiment, after the preset proxy pool is constructed, the embodiment of the present invention periodically updates the proxy addresses in the preset proxy pool. Specifically, when it is determined that the update cycle is met, the embodiment of the present invention verifies the validity of each preset proxy address in the preset proxy pool to obtain a second verification result, and then updates the preset proxy pool according to the second verification result. Accordingly, the update cycle in the embodiment of the present invention can be customized, such as every 24 hours, every week, or every month. When it is determined that the time from the last update of the preset proxy pool reaches the update cycle, the embodiment of the present invention verifies the validity and updates each proxy IP (preset proxy address) in the preset proxy pool to ensure the stability and effectiveness of the preset proxy pool. For example, when the second verification result is verification passed, that is, the preset proxy address is valid, the preset proxy address is retained. Conversely, when the second verification result is verification failed, the corresponding preset proxy pool is removed from the preset proxy pool. Among them, the overall process of constructing the preset proxy pool in the embodiment of the present invention is as follows. Figure 4 shown.
[0089] In some embodiments of the present invention, the first preprocessed data includes a word cloud image and a word meaning network image. Figure 5 In the embodiment of the present invention, the first data preprocessing is performed on the data to be analyzed to obtain the first preprocessed data, including but not limited to the following steps:
[0090] Step S410: Perform word segmentation processing on the data to be analyzed by using a preset word segmentation function to obtain a word segmentation result list, wherein the word segmentation result list includes a number of preset word segmentations.
[0091] Step S420: traverse the word segmentation result list to perform a preset condition check on each preset word segmentation to obtain a condition check result. The preset condition check includes determining whether the preset word segmentation is not in the preset stop word character table and determining whether the word length of the preset word segmentation is greater than a preset length.
[0092] Step S430: When the conditional check result is that the preset word segmentation is not in the preset stop word character table, and the word length of the preset word segmentation is greater than the preset length, the preset word segmentation is added to the check result list.
[0093] Step S440: Generate a word cloud image using a preset word cloud function according to the inspection result list.
[0094] Step S450: Perform text filtering on the word segmentation result list according to the verification result list to obtain target text content.
[0095] Step S460: Calculate the word frequency of the target text content to determine the target keyword through the calculated word frequency data.
[0096] Step S470: When it is determined that the number of target keywords in the first text of the target text content is greater than the preset threshold, adjust the association matrix corresponding to the first text according to the preset change value.
[0097] Step S480: Construct a semantic network image based on the association matrix.
[0098] In this specific embodiment, the first preprocessed data includes a word cloud image and a semantic network image. Correspondingly, the embodiment of the present invention first performs word segmentation on the data to be analyzed through a preset word segmentation function to obtain a list of word segmentation results. Specifically, in the embodiment of the present invention, the list of word segmentation results includes several preset word segments. Among them, the preset word segmentation function in the embodiment of the present invention refers to a function for performing word segmentation on the input text, such as the word segmentation function in the Jieba library. Correspondingly, the embodiment of the present invention inputs the data to be analyzed into the preset word segmentation function to perform word segmentation on the corresponding text data, and obtains a segmented list, that is, a list of word segmentation results. Then, the embodiment of the present invention traverses the list of word segmentation results to perform a preset condition check on each preset word segment to obtain a condition check result. Specifically, the preset condition check in the embodiment of the present invention includes determining whether the preset word segment is not in the preset stop word character table, and whether the word length of the preset word segment is greater than the preset length. Among them, the preset stop word character table in the embodiment of the present invention refers to a set containing several stop words, and stop words refer to words that frequently appear in the text but contribute little to the meaning of the text, such as "de", "le", etc. The embodiment of the present invention determines whether the preset word segment is a meaningful word segment by determining whether the preset word segment is not in the preset stop word character table. Correspondingly, in the embodiment of the present invention, by determining whether the word length of the preset word segment is greater than the preset length, it is determined whether the information contained in the preset word segment is sufficient. For example, as Figure 6 shown, the preset length in the embodiment of the present invention is 1, and by determining whether the word length of the preset word segment is greater than 1, words with a single character are excluded. Then, when it is determined that the condition check result is that the preset word segment is not in the preset stop character table and the word length of the preset word segment is greater than the preset length, add the preset word segment to the check result list. Specifically, when the preset word segment is not in the preset stop character table, that is, the preset word segment is a meaningful word segment, and at the same time the word length of the preset word segment is greater than the preset length, that is, the preset word segment has corresponding information, then select the preset word segment and add it to a new list to construct a check result list. Then, the embodiment of the present invention generates a word cloud image according to the check result list through a preset word cloud function. Specifically, the embodiment of the present invention converts the generated check result list into a long text in string format and generates a word cloud image through a preset word cloud function, such as a function in the WordCloud library. In addition, as Figure 6 shown, after generating the word cloud image in the embodiment of the present invention, save the word cloud image as a PNG file.
[0099] Furthermore, the embodiment of the present invention performs text filtering on the word segmentation result list according to the test result list to obtain the target text content, and then performs word frequency calculation on the target text content to determine the target keyword through the corresponding word frequency data, so that when it is determined that the number of target keywords in the first text in the target text content is greater than a preset threshold, the association matrix of the first text is adjusted according to the preset change value, and a word meaning network image is constructed according to the association matrix. Specifically, when performing word meaning network analysis, the embodiment of the present invention first performs word segmentation on the entire text, and performs preset condition tests on each word segmentation obtained to obtain a test result list, that is, the word segments in the test result list are all not in the stop word character table, and the length of the word segmentation is greater than the preset length. Accordingly, the embodiment of the present invention filters the text content in the word segmentation result list according to each preset word segmentation in the test result list, so as to screen out the target text content that satisfies the requirement of not being in the stop word character table and having a word length greater than the preset length. For example, if Figure 7 As shown, the embodiment of the present invention performs text segmentation by calling the segmentation function in the Jieba library, and filters out the words in the stop character table and the words with a length less than 1, so as to obtain the target text content. Then, the embodiment of the present invention performs word frequency calculation on the text content (target text content) after the entire segmentation and stop word filtering, and obtains word frequency data. Since all the segmentations cannot be reflected in the graph, the embodiment of the present invention uses the first N words with the highest word frequency as keywords, i.e., target keywords. Then, the target text content is traversed, and when the number of target keywords appearing in the same text (such as the same sentence) is greater than the preset threshold, such as two keywords appearing in the same sentence, the value of the corresponding position in the association matrix of the embodiment of the present invention increases by a set value (preset change value), thereby realizing the adjustment of the association matrix. Finally, the embodiment of the present invention converts the association matrix between words into a graph, uses the values in the association matrix as the weights of the edges, uses the target keywords as nodes, and uses the Networkx library to construct a graph structure, thereby generating a PNG file.
[0100] Combination Figure 1 , refer to Figure 8 In some embodiments of the present invention, before performing a first data preprocessing on the data to be analyzed to obtain the first preprocessed data, the data analysis method provided by the embodiment of the present invention further includes but is not limited to the following steps:
[0101] Step S510: Perform second data preprocessing on the data to be analyzed to obtain second preprocessed data, wherein the second data preprocessing includes data deduplication processing and outlier processing.
[0102] In this specific embodiment, before the first data preprocessing is performed on the data to be analyzed, the embodiment of the present invention also performs a second data preprocessing on the data to be analyzed to obtain the second preprocessed data, and then performs the first data preprocessing on the second preprocessed data to obtain the first preprocessed data. Specifically, there may be data anomalies in the data to be analyzed collected by the embodiment of the present invention, such as duplicate data, erroneous data, etc. Therefore, the embodiment of the present invention performs the second data preprocessing before the text preprocessing is performed on the data to be analyzed to improve the data quality and improve the accuracy and reliability of data analysis. Accordingly, the second data preprocessing in the embodiment of the present invention includes data deduplication processing and outlier processing. For example, the embodiment of the present invention uses the DISTINCT keyword or the GROUP BY clause through SQL statements to remove duplicate data and retain unique data. At the same time, the processing of outliers in the embodiment of the present invention includes deleting erroneous data, replacing outliers, and supplementing data by the mean replacement method.
[0103] Reference Fig. 9 In some embodiments of the present invention, performing predictive analysis on the first preprocessed data to obtain data analysis results includes but is not limited to the following steps:
[0104] Step S610: Aggregate the first pre-processed data corresponding to each target proxy address to obtain preset aggregated data.
[0105] Step S620: input the preset aggregated data into the preset prediction model for prediction analysis to obtain data analysis results, wherein the weight data of the preset prediction model is determined by an optimal weighting algorithm.
[0106] In this specific embodiment, the embodiment of the present invention first aggregates the first preprocessed data corresponding to each target proxy address to obtain preset aggregated data. Specifically, the embodiment of the present invention can collect corresponding data to be analyzed through each target proxy address. Among them, the embodiment of the present invention obtains the corresponding first preprocessed data after preprocessing the data to be analyzed. Therefore, before predicting and analyzing the data, the embodiment of the present invention first preprocesses the data collected by multiple proxy IPs (target proxy addresses), and then summarizes and aggregates them to obtain preset aggregated data. Then, the embodiment of the present invention inputs the preset aggregated data into the preset prediction model for prediction analysis to obtain data analysis results. Specifically, in the data prediction process, the embodiment of the present invention obtains a time series data set {X1, X2, X3.....X n}, the input source St at the current moment T includes the current input value Xt and the inflow information Xt-1 at the previous moment, and the output includes the predicted value Nt at the moment T and the information flow dt flowing into the next moment. Accordingly, the embodiment of the present invention determines the weight of the combined prediction model (preset prediction model) through the optimal weighting algorithm, and obtains the error matrix F, where F=[Y m -Y 1m ,Y2-Y 2m ], m represents the time point of the mth prediction and the trend of the predicted data.
[0107] Combination Fig. 9 , refer to Fig.10 In some embodiments of the present invention, after inputting the preset aggregated data into the preset prediction model for prediction analysis and obtaining the data analysis results, the data analysis method provided by the embodiment of the present invention also includes but is not limited to the following steps:
[0108] Step S710: Visualize the preset aggregated data and data analysis results to obtain visualized data information.
[0109] In this specific embodiment, after the prediction analysis obtains the data analysis results, the embodiment of the present invention performs data visualization on the preset aggregated data and the data analysis results to obtain visualized data information. Specifically, in order to be able to display the data analysis results more intuitively and in-depth, the embodiment of the present invention performs data visualization on the relevant data information. Accordingly, the embodiment of the present invention performs data visualization on the aggregated preset aggregated data and the predicted data information (data analysis results), such as generating corresponding graphs, tables, and text information, thereby obtaining visualized data information.
[0110] The following describes the solution of the embodiment of the present invention in detail in combination with a specific data analysis scenario:
[0111] For example, refer to Fig.11 , Fig.11It is an overall flow chart of data analysis provided by an embodiment of the present invention. Specifically, the embodiment of the present invention first constructs an IP proxy pool to randomly or strategically allocate proxy IPs through the IP proxy pool to obtain a target proxy IP, and then uses the target proxy IP to collect basic information, business data, and real-time policy information of users from the cloud server. Next, the embodiment of the present invention preprocesses the acquired data information, including data deduplication, outlier processing, word frequency analysis, and word meaning network analysis. Then, the embodiment of the present invention summarizes the data collected by each proxy IP, thereby aggregating to obtain preset aggregate data. Furthermore, the embodiment of the present invention performs data prediction on the preset combined data to obtain data analysis results. Accordingly, the embodiment of the present invention visualizes the summarized data information and the predicted data analysis results, and presents them in the form of graphs, tables, and text.
[0112] It is easy to understand that the embodiment of the present invention can improve the stability of IP address access and data collection efficiency while reducing the access pressure of a single proxy IP by constructing a proxy IP pool, using multiple proxy IPs to make data requests at the same time, and reasonably allocating proxy IPs. Among them, the embodiment of the present invention can reduce the problem of IP being blocked and unable to be used normally by constantly changing proxy IPs, and can achieve efficient management and utilization of proxy IPs, and prevent instability during data collection. At the same time, the embodiment of the present invention performs in-depth analysis and mining of the collected data through word meaning network analysis and word frequency analysis to improve the integrity and adequacy of data processing. For example, the example of the present invention uses advanced text analysis technologies such as Jieba word segmentation and BERT model to deeply mine the semantic information of the collected data text, generate word cloud diagrams and word meaning network diagrams, and perform predictive analysis on the data, so that decision makers can understand the policy content and potential impact more intuitively and deeply, and improve the depth and breadth of policy analysis. In addition, the embodiment of the present invention integrates new dynamic data sources and analysis tools, has good scalability, can flexibly adapt to new needs and challenges in the field of communications, and maintain long-term technical competitiveness.
[0113] See also Fig.12 The present application also provides a data analysis system that can implement the above data analysis method. The system includes:
[0114] The first module 810 is used to determine a target proxy address from a preset proxy pool, so as to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address.
[0115] The second module 820 is used to perform a first data preprocessing on the data to be analyzed to obtain first preprocessed data, wherein the first data preprocessing includes word segmentation processing and word meaning network analysis processing.
[0116] The third module 830 is used to perform predictive analysis on the first preprocessed data to obtain data analysis results.
[0117] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0118] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above data analysis method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0119] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0120] See also Fig.13 , Fig.13 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0121] The processor 910 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0122] The memory 920 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 920 can store an operating system and other application programs. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 920, and the processor 910 calls and executes the data analysis method of the embodiment of the present application;
[0123] Input / output interface 930, used to implement information input and output;
[0124] Communication interface 940, used to realize communication interaction between the device and other devices, which can be realized through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0125] bus 950 , which transmits information between the various components of the device (e.g., processor 910 , memory 920 , input / output interface 930 , and communication interface 940 );
[0126] The processor 910 , the memory 920 , the input / output interface 930 , and the communication interface 940 are connected to each other in communication within the device via a bus 950 .
[0127] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned data analysis method is implemented.
[0128] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0129] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0130] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0131] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0132] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0133] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0134] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0135] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0136] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0137] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0139] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0140] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A data analysis method, characterized in that: The method comprises the following steps: Determine a target proxy address from a preset proxy pool, so as to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address; Performing a first data preprocessing on the data to be analyzed to obtain first preprocessed data; wherein the first data preprocessing includes word segmentation processing and word meaning network analysis processing; Performing predictive analysis on the first preprocessed data to obtain data analysis results.
2. The method according to claim 1, characterized in that: The step of determining a target proxy address from a preset proxy pool to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address includes: Obtaining a preset proxy address to verify the validity of the preset proxy address and obtain a first verification result; When it is determined that the first verification result corresponding to the preset proxy address is verification passed, performing a risk assessment on the preset proxy address; When it is determined that the risk level of the preset proxy address is lower than the preset risk threshold, the preset proxy address is stored in a preset database to construct the preset proxy pool; Allocating the target proxy address from the preset proxy pool according to a preset allocation strategy; Data is collected through the target proxy address to obtain the data to be analyzed.
3. The method according to claim 2, characterized in that After executing the step of determining that the risk level of the preset proxy address is lower than the preset risk threshold, storing the preset proxy address in a preset database to construct the preset proxy pool, the method further includes: When it is determined that the update period is satisfied, performing validity verification on each of the preset proxy addresses in the preset proxy pool to obtain a second verification result; The preset proxy pool is updated according to the second verification result.
4. The method according to claim 1, characterized in that: The first preprocessed data includes a word cloud image and a word meaning network image; The step of performing first data preprocessing on the data to be analyzed to obtain first preprocessed data includes: Performing word segmentation processing on the data to be analyzed by using a preset word segmentation function to obtain a word segmentation result list; wherein the word segmentation result list includes a plurality of preset word segmentations; Traversing the word segmentation result list to perform a preset condition check on each of the preset word segments to obtain a condition check result; wherein the preset condition check includes determining whether the preset word segment is not in a preset stop word character table, and determining whether the word length of the preset word segment is greater than a preset length; When it is determined that the conditional test result is that the preset word segmentation is not in the preset stop word character table, and the word length of the preset word segmentation is greater than the preset length, the preset word segmentation is added to the test result list; Generate a word cloud image by using a preset word cloud function according to the inspection result list; Performing text filtering on the word segmentation result list according to the inspection result list to obtain target text content; Performing word frequency calculation on the target text content to determine the target keyword through the word frequency data obtained by calculation; When it is determined that the number of the target keywords in the first text in the target text content is greater than a preset threshold, adjusting the association matrix corresponding to the first text according to a preset change value; A word meaning network image is constructed based on the association matrix.
5. The method according to claim 1, characterized in that: Before performing the first data preprocessing on the data to be analyzed to obtain the first preprocessed data, the method further includes: A second data preprocessing is performed on the data to be analyzed to obtain second preprocessed data; wherein the second data preprocessing includes data deduplication processing and outlier processing.
6. The method according to claim 1, characterized in that The performing predictive analysis on the first preprocessed data to obtain a data analysis result includes: Aggregating the first pre-processed data corresponding to each of the target proxy addresses to obtain preset aggregated data; The preset aggregated data is input into a preset prediction model for prediction analysis to obtain the data analysis result; wherein the weight data of the preset prediction model is determined by an optimal weighting algorithm.
7. The method according to claim 6, characterized in that After inputting the preset aggregated data into a preset prediction model for prediction analysis to obtain the data analysis result, the method further includes: The preset aggregated data and the data analysis results are visualized to obtain visualized data information.
8. A data analysis system, characterized in that: The system comprises: The first module is used to determine a target proxy address from a preset proxy pool, so as to obtain the data to be analyzed from a target cloud server cluster according to the target proxy address; The second module is used to perform a first data preprocessing on the data to be analyzed to obtain first preprocessed data; wherein the first data preprocessing includes word segmentation processing and word meaning network analysis processing; The third module is used to perform predictive analysis on the first preprocessed data to obtain data analysis results.
9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.