Webpage source analysis method and device, equipment and medium

By accessing JSSDK to the website to monitor web page jump behavior and performing real-time statistical analysis and offline integration processing, the impact on website performance and security in the existing technology is solved, and an efficient web source analysis is achieved without intrusiveness.

CN119996146APending Publication Date: 2025-05-13ISOFTSTONE INFORMATION TECHNOLOGY (GROUP) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510073052.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, embedding original tracking code in the website for web source analysis may affect the website loading speed and performance, and there are security risks.

Method used

Connect the JSSDK to the website, monitor the web page jump behavior through the JSSDK, obtain user operation information and store it in a distributed message system, and use Spark to perform real-time statistical analysis and offline integration processing to realize web page source analysis.

Benefits of technology

It realizes non-invasive web source analysis, avoids the impact on website performance and security, and ensures the efficient and safe operation of the website.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996146A_ABST
    Figure CN119996146A_ABST
Patent Text Reader

Abstract

The invention discloses a webpage source analysis method and device, equipment and a medium. The method comprises the following steps: accessing a JS SDK (JavaScript Software Development Kit) into a target website, and monitoring whether the target website has a webpage jump behavior or not through the JS SDK; if yes, obtaining user operation information corresponding to the webpage jump behavior, generating a log record from the user operation information, and storing the log record in the target distributed message system; obtaining candidate log records from the target distributed message system based on the first time interval, and grouping the candidate log records by using Spark to obtain a plurality of candidate RDDs; performing real-time statistical analysis on the target log record in each candidate RDD to obtain a target RDD, and storing the target RDD in a target distributed storage system; and obtaining a plurality of target RDDs from the target distributed storage system based on a second time interval, and performing offline integration processing on target log records in the target RDDs to obtain a webpage source analysis result of the target website. According to the scheme, the safety and the high efficiency of website operation can be effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to a web page source analysis method, device, equipment and medium. Background Art

[0002] Web page source analysis is crucial for understanding user access paths, evaluating marketing channel effectiveness, optimizing website content and structure, and monitoring potential security threats. It is the key basis for formulating effective network strategies and ensuring the healthy development of websites.

[0003] Web page source analysis can effectively track the internal and external reference links of a web page. In the related art, traffic analysis tools are used to perform web page source analysis. However, traffic analysis tools need to embed the original tracking code in the website, and the connection is not lightweight enough, which may affect the loading speed of the website, especially at high throughput. The impact on website performance is greater. In addition, when there are errors in the embedded original tracking code, it may affect the normal operation of the website, so security cannot be guaranteed. Summary of the invention

[0004] The present invention provides a web page source analysis method, device, equipment and medium, which can perform non-intrusive web page source analysis on a website by connecting JSSDK to the website, without affecting the performance and operation of the website itself, thereby ensuring the security and efficiency of the website operation.

[0005] According to one aspect of the present invention, a method for analyzing the source of a web page is provided, the method comprising:

[0006] Connecting JSSDK to the target website, and monitoring whether the target website has web page jump behavior through the JSSDK; wherein the JSSDK realizes monitoring based on user operations of the target website;

[0007] If it exists, then obtain the user operation information corresponding to the web page jump behavior, and generate a log record of the user operation information and store it in the target distributed messaging system;

[0008] Obtain candidate log records from the target distributed messaging system based on a first time interval, and group the candidate log records using Spark to obtain a plurality of candidate RDDs;

[0009] Perform real-time statistical analysis on the target log records in each candidate RDD to obtain a target RDD, and store the target RDD in a target distributed storage system;

[0010] Based on a second time interval, multiple target RDDs are obtained from the target distributed storage system, and target log records in the target RDDs are offline integrated to obtain web page source analysis results of the target website; wherein the second time interval is greater than the first time interval.

[0011] According to another aspect of the present invention, a web page source analysis device is provided, the device comprising:

[0012] A web page jump behavior judgment module, used to access JSSDK in the target website, and monitor whether the target website has web page jump behavior through the JSSDK; wherein the JSSDK realizes monitoring based on user operations of the target website;

[0013] A user operation information acquisition module, used to acquire the user operation information corresponding to the web page jump behavior, if any, and generate a log record of the user operation information and store it in the target distributed message system;

[0014] A log record grouping module, configured to obtain candidate log records from the target distributed messaging system based on a first time interval, and group the candidate log records using Spark to obtain a plurality of candidate RDDs;

[0015] A log record real-time analysis module is used to perform real-time statistical analysis on the target log record in each candidate RDD to obtain a target RDD, and store the target RDD in a target distributed storage system;

[0016] A log record offline analysis module is used to obtain multiple target RDDs from the target distributed storage system based on a second time interval, and perform offline integration processing on the target log records in the target RDD to obtain the web page source analysis results of the target website; wherein the second time interval is greater than the first time interval.

[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0018] at least one processor; and,

[0019] a memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the web page source analysis method described in any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the web page source analysis method described in any embodiment of the present invention when executed.

[0022] The technical solution of the embodiment of the present invention is to connect JSSDK in the target website, and monitor whether the target website has web page jump behavior through JSSDK; wherein JSSDK realizes monitoring based on the user operation of the target website; if it exists, the user operation information corresponding to the web page jump behavior is obtained, and the log record generated by the user operation information is stored in the target distributed message system; based on the first time interval, the candidate log record is obtained from the target distributed message system, and the candidate log record is grouped by Spark to obtain multiple candidate RDDs; the target log record in each candidate RDD is subjected to real-time statistical analysis to obtain the target RDD, and the target RDD is stored in the target distributed storage system; based on the second time interval, multiple target RDDs are obtained from the target distributed storage system, and the target log record in the target RDD is subjected to offline integration processing to obtain the web page source analysis result of the target website; wherein the second time interval is greater than the first time interval. This technical solution can perform non-intrusive web page source analysis on the website by connecting JSSDK in the website, which will not affect the performance and operation of the website itself, and ensure the security and efficiency of the website operation.

[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0025] Figure 1 is a flow chart of a web page source analysis method provided according to the first embodiment of the present invention;

[0026] Figure 2 is a schematic diagram of a web page source analysis method provided according to Embodiment 1 of the present invention;

[0027] Figure 3 is a flow chart of a web page source analysis method provided according to Embodiment 2 of the present invention;

[0028] Figure 4is a structural diagram of a web page source analysis device provided according to Embodiment 3 of the present invention;

[0029] Figure 5 It is a structural schematic diagram of an electronic device for implementing a web page source analysis method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first", "second", "target", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment 1

[0033] Figure 1 This is a flow chart of a web page source analysis method provided in the first embodiment of the present invention. This embodiment is applicable to the case of non-intrusive analysis of the source of a website page. The method can be executed by a web page source analysis device. The web page source analysis device can be implemented in the form of hardware and / or software. The web page source analysis device can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:

[0034] S110, connecting JSSDK to the target website, and monitoring whether there is web page jump behavior in the target website through JSSDK; wherein JSSDK realizes monitoring based on user operations of the target website.

[0035] The target website may refer to a website that needs to be analyzed for the source of a web page, which may be one or more websites, and may be flexibly set according to actual needs. JSSDK (JavaScript Software Development Kit) provides a series of pre-written JavaScript codes that implement certain functions, such as user authentication, data analysis, social sharing, etc. Exemplarily, user operations may include a user entering a URL, i.e., a URL (Uniform Resource Locator), in a website, or a user clicking a tag in a website page (such as forward, backward, link, etc.).

[0036] In this embodiment, only one line of reference code needs to be added to the target website to directly call the JSSDK pre-stored in a public location inside or outside the website, thereby connecting the JSSDK to the target website, so as to realize web page jump behavior monitoring by driving the triggering of web page embedding through the user's operation behavior on the target website, without the need to embed the original tracking code in the target website. Therefore, the lightweight JSSDK docking method can be used to realize non-intrusive website monitoring, which will not affect the website performance (such as loading speed, etc.) and the safe operation of the website.

[0037] In addition, the prior art usually uses traversal retrieval (i.e., traversal retrieval from the website homepage layer by layer) to search for incremental pages when performing website monitoring. When there are many page levels, the search speed of incremental pages (i.e., changed web pages) will be greatly reduced, and it is easy to be intercepted by the interceptor. In this embodiment, website monitoring driven by user behavior can effectively avoid being intercepted during active traversal retrieval and being affected by the depth of the web page level, and can capture incremental pages in time, improve website monitoring efficiency, and reduce website monitoring complexity. After accessing JSSDK in the target website, you can use JSSDK to monitor whether the target website has web page jump behavior.

[0038] In this embodiment, optionally, JSSDK is used to monitor whether the target website has web page jump behavior, including: if JSSDK monitors that the target website has web page reloading behavior, it is determined that the target website has web page jump behavior.

[0039] It should be noted that when a web page reload occurs, it indicates that a web page jump occurs. Therefore, if the target website is detected to have a web page reload through JSSDK, it can be determined that the target website has a web page jump, that is, the target website has a web page jump.

[0040] In this embodiment, optionally, JSSDK is used to monitor whether the target website has any web page jump behavior, including: if JSSDK monitors that the target website has a user clicking a preset tag, it is determined that the target website has a web page jump behavior; wherein, the preset tag refers to a pre-set tag that can achieve web page jump.

[0041] Exemplarily, the preset tag can be a forward / back button of a browser or some preset link (such as a hyperlink), and a web page jump can be realized by clicking on the preset tag. Specifically, when the JSSDK monitors that a user clicks on a preset tag on the target website, it can be determined that the target website has a web page jump behavior.

[0042] In this embodiment, optionally, JSSDK is used to monitor whether the target website has web page jump behavior, including: if JSSDK monitors that the target website has internal routing jump behavior, it is determined that the target website has web page jump behavior.

[0043] It should be noted that for the page jump behavior performed by internal routing jumps for single-page applications using technology stacks such as Vue and React, this jump does not increase the history stack and does not trigger page loading. Currently, there is no special event monitoring for this jump behavior in external JS monitoring. Failure to capture internal routing jumps will cause the internal reference link analysis data to be missing, resulting in incomplete and inaccurate web page source analysis. In response to the above problems, in this embodiment, JSSDK can be used to monitor whether the target website has internal routing jump behavior. If internal routing jump behavior is detected, it can be determined that the target website has web page jump behavior.

[0044] In this embodiment, optionally, internal route jump behavior of the target website is monitored through JSSDK, including: judging through the popstate event listener in the target website whether to change the web page history record of the target website by calling a preset function through JSSDK; wherein the preset function includes history.pushState and history.replaceState; if so, obtaining the web page URL before and after the web page history record is changed for comparative analysis to obtain a comparison result; wherein the comparison result is consistent or inconsistent; if the comparison result is inconsistent, it is determined that internal route jump behavior occurs in the target website.

[0045] Among them, the popstate event is an event provided by the HTML5 History API, which can be used to monitor changes in the browser history stack. pushState and replaceState are two methods provided in the HTML5 History API that allow developers to operate the browser's history stack without reloading the page. Among them, the pushState method can be used to add a new state to the history stack, which can be implemented by calling the history.pushState() function. When this method is called, the new state is added to the top of the current state, and the URL in the browser address bar is updated accordingly, but the page is not reloaded. The replaceState method can be used to modify the current history entry instead of adding a new entry to the history stack, which can be implemented by calling the history.replaceState() function. When this method is called, the current state is replaced with the new state, and the URL in the browser address bar is updated accordingly, but the page is not reloaded. The popstate event is triggered when the history entry changes, such as when the user clicks the browser's forward or back button, or when the history.pushState() or history.replaceState() function is called through JS to change the browser's history.

[0046] In this embodiment, it is necessary to add a popstate event listener in the target website in advance, and determine whether the web page history record of the target website has been changed by calling history.pushState() or history.replaceState() through the popstate event listener. If so, the web page URLs before and after the web page history record is changed are obtained, and the two web page URLs are compared and analyzed to obtain the comparison result (consistent or inconsistent). If the comparison result is inconsistent, it can be determined that the target website has an internal route jump behavior; if the comparison result is consistent, it can be determined that the target website has not had an internal route jump behavior. In addition, the popstate event listener can also capture the navigation triggered by the forward / back button of the browser, and accurately report the user's operation log. In this way, almost all types of front-end route changes can be captured, which is independent of a specific front-end framework and can be used with React, Vue, Angular or other frameworks using HTML5 HistoryAPI. And it is non-invasive and does not require modification of the application code.

[0047] Through such a setting, this solution can effectively capture the web page jump behavior caused by internal routing jumps, thereby ensuring the accuracy and comprehensiveness of web page source analysis.

[0048] S120: If it exists, obtain the user operation information corresponding to the web page jump behavior, and generate a log record of the user operation information and store it in the target distributed message system.

[0049] In this embodiment, if the target website is detected to have a web page jump behavior through JSSDK, it is necessary to obtain the user operation information corresponding to the web page jump behavior. Exemplarily, the user operation information may include user UUID (Universally Unique Identifier), user IP (Internet Protocol), current web page URL and jump Refer link, etc. It should be noted that when there are many websites and web pages to be monitored or there are many users visiting the website, the corresponding user operation information is numerous and messy. If the user operation information is directly processed, there may be insufficient processing capacity and security issues. Therefore, after obtaining the user operation information corresponding to the web page jump behavior, the user operation information can be generated and stored in the target distributed message system. Exemplarily, the target distributed message system can be a Kafka system. Among them, the Kafka system is a distributed, partitioned, multi-copy, Zookeeper-coordinated distributed message system. As a high-throughput, scalable message queue system, it is designed to process a large amount of real-time data streams.

[0050] S130, obtaining candidate log records from the target distributed messaging system based on the first time interval, and grouping the candidate log records using Spark to obtain a plurality of candidate RDDs.

[0051] In this embodiment, after the user operation information generates a log record and stores it in the target distributed messaging system, multiple log records can be obtained from the target distributed messaging system as candidate log records based on the first time interval. Among them, the first time interval can refer to a shorter time interval preset according to actual needs, such as 30 seconds or 1 minute. Then Spark is used to divide the candidate log records into DStream (DiscretizedStream, discretized data stream) discretized data streams of preset time window size, and multiple discrete RDDs (Resilient Distributed Datasets) are obtained as candidate RDDs. Among them, Spark is an open source computing framework based on memory, which can be used for real-time stream computing of real-time data streams; RDD is the core component in Spark, which can be used to realize distributed storage and processing of data. Each RDD contains a batch of data within a specific time interval, and multiple RDDs can be processed in parallel, thereby improving data processing efficiency.

[0052] S140, performing real-time statistical analysis on the target log records in each candidate RDD to obtain a target RDD, and storing the target RDD in a target distributed storage system.

[0053] In this embodiment, each target log record in each candidate RDD can be subjected to real-time statistical analysis based on preset real-time analysis indicators to obtain real-time analysis results. Exemplarily, data indicators that need to be analyzed in real time (i.e., with high real-time requirements) can be pre-set as preset real-time analysis indicators according to actual needs, such as the number of IP visits, the number of views, and the number of visitors of the current web page. Then, the candidate RDD is updated according to the real-time analysis results to obtain the target RDD, and the target RDD is stored in batches in the target distributed storage system according to the time dimension as a data source for subsequent offline analysis. Exemplarily, the target distributed storage system can be set to an HBase cluster. Among them, the HBase cluster is a highly reliable, high-performance, column-oriented, scalable distributed storage system, which is built based on the Hadoop distributed file system (HDFS) and coordinated and managed by Zookeeper. By performing Spark streaming calculation and analysis on the basis of RDD, quasi-real-time calculation and analysis can be achieved for data in minute-level windows.

[0054] S150, obtaining multiple target RDDs from the target distributed storage system based on a second time interval, and performing offline integration processing on the target log records in the target RDD to obtain a web page source analysis result of the target website; wherein the second time interval is greater than the first time interval.

[0055] In this embodiment, multiple target RDDs can be obtained from the target distributed storage system based on the second time interval. The second time interval may refer to a longer time interval preset according to actual needs, such as 1 hour or 1 day. It should be noted that since the webpage source analysis does not require high real-time performance, it can be processed in an offline analysis manner. After obtaining multiple target RDDs, the target log records in the target RDD can be offline integrated based on the preset offline analysis indicators to obtain the webpage source analysis results of the target website. Among them, data indicators that need to be analyzed offline (i.e., low real-time performance requirements) can be pre-set as preset offline analysis indicators according to actual needs, such as the current webpage URL, the source webpage URL, and the number of jumps (used to describe the number of jumps from the source webpage to the current webpage). Exemplarily, taking the number of jumps as an example, the four monitoring indicators (website id) sysId + (user id) uuid + (webpage url) web + (refer link) refer can be grouped and aggregated to obtain an RDD subset, and the number of jumps for each link in this batch and the corresponding source link can be obtained after grouping and aggregating the data in this batch. By using the lineage mechanism of RDD, data replication can be reduced, partition parallel processing can be supported, and complex correlation relationship calculations of indicators can be realized. After obtaining the source analysis results of the target website's web pages, the HBase distributed database can be used to save the data analysis results, and each HBase table can be pre-partitioned and fixed with partition points to realize massive data storage and real-time query.

[0056] It should be noted that the URL and Refer aggregation analysis realizes a bidirectional queryable link relationship map. Among them, the bidirectional query includes forward query and reverse query. Specifically, forward query refers to retrieving all reference sources (Refer) through the target URL; reverse query refers to retrieving all target pages referenced by it through the Refer URL. In addition. This embodiment can support parsing and visualizing multi-level call relationships between pages for complex call relationship analysis, which is particularly suitable for complex single-page applications (SPA) or front-end applications of microservice architecture; detect and count independent pages without entry links for isolated page identification, and find incorrectly indexed content or potential security issues.

[0057] The technical solution of the embodiment of the present invention is to connect JSSDK in the target website, and monitor whether the target website has web page jump behavior through JSSDK; wherein JSSDK realizes monitoring based on the user operation of the target website; if it exists, the user operation information corresponding to the web page jump behavior is obtained, and the user operation information is generated into a log record and stored in the target distributed message system; based on the first time interval, the candidate log record is obtained from the target distributed message system, and the candidate log record is grouped by Spark to obtain multiple candidate RDDs; the target log record in each candidate RDD is subjected to real-time statistical analysis to obtain the target RDD, and the target RDD is stored in the target distributed storage system; based on the second time interval, multiple target RDDs are obtained from the target distributed storage system, and the target log record in the target RDD is subjected to offline integration processing to obtain the web page source analysis result of the target website; wherein the second time interval is greater than the first time interval. This technical solution can realize non-intrusive website detection by connecting JSSDK in the website, and the website page source analysis is performed by combining real-time analysis with offline analysis, which will not affect the performance and operation of the website itself, and ensure the security and efficiency of the website operation.

[0058] In this embodiment, optionally, after obtaining candidate log records from the target distributed messaging system based on the first time interval, it also includes: performing data preprocessing on the candidate log records; wherein the data preprocessing includes data cleaning and format conversion; accordingly, using Spark to group the candidate log records to obtain multiple candidate RDDs, including: using Spark to group the candidate log records after data preprocessing to obtain multiple candidate RDDs.

[0059] It should be noted that due to the diversity and uncontrollability of user operations, abnormal data, such as erroneous data or interference data, may exist in the acquired user operation information, resulting in the candidate log record being an invalid record. In order to ensure the availability of user operation information and candidate log records, this embodiment adds a data preprocessing link. After obtaining the candidate log record from the target distributed message system based on the first time interval, each candidate log record is subjected to data preprocessing (including data cleaning and format conversion) so that Spark can be used to group the candidate log records after data preprocessing to obtain multiple candidate RDDs. Among them, data cleaning can specifically include data filtering and data repair. For example, when duplicate data appears, data filtering can be performed; when data is missing or abnormal, data filtering or data repair processing can be selected according to actual needs. When the data length is too long or the data format does not meet the current application requirements, the data format can be converted.

[0060] Through such a setting, this solution can ensure the validity and availability of log records by performing data preprocessing on log records, which helps to improve the accuracy of web page source analysis.

[0061] Figure 2 Schematic diagram of a web page source analysis method provided by Embodiment 1 of the present invention. Figure 2 As shown, the whole solution includes four parts: data transmission, data processing, data storage and traffic analysis. Specifically, firstly, jssdk is connected to the website, and the route jump behavior (i.e., web page jump behavior) is identified based on the page embedding point, and the corresponding user operation information is obtained after the route jump behavior is identified, and the user operation information is cleaned. Then, Spark is used to perform real-time streaming calculation and analysis on the user operation information after data cleaning based on the preset real-time analysis indicators, and the real-time analysis results are stored in the Hbase distributed database in the form of log records to realize data distributed storage. Then, the real-time analysis result data set is obtained from the Hbase distributed database with the hour as the time dimension, and the obtained data set is offline batch calculated and analyzed based on the preset offline analysis indicators, and the offline analysis results are stored in the Hbase distributed database in the form of log records. Finally, traffic analysis is performed based on the offline analysis results stored in the Hbase distributed database, so as to realize the source analysis of website pages based on traffic analysis.

[0062] Among them, web page source analysis can include forward analysis and reverse analysis. Specifically, forward analysis can be used to describe the reference sources of website web links (including direct sources, internal sources and external sources), and reverse analysis can be used to describe which website web links (such as web link 1, web link 2...web link n) reference the source link.

[0063] Embodiment 2

[0064] Figure 3 This is a flow chart of a web page source analysis method provided in the second embodiment of the present invention. This embodiment is optimized based on the above embodiment. The specific optimization is: after using Spark to group the candidate log records to obtain multiple candidate RDDs, it also includes: determining the RDD generation time of each candidate RDD, and generating a log mark number for each target log record in each candidate RDD; according to the RDD generation time and log mark number corresponding to the target log record, determining the identification information of the target log record.

[0065] like Figure 3 As shown, the method of this embodiment specifically includes the following steps:

[0066] S210, connecting JSSDK to the target website, and monitoring whether there is any web page jump behavior on the target website through JSSDK; wherein JSSDK implements monitoring based on user operations on the target website.

[0067] S220: If it exists, obtain the user operation information corresponding to the web page jump behavior, and generate a log record of the user operation information and store it in the target distributed messaging system.

[0068] S230, obtaining candidate log records from the target distributed messaging system based on the first time interval, and grouping the candidate log records using Spark to obtain multiple candidate RDDs.

[0069] S240, determining the RDD generation time of each candidate RDD, and generating a log mark sequence number for each target log record in each candidate RDD.

[0070] It should be noted that Spark streaming computing needs to ensure the uniqueness of each user's operation data, that is, to ensure the uniqueness of log records. The existing technology generally adopts a distributed ID generation method. The amount of log record data is uncertain, and a large number of log record IDs may need to be generated at the same time. The blocking ID generation method will affect the speed of streaming computing, and the generation efficiency is low. The complexity of generating IDs is also inconvenient to store as rowkeys.

[0071] In this embodiment, Spark's RDD native API is used to mark the serial numbers of the internal elements of the RDD, and the elements in the RDD are determined in combination with the uniqueness of the time object in the Dstream. Since the RDD generation time of each candidate RDD is specific, that is to say, each candidate RDD has a unique RDD generation time. Therefore, the RDD generation time can be used as the unique identifier of the corresponding candidate RDD, that is, each candidate RDD is uniquely characterized by the RDD generation time. Using Spark's RDD native API, log mark serial numbers can be generated for each target log record in each candidate RDD in order from small to large. For example, the log mark serial numbers corresponding to each target log record in each candidate RDD are 0, 1, 2,... At this time, the log mark serial numbers corresponding to each target log record in each candidate RDD are also unique.

[0072] S250, determining identification information of the target log record according to the RDD generation time and the log mark sequence number corresponding to the target log record.

[0073] It is understandable that the target log record can be determined from which candidate RDD the target log record comes according to the RDD generation time corresponding to the target log record, and the target log record can be further determined from which candidate RDD the target log record is in combination with the log mark number corresponding to the target log record, thereby clarifying the specific identity of the target log record. Therefore, the identification information of the target log record can be determined according to the RDD generation time and log mark number corresponding to the target log record, and the identification information can be used to uniquely characterize the target log record.

[0074] S260, performing real-time statistical analysis on the target log records in each candidate RDD to obtain a target RDD, and storing the target RDD in a target distributed storage system.

[0075] S270, obtaining multiple target RDDs from the target distributed storage system based on a second time interval, and performing offline integration processing on the target log records in the target RDD to obtain a web page source analysis result of the target website; wherein the second time interval is greater than the first time interval.

[0076] Among them, the specific implementation methods of S210-S230 and S260-S270 can refer to the relevant description in the above-mentioned embodiment 1, and will not be repeated here.

[0077] The technical solution of the embodiment of the present invention determines the identification information of the target log record according to the RDD generation time and log mark sequence number corresponding to the target log record, which can effectively overcome the adverse effect of the blocking ID generation method on the streaming computing speed and help improve the efficiency of log ID generation.

[0078] Embodiment 3

[0079] Figure 4 This is a schematic diagram of the structure of a web page source analysis device provided in the third embodiment of the present invention. The device can execute the web page source analysis method provided in any embodiment of the present invention and has the corresponding functional modules and beneficial effects of the execution method. Figure 4 As shown, the device comprises:

[0080] The web page jump behavior judgment module 310 is used to access the JSSDK in the target website and monitor whether the target website has a web page jump behavior through the JSSDK; wherein the JSSDK implements monitoring based on user operations of the target website;

[0081] The user operation information acquisition module 320 is used to acquire the user operation information corresponding to the web page jump behavior, if any, and generate a log record of the user operation information and store it in the target distributed messaging system;

[0082] A log record grouping module 330, configured to obtain candidate log records from the target distributed messaging system based on a first time interval, and group the candidate log records using Spark to obtain a plurality of candidate RDDs;

[0083] A log record real-time analysis module 340 is used to perform real-time statistical analysis on the target log record in each candidate RDD to obtain a target RDD, and store the target RDD in a target distributed storage system;

[0084] The log record offline analysis module 350 is used to obtain multiple target RDDs from the target distributed storage system based on a second time interval, and perform offline integration processing on the target log records in the target RDD to obtain the web page source analysis results of the target website; wherein the second time interval is greater than the first time interval.

[0085] Optionally, the web page jump behavior determination module 310 is used to:

[0086] If the JSSDK detects that the target website has a web page reloading behavior, it is determined that the target website has a web page jump behavior.

[0087] Optionally, the web page jump behavior determination module 310 is further configured to:

[0088] If the JSSDK monitors that a user clicks on a preset tag on the target website, it is determined that a web page jump occurs on the target website; wherein the preset tag refers to a pre-set tag that can achieve web page jump.

[0089] Optionally, the web page jump behavior determination module 310 is further configured to:

[0090] If the JSSDK detects that the target website has an internal route jump behavior, it is determined that the target website has a web page jump behavior.

[0091] Optionally, the web page jump behavior determination module 310 is further configured to:

[0092] Determine, through the popstate event listener in the target website, whether to change the webpage history record of the target website by calling a preset function through the JSSDK; wherein the preset function includes history.pushState and history.replaceState;

[0093] If yes, then obtaining the webpage URLs before and after the webpage history record is changed for comparative analysis to obtain a comparison result; wherein the comparison result is consistent or inconsistent;

[0094] If the comparison result is inconsistent, it is determined that the target website has an internal routing jump behavior.

[0095] Optionally, the device further includes: a log record identification information determination module, configured to:

[0096] After using Spark to group the candidate log records to obtain multiple candidate RDDs, determine the RDD generation time of each candidate RDD, and generate a log mark sequence number for each target log record in each candidate RDD;

[0097] Determine identification information of the target log record according to the RDD generation time and log mark sequence number corresponding to the target log record.

[0098] Optionally, the device further includes: a data preprocessing module, configured to:

[0099] After obtaining the candidate log record from the target distributed messaging system based on the first time interval, performing data preprocessing on the candidate log record; wherein the data preprocessing includes data cleaning and format conversion;

[0100] Accordingly, the log record grouping module 330 is used to:

[0101] Spark is used to group the candidate log records after data preprocessing to obtain multiple candidate RDDs.

[0102] A web page source analysis device provided in an embodiment of the present invention can execute a web page source analysis method provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.

[0103] Embodiment 4

[0104] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0105] like Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0106] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0107] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a web page source analysis method.

[0108] In some embodiments, the web page source analysis method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the web page source analysis method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the web page source analysis method in any other appropriate manner (e.g., by means of firmware).

[0109] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0110] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0111] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0112] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0113] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0114] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0115] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0116] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A web page source analysis method, characterized in that: The method comprises: Connecting JSSDK to the target website, and monitoring whether the target website has web page jump behavior through the JSSDK; wherein the JSSDK realizes monitoring based on user operations of the target website; If it exists, obtaining the user operation information corresponding to the web page jump behavior, and generating a log record of the user operation information and storing it in the target distributed messaging system; Obtain candidate log records from the target distributed messaging system based on a first time interval, and group the candidate log records using Spark to obtain a plurality of candidate RDDs; Perform real-time statistical analysis on the target log records in each candidate RDD to obtain a target RDD, and store the target RDD in a target distributed storage system; Based on a second time interval, multiple target RDDs are obtained from the target distributed storage system, and target log records in the target RDDs are offline integrated to obtain web page source analysis results of the target website; wherein the second time interval is greater than the first time interval.

2. The method according to claim 1, characterized in that Monitoring the target website through the JSSDK to see if there is a web page jump behavior includes: If the JSSDK detects that the target website has a web page reloading behavior, it is determined that the target website has a web page jump behavior.

3. The method according to claim 1, characterized in that Monitoring the target website through the JSSDK to see if there is a web page jump behavior includes: If the JSSDK monitors that a user clicks on a preset tag on the target website, it is determined that a web page jump occurs on the target website; wherein the preset tag refers to a pre-set tag that can achieve web page jump.

4. The method according to claim 1, characterized in that: Monitoring the target website through the JSSDK to see if there is a web page jump behavior includes: If the JSSDK detects that the target website has an internal route jump behavior, it is determined that the target website has a web page jump behavior.

5. The method according to claim 4, characterized in that The JSSDK monitors the target website for internal route jumps, including: Determine, through the popstate event listener in the target website, whether to change the webpage history record of the target website by calling a preset function through the JSSDK; wherein the preset function includes history.pushState and history.replaceState; If yes, then obtaining the webpage URLs before and after the webpage history record is changed for comparative analysis to obtain a comparison result; wherein the comparison result is consistent or inconsistent; If the comparison result is inconsistent, it is determined that the target website has an internal routing jump behavior.

6. The method according to any one of claims 1 to 5, characterized in that After using Spark to group the candidate log records to obtain multiple candidate RDDs, the method further includes: Determine the RDD generation time of each candidate RDD, and generate a log mark sequence number for each target log record in each candidate RDD; Determine identification information of the target log record according to the RDD generation time and log mark sequence number corresponding to the target log record.

7. The method according to any one of claims 1 to 5, characterized in that After acquiring the candidate log record from the target distributed messaging system based on the first time interval, the method further includes: Performing data preprocessing on the candidate log records; wherein the data preprocessing includes data cleaning and format conversion; Accordingly, Spark is used to group the candidate log records to obtain multiple candidate RDDs, including: Spark is used to group the candidate log records after data preprocessing to obtain multiple candidate RDDs.

8. A web page source analysis device, characterized in that: The device comprises: A web page jump behavior judgment module, used to access JSSDK in the target website, and monitor whether the target website has web page jump behavior through the JSSDK; wherein the JSSDK realizes monitoring based on user operations of the target website; A user operation information acquisition module, used to acquire the user operation information corresponding to the web page jump behavior, if any, and generate a log record of the user operation information and store it in the target distributed message system; A log record grouping module, configured to obtain candidate log records from the target distributed messaging system based on a first time interval, and group the candidate log records using Spark to obtain a plurality of candidate RDDs; A log record real-time analysis module is used to perform real-time statistical analysis on the target log record in each candidate RDD to obtain a target RDD, and store the target RDD in a target distributed storage system; A log record offline analysis module is used to obtain multiple target RDDs from the target distributed storage system based on a second time interval, and perform offline integration processing on the target log records in the target RDD to obtain the web page source analysis results of the target website; wherein the second time interval is greater than the first time interval.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the web page source analysis method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the web page source analysis method according to any one of claims 1 to 7 when executed.