Access Path Analysis Method, Device, Equipment and Computer Storage Medium
By using compressed bitmap data and hash algorithms to convert user identity identity formats and combining with Druid databases for real-time query, the problem of inefficient existing access path analysis methods is solved, real-time query of large data volumes and flexible path analysis of specific user groups is realized, and user experience is improved.
Patent Information
- Application Number
- CN202210675129.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-15
AI Technical Summary
The existing access path analysis methods are inefficient and cannot support query and calculation of large data volumes, and are not flexible enough to meet the needs of specific page path analysis.
The user identity identity is converted in format using compressed bitmap data structures (such as RoaringBitMap) and hashing algorithms, stored as compressed bitmap data, combined with the Druid database for real-time query, and supports access path analysis of specific user groups through custom extensions.
It improves the efficiency and flexibility of access path analysis, supports real-time query of large data volumes, saves storage space and improves user experience.
Smart Images

Figure CN115017392B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of data analysis, and in particular, to a method, device, equipment, and computer storage medium for access path analysis. Background Art
[0002] Analyzing the page access path situation of users and optimizing each page entry according to the analysis situation can improve the user experience. Currently, the mainstream page path analysis generally uses the following methods: writing code to calculate in memory, using relational databases such as mysql and postgrel to write sql for calculation, or using hive, sparksql, etc. for calculation.
[0003] The inventors found during the implementation of the embodiments of the present invention that the existing access path analysis has the problem of low analysis efficiency. Summary of the Invention
[0004] In view of the above problems, the embodiments of the present invention provide an access path analysis method for solving the problem of low access path analysis efficiency in the prior art.
[0005] According to one aspect of the embodiments of the present invention, an access path analysis method is provided, and the method includes:
[0006] Obtain a path query request; the path query request includes an access path parameter and a target user identifier;
[0007] Query in the index data according to the target user identifier to obtain a target user group corresponding to the target user identifier; at least one optional user identifier corresponding to each of a plurality of optional user groups is stored in the index data; the target user group is at least one of the optional user groups;
[0008] Process the page access data corresponding to the target user group according to the access path parameter to obtain a path analysis result corresponding to the target user identifier.
[0009] In an optional manner, the index data includes compressed bitmap data; the method further includes:
[0010] Obtain user grouping information; the user grouping information includes user identity identifiers of at least one optional user corresponding to each of the plurality of optional user groups; the optional user is a user within the optional user group;
[0011] Convert the format of the user identity identifier to obtain a converted identifier;
[0012] Compress and store the converted identifiers corresponding to each of the optional user groups to obtain the compressed bitmap data.
[0013] In an alternative manner, the method further includes:
[0014] Performing a hash calculation on the user identity identifier to obtain the converted identifier with a target number of bits; wherein, the target number of bits is determined according to the data structure of the compressed bitmap data.
[0015] In an alternative manner, the method further includes:
[0016] Performing an anti-temporal ordering process on the compressed bitmap data to obtain the user identifiers corresponding to each optional user group identifier;
[0017] Querying in the compressed bitmap data according to the target user identifier to obtain the target user group.
[0018] In an alternative manner, the method further includes:
[0019] Obtaining the original page access data;
[0020] Sorting the original page access data according to the access time to obtain the access page identifier time sequence;
[0021] Filtering the access page identifier time sequence according to the access user identifier to obtain the page access data respectively corresponding to each optional user identifier; the target user identifier is one of the optional user identifiers.
[0022] In an alternative manner, the access path parameter includes a start page, an access path direction, and a path depth; the path analysis result includes an access page chain; the method further includes:
[0023] Searching in the access page identifier time sequence corresponding to the target user identifier according to the start page to obtain a target page node;
[0024] Starting from the target page node, searching for the page identifiers with the number of the path depth in the access path direction of the access page identifier time sequence corresponding to the target user identifier to obtain the access page chain.
[0025] In an alternative manner, the method further includes:
[0026] Sequentially saving the found page identifiers in a tree data structure, the tree data structure includes multiple nodes, respectively establishing a corresponding node counter for each node in the access page chain, and updating the node counter according to the number of times the page identifier corresponding to the node appears;
[0027] Performing a visualization process on the tree data structure to obtain the access page chain.
[0028] According to another aspect of the embodiments of the present invention, there is provided an access path analysis device, including:
[0029] An acquisition module, configured to acquire a path query request; the path query request includes access path parameters and a target user identifier;
[0030] A query module, configured to query in index data according to the target user identifier to obtain a target user group corresponding to the target user identifier; at least one optional user identifier corresponding to each of a plurality of optional user groups is stored in the index data; the target user group is at least one of the optional user groups;
[0031] A processing module, configured to process page access data corresponding to the target user group according to the access path parameters to obtain a path analysis result corresponding to the target user identifier.
[0032] According to another aspect of the embodiments of the present invention, there is provided an access path analysis device, including:
[0033] A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0034] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations of the access path analysis method as described in any one of the foregoing.
[0035] According to still another aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, where at least one executable instruction is stored in the storage medium, and the executable instruction causes an access path analysis device to perform the following operations:
[0036] Acquire a path query request; the path query request includes access path parameters and a target user identifier;
[0037] Query in index data according to the target user identifier to obtain a target user group corresponding to the target user identifier; at least one optional user identifier corresponding to each of a plurality of optional user groups is stored in the index data; the target user group is at least one of the optional user groups;
[0038] Process page access data corresponding to the target user group according to the access path parameters to obtain a path analysis result corresponding to the target user identifier.
[0039] In an embodiment of the present invention, a path query request is obtained; the path query request includes an access path parameter and a target user identifier; the target user group identifier is obtained by querying in the index data according to the target user identifier; the index data stores multiple optional user groups and the corresponding optional user identifiers of each of the optional user groups; by representing the user grouping relationship with the user group identifier and the user identifier corresponding to the user group and storing it in the index data, the storage space is effectively saved and the search efficiency is improved. Finally, the page access data corresponding to the target user group is processed according to the access path parameter to obtain the path analysis result corresponding to the target user identifier, so as to be able to provide in real time the access path analysis result of the user group where the specific user is located, and improve the efficiency of access path analysis and the user experience.
[0040] The above description is only an overview of the technical solution of the embodiment of the present invention. In order to be able to understand the technical means of the embodiment of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiment of the present invention more obvious and understandable, the specific embodiments of the present invention are given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings are only used to illustrate the embodiments and are not considered as a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0042] Figure 1 shows a schematic flow chart of the access path analysis method provided by the embodiment of the present invention;
[0043] Figure 2 shows a schematic diagram of the display interface of the access path analysis method provided by another embodiment of the present invention;
[0044] Figure 3 shows a schematic diagram of the display interface of the access path analysis method provided by another embodiment of the present invention;
[0045] Figure 4 shows a schematic structural diagram of the access path analysis device provided by the embodiment of the present invention;
[0046] Figure 5 shows a schematic structural diagram of the access path analysis device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Hereinafter, the exemplary embodiments of the present invention will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein.
[0048] Before describing the embodiments of the present invention, the following related terms are described:
[0049] Apache Druid: A real-time analytics database designed for fast slicing and dicing analysis (OLAP queries) of large datasets. Druid is most commonly used as a database, and its main features mainly include: supporting real-time data import, which can be queried immediately after import, supporting high-concurrency import; interactive queries with sub-second response, supporting relatively high concurrency; long-term normal operation, etc. Druid is usually used for the front-end of application data analysis or as the back-end of high-concurrency APIs that require fast aggregation.
[0050] RoaringBitMap: A traditional bitmap, which is a very fast data structure, but the disadvantage is that it consumes too much memory. In order to reduce memory consumption, compressed bitmaps are usually used. RoaringBitMap is a compressed bitmap with excellent performance.
[0051] Murmurhash3: A non-cryptographic hash function suitable for general hash retrieval operations. The current version is MurmurHash3, which can generate 32-bit or 128-bit hash values.
[0052] Hdfs: Hadoop Distribute File System, a master-slave (Master / Slave) structure model. An HDFS cluster consists of a NameNode (name node) and several DataNodes (data nodes). The NameNode serves as the main server, managing the namespace of the file system and the client's access operations to files; the DataNodes in the cluster manage the stored data.
[0053] Before describing the embodiments of the present invention, the existing access path analysis methods and their existing problems are introduced.
[0054] Currently, the mainstream page path analysis generally uses the following methods: 1. Write code to calculate in memory, which is only suitable for simple logic and small data volumes; 2. Use relational databases such as mysql and postgrel to write SQL for calculation. The data volume should not be too large, and the path level should not exceed two levels; 3. Use calculations such as hive and sparksql, which are suitable for large data volumes. The disadvantage is that they are only suitable for offline queries and not suitable for real-time queries in OLAP scenarios.
[0055] Therefore, the current page path analysis mainly faces the following two problems: it does not support a large amount of data and is not flexible enough. With the advent of the big data era, the exponential growth of data volume has also brought difficulties to data analysis. Using traditional databases, it is impossible to support queries and calculations of large amounts of data. Using big data databases such as Hive, ad-hoc queries cannot be achieved. Moreover, most of the existing technologies only support page path analysis with predefined conditions, such as subsequent conversions after adding a product to the shopping cart, click conversion analysis of the recommended banner page, etc. If it is necessary to support the path conversion of a specific page and specify the path conversion with a good number of steps, custom development is required, which is not flexible enough. This leads to problems of low query efficiency and poor query experience in the existing access path analysis methods.
[0056] Figure 1 The flowchart of the access path analysis method provided by an embodiment of the present invention is shown. This method is executed by a computer processing device. The computer processing device may include a mobile phone, a laptop computer, etc. As Figure 1 shown, the method includes the following steps:
[0057] Step 10: Obtain a path query request; the path query request includes access path parameters and a target user identifier.
[0058] In an embodiment of the present invention, the path query request is used to request to query the path information of the accessed pages of one or more users. For example, the subsequent pages after a user adds a product to the shopping cart, the subsequent pages after a user clicks on the recommended banner page, etc. The path information includes the information of at least one page, and the page information includes at least information such as entry and exit times, upstream and downstream pages, and in-page operations. The path query request may be sent by a person who needs to analyze user behavior, such as an operator.
[0059] The access path parameters are used to characterize the structure-related parameters of the path to be constructed, and may include path depth, path direction, the starting page of the path, and page access time, etc. According to the access path parameters and the user's historical page access data, an access path chain with a corresponding structure can be constructed.
[0060] Optionally, the access path parameters may further include filtering conditions, and the filtering conditions are used to screen the historical page access data to obtain data related to path construction.
[0061] The user identifier is used to specifically characterize the user, and may be a user identity identifier, a device identifier, an IP address, etc. Further, the user may belong to one or more user groups, and the user groups may be pre-divided according to the characteristics of the user or operational requirements. For example, users can be grouped according to regions, preferences, activity levels, or specific behaviors and events. Among them, specific behaviors and events may include whether to recharge, whether to follow, etc.
[0062] Step 20: Query in the index data according to the target user identifier to obtain the target user group corresponding to the target user identifier; at least one optional user identifier corresponding to each of multiple optional user groups is stored in the index data; the target user group is at least one of the optional user groups.
[0063] In an embodiment of the present invention, an optional user group corresponds to a group identifier specifically characterizing the optional user group. An optional user group includes at least one user, and the user identifiers of the users included in the optional user group are the user identifiers corresponding to the optional group identifier. Among them, the user identifier may include a user ID, a device ID, a network IP address, etc.
[0064] To save the storage cost of the index data and improve the search efficiency, the data structure of the index data may be a bitmap, that is, a bitmap. Optionally, the bitmap data may be compressed bitmap data.
[0065] Considering that the existing page path analysis does not support statistical analysis for specific user groups, when it is necessary to perform path analysis on the accessed pages for the user group to which a specific user belongs, it is necessary to separately look up tables for all users in the user group where the specific user is located, and its search efficiency is low while the storage cost is high. In the embodiment of the present invention, by setting up a sub-group identifier for the user group and pre-storing the user identifiers corresponding to the sub-group identifier in the bitmap data, and the search speed of the bitmap data is faster, thereby the query performance can be improved.
[0066] In an embodiment of the present invention, the index data includes compressed bitmap data; with the high compression of RoaringBitMap, the sub-group information of millions or tens of millions of users can be easily loaded into the memory and serialized to Hdfs. Therefore, when it is necessary to obtain the user identifier, the data is obtained from Hdfs and deserialized, saving memory and also improving the access speed.
[0067] Step 20 further includes: Step 201: Obtain user sub-group information; the user sub-group information includes the user identity identifiers of at least one optional user corresponding to each of the multiple optional user groups; the optional user is a user within the optional user group.
[0068] In an embodiment of the present invention, the basis for user sub-grouping may be operational requirements or user characteristic dimensions, such as activity, preference, and region. For example, users who are active within 7 days, users in a certain region, and users who follow a certain singer can each form an optional user group.
[0069] Step 202: Perform format conversion on the user identity identifier to obtain a converted identifier.
[0070] In an embodiment of the present invention, considering that the data structure of the compressed bitmap has specific format requirements for data, it is necessary to first perform corresponding format conversion on the user identity identifier before it can be stored in the compressed bitmap data. Specifically, considering that the user identity identifier generally uses the device number of the user, and the user device number is a very long string, so through hash calculation, the string can be converted into a corresponding numerical type, such as a 32-bit or 64-bit long type. Among them, a 32-bit id can accommodate 2 to the 32nd power of numerical values. A 64-bit id can accommodate 2 to the 64th power of numerical values. According to different precision requirements, an id with a suitable length is selected, effectively saving memory space and improving the query efficiency.
[0071] In an embodiment of the present invention, step 202 further includes:
[0072] Step 2021: Perform hash calculation on the user identity identifier to obtain the converted identifier with the target number of bits; wherein, the target number of bits is determined according to the data structure of the compressed bitmap data.
[0073] In an embodiment of the present invention, the hash calculation can adopt the hash algorithm of Murmurhash3, so as to convert the device identifier in string format of the user into a numerical type in 64-bit long format, so that the screening of clustering can be combined with the 64-bit bitmap function of RoaringBitMap.
[0074] Step 203: Compress and store the converted identifiers corresponding to each optional user group to obtain the compressed bitmap data.
[0075] In an embodiment of the present invention, the converted identifiers corresponding to the optional user groups are stored in the compressed bitmap data structure with the corresponding number of bits to obtain the compressed bitmap data corresponding to the optional user groups.
[0076] For example, the conversion of the user id of ios before and after is as follows:
[0077] Id before conversion: 5b19d5afa15c4084907e2ccd7a1759fc
[0078] 32-bit id after conversion: 200030963
[0079] 64-bit id after conversion: -5102500264101004557
[0080] If there are one million such IDs, directly storing them in a file would require approximately 36M of space, while storing the converted IDs in a 32-bit Roaring Bitmap only requires about 2.4M. If a high-precision 64-bit Roaring Bitmap is needed, only about 26M of space is required. When concurrently querying the user access path chains with clustering, using compression can save more cluster resources and improve the query speed. As the number of clustered users increases, the effect of saving space becomes more obvious.
[0081] In an embodiment of the present invention, step 20 further includes:
[0082] Step 204: Perform anti-temporal processing on the compressed bitmap data to obtain the user identifiers corresponding to each optional user group identifier.
[0083] In an embodiment of the present invention, the compressed bitmap data is calculated according to a daily scheduled task and stored on Hdfs. When it is necessary to query which users correspond to the target user group identifier, the user clustering data on Hdfs can be anti-temporally converted into a RoaringBitmap, and then corresponding queries can be performed in the RoaringBitmap.
[0084] Specifically, first obtain the file input stream of Hdfs, and call the deserialize anti-temporalization (FSDataInputStream in) method of RoaringBitmap, and pass the input stream as a parameter into the method, so that RoaringBitmap loads the file according to the passed-in file stream, anti-temporally processes the compressed bitmap data file, and loads it into memory, thereby improving the query efficiency.
[0085] Step 205: Query in the compressed bitmap data according to the target user identifier to obtain the target user group.
[0086] In an embodiment of the present invention, the user group corresponding to the user group identifier corresponding to the target user identifier in the bitmap data is determined as the target user group. For example, when the target user identifier is 001, the user group identifiers corresponding to the optional user identifier 001 queried in the bitmap data are 100, 101, 102, 103, 105, 109, and 120, then the target user group is the user groups corresponding to the user group identifiers 100, 101, 102, 103, 105, 109, and 120 respectively.
[0087] Step 30: Process the page access data corresponding to the target user group according to the access path parameter to obtain the path analysis result corresponding to the target user identifier.
[0088] In an embodiment of the present invention, the page access data corresponding to the target user group refers to the relevant data of the pages accessed by all users within the target user group during a preset historical period, such as page identifiers, access times, adjacent page information, etc. The page access data can be collected in advance.
[0089] In order to further improve the query efficiency, the page access data can also be preprocessed first. Considering that the page access chain mainly represents the temporal association relationship of the accessed pages, that is, the page nodes in the page access chain are connected according to the access order to form the access chain. Therefore, the page access timing information corresponding to each user can be obtained by processing the page access data. Thus, when constructing the page access chain corresponding to the user, it can be directly extracted according to the page access timing information. Therefore, in still another embodiment of the present invention, before step 30, it further includes:
[0090] Step 301: Obtain the original page access data.
[0091] In an embodiment of the present invention, the original page access data includes information such as user identifiers, session identifiers, access times, and page identifiers. The collection method of the original page access data can be to collect the user access page id, session id (session identifier), and basic device information, etc. by performing data tracking on the APP, collecting data without data tracking, or using the user selection function.
[0092] Step 302: Sort the original page access data according to the access time to obtain the access page identifier timing sequence.
[0093] In an embodiment of the present invention, the access page identifier timing sequence includes the identifiers of multiple accessed pages arranged in the order of access time. First, the identifiers of each accessed page and the corresponding access times are extracted from the original page access data, and sorted according to the access time to obtain the access page identifier timing sequence.
[0094] Step 303: Filter the access page identifier timing sequence according to the access user identifier to obtain the page access data corresponding to each optional user identifier.
[0095] In an embodiment of the present invention, using the access user as the primary key, the access page identifiers under this primary key are aggregated to obtain the page access data corresponding to the optional user identifier. Optionally, the minimum unit of aggregation can be a session or a time unit, such as a day, a week, etc.
[0096] In one embodiment of the present invention, the access path parameter includes a start page, an access path direction, and a path depth; wherein, the start page refers to the starting page of the access path, which can generally be a startup page, a welcome page, a login page, etc. The access path direction refers to forward or backward from the start page, where the front and back refer to the sequence of page access times. The unit of the path depth can be the number of pages, that is, the constructed access path needs to include information of multiple accessed pages of the user. The path analysis result includes an access page chain, which is a chain formed by connecting multiple accessed pages according to the access time and the transition relationship. The access page chain includes multiple nodes, and one node corresponds to a historically accessed page. Among them, the transition relationship includes accessing page B after accessing page A, that is, it is regarded as transitioning from node A to node B.
[0097] Step 30 further includes: Step 310: Search in the access page identifier time sequence corresponding to the target user identifier according to the start page to obtain a target page node.
[0098] In one embodiment of the present invention, a page whose page identifier in the access page identifier time sequence corresponding to the target user identifier is the same as the identifier of the start page is determined as the target page node. The access page chain can be formed in the form of a tree, and the tree includes at least one node, and the target page node is the root node of the tree.
[0099] Step 311: Starting from the target page node, search for the page identifiers of the number of the path depth in the access path direction of the access page identifier time sequence corresponding to the target user identifier to obtain the access page chain.
[0100] In one embodiment of the present invention, search in the access page identifier time sequence according to the path depth and the access path direction, and sequentially add the found access page identifiers as nodes to the access page chain.
[0101] Further, considering that the user may return to the previous page and repeat access when accessing multiple pages, that is, a node may be accessed multiple times. Therefore, in one embodiment of the present invention, step 311 further includes:
[0102] Step 3111: Sequentially save the found page identifiers in a tree data structure. The tree data structure includes multiple nodes. For each node in the access page chain, a corresponding node counter is established, and the node counter is updated according to the number of times the page identifier corresponding to the node appears.
[0103] In one embodiment of the present invention, the method for updating the tree structure may be as follows: for the subsequent page of the currently found page, if the subsequent page is not in the access page identifier time sequence, a new node is added to the tree data structure; if it exists, the counter of the node corresponding to the subsequent page is incremented by one.
[0104] Step 3112: Perform visualization processing on the tree data structure to obtain the access page chain.
[0105] In one embodiment of the present invention, the access page chain obtained by performing visualization processing on the tree data structure may be Figure 2 and Figure 3 as shown on the right side of.
[0106] As Figure 2 and Figure 3 shown, each node in the access page chain includes the corresponding page event and the percentage of users participating in this event in the current user group. For example, for Figure 2 the input starting page is APP startup, the path direction is backward, the path level is 3, and the target users are all users (that is, all users are included in the current user group). In the output access page chain, 100% of the users start from APP startup. Among them, 65.18% of the users visit the music home page, 27.73% of the users visit the play test page. Among the users who visit the music home page, 60.32% of the users exit from the APP, and 26.66% of the users visit the play test page.
[0107] In yet another embodiment of the present invention, it is also possible to directly perform access path analysis on all users corresponding to the group according to the user group identifier. The specific process may be as follows:
[0108] First, a query interface as shown on the left side of Figure 2 and Figure 3 is displayed on the front end. The query conditions are received through the query interface. The query conditions at least include the starting page, access path direction, path depth, time range, user grouping, and filtering conditions.
[0109] Then, according to the conditions selected on the front end, the server concatenates the starting page, access path direction, path depth, time range, user grouping, and filtering conditions as parameters to form a query JSON for Druid and sends it to the broker of Druid for query. The broker selects the custom extension of the present invention embodiment for query according to the name of the aggregator.
[0110] Among them, the custom extension for Druid includes the following:
[0111] First, load the user page access data that meets the filtering conditions within the specified time range, and then determine whether the user has selected a user group. If the user selects a user group on the interface, the group ID of the selected group will be used as a query parameter and sent to the background. Therefore, in this solution, it is possible to determine whether the user has selected a group by checking whether there is a group ID in the query parameter.
[0112] User group data is the group division of users by operations based on user characteristics, such as region, activity level, and specific behaviors. The group data contains the device information of users. Through the group data, the performance of specific user groups on the path analysis module can be viewed.
[0113] The advantage of supporting custom group division is that operations personnel can comprehensively understand different groups of people. When accessing the same page, they can then improve the page layout distribution based on the subsequent conversions on each page. They can also perform customized page distributions for different groups, and preferentially expose the page entrances with higher access frequencies in the access path. If the user does not select a user group, there is no need to load the user group data, and the access paths of all users can be directly calculated, saving the time and resource consumption of loading user groups. Only when a user group is selected, the corresponding data of the selected group is loaded into the memory for subsequent calculations. Resources are reasonably allocated to maximize the cluster computing power.
[0114] If a user group is selected, the data of the user group is deserialized from Hdfs into a RoaringBitmap.
[0115] Specifically, first obtain the file input stream of Hdfs, call the deserialize(FSData Input Stream in) method of Roaring Bitmap, and pass the input stream as a parameter into the method. Roaring Bitmap loads the file according to the passed-in file stream, deserializes the compressed group file, and loads it into the memory, thus saving the time consumption of directly using druid to query group data in the prior art.
[0116] The user segmentation data is pre-computed through a daily scheduled task, serialized into a bitmap format file, and saved on Hdfs. After the data is loaded, first check whether the user is in the selected segmentation. If not, discard this piece of data. If in the selected user segmentation, continue with the indexing of the starting page. Through sequential indexing of the page ID list, if the starting page is indexed, then according to the direction configured in the query, look backward or forward by the number of steps according to the depth configured in the path, and save the page IDs in a tree data structure for caching. For the subsequent pages of a page, if they do not exist, add a new node. If they exist, increment the counter of the existing node by 1. After the query is completed, return the data in the tree data structure in JSON format.
[0117] Finally, the page performs data visualization based on the returned JSON data.
[0118] The embodiments of the present invention can be applied to the operation analysis platform, user single graph, and path analysis module. With the help of this module, operators can more intuitively analyze the page access path of users, so as to optimize each page entry. The query speed is doubled compared with the prior art, saving space and reducing the query pressure. It can be extended later and applied to more query analysis modules that require user segmentation screening.
[0119] The embodiments of the present invention implement a custom Druid query extension for the user access page path chain by writing code. It cleverly uses the Hash algorithm to convert the segmented user ID into a numerical value, and combines the compressed bitmap technology, effectively reducing the memory size occupied by the user segmentation data, and having a higher efficiency in screening segmented users. The user access page path chain query implemented by this solution has a significant improvement in query efficiency. It also supports real-time queries with custom starting access pages, query depths, and access directions, improving the user's query experience. It is convenient for operations to customize different strategies for data analysis at any time, and the analysis results can be saved and downloaded, thus improving the efficiency of access path analysis and the user experience.
[0120] The access path analysis method provided by the embodiments of the present invention obtains a path query request; the path query request includes access path parameters and a target user identifier; queries in the index data according to the target user identifier to obtain a target user group identifier; the index data stores multiple optional user groups and the optional user identifiers respectively corresponding to each of the optional user groups; by representing the user grouping relationship with the user group identifier and the user identifier corresponding to the user group and storing it in the index data, the storage space is effectively saved and the search efficiency is improved. Finally, the page access data corresponding to the target user group is processed according to the access path parameters to obtain the path analysis result corresponding to the target user identifier, so as to be able to provide the access path analysis result of the user group where the specific user is located in real time, and improve the efficiency of access path analysis and the user experience.
[0121] Figure 4 FIG. shows a schematic structural diagram of an access path analysis device provided by an embodiment of the present invention. As Figure 4 shown, the device 40 includes: an obtaining module 401, a query module 402, and a processing module 403.
[0122] Among them, the obtaining module 401 is configured to obtain a path query request; the path query request includes access path parameters and a target user identifier;
[0123] The query module 402 is configured to query in the index data according to the target user identifier to obtain the target user group corresponding to the target user identifier; the index data stores at least one optional user identifier respectively corresponding to multiple optional user groups; the target user group is at least one of the optional user groups;
[0124] The processing module 403 is configured to process the page access data corresponding to the target user group according to the access path parameters to obtain the path analysis result corresponding to the target user identifier.
[0125] The operation process executed by the access path analysis device provided by the embodiments of the present invention is substantially the same as the signing method and will not be elaborated here.
[0126] The access path analysis device provided by the embodiment of the present invention obtains a path query request; the path query request includes access path parameters and a target user identifier; queries the index data according to the target user identifier to obtain a target user group identifier; the index data stores multiple optional user groups and the corresponding optional user identifiers of each optional user group; by representing the user grouping relationship with the user group identifier and the user identifier corresponding to the user group and storing it in the index data, the storage space is effectively saved and the search efficiency is improved. Finally, the page access data corresponding to the target user group is processed according to the access path parameters to obtain the path analysis result corresponding to the target user identifier, so as to be able to provide the access path analysis result of the user group where the specific user is located in real time, and improve the efficiency of access path analysis and the user experience.
[0127] Figure 3 FIG. shows a schematic structural diagram of the access path analysis device provided by the embodiment of the present invention. The specific implementation of the access path analysis device is not limited in the specific embodiment of the present invention.
[0128] As Figure 3 shown, the access path analysis device may include: a processor 402, a communication interface 404, a memory 406, and a communication bus 408.
[0129] Among them: the processor 402, the communication interface 404, and the memory 406 communicate with each other through the communication bus 408. The communication interface 404 is used to communicate with network elements of other devices such as clients or other servers. The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the above-mentioned embodiment of the access path analysis method.
[0130] Specifically, the program 410 may include program code, and the program code includes computer executable instructions.
[0131] The processor 402 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiment of the present invention. One or more processors included in the access path analysis device may be of the same type of processor, such as one or more CPUs; or different types of processors, such as one or more CPUs and one or more ASICs.
[0132] A memory 406 for storing a program 410. The memory 406 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0133] Specifically, the program 410 can be called by the processor 402 to cause the access path analysis device to perform the following operations:
[0134] Obtain a path query request; the path query request includes access path parameters and a target user identifier;
[0135] Query in the index data according to the target user identifier to obtain a target user group corresponding to the target user identifier; the index data stores at least one optional user identifier corresponding to each of multiple optional user groups; the target user group is at least one of the optional user groups;
[0136] Process the page access data corresponding to the target user group according to the access path parameters to obtain a path analysis result corresponding to the target user identifier.
[0137] The operation process and signing method performed by the access path analysis device provided by the embodiments of the present invention are reasonably substantially the same and will not be elaborated here.
[0138] The access path analysis device provided by the embodiments of the present invention obtains a path query request; the path query request includes access path parameters and a target user identifier; queries in the index data according to the target user identifier to obtain a target user group identifier; the index data stores multiple optional user groups and the optional user identifiers corresponding to each of the optional user groups; by representing the user grouping relationship with the user group identifier and the user identifier corresponding to the user group and storing it in the index data, the storage space is effectively saved and the search efficiency is improved. Finally, the page access data corresponding to the target user group is processed according to the access path parameters to obtain a path analysis result corresponding to the target user identifier, so as to be able to provide in real time the access path analysis result of the user group where a specific user is located, and improve the efficiency of access path analysis and the user experience.
[0139] The embodiments of the present invention provide a computer-readable storage medium, and the storage medium stores at least one executable instruction. When the executable instruction runs on the access path analysis device, it causes the access path analysis device to execute the access path analysis method in any of the above method embodiments.
[0140] The executable instruction can specifically be used to cause the access path analysis device to perform the following operations:
[0141] Obtain a path query request; the path query request includes an access path parameter and a target user identifier;
[0142] Query in the index data according to the target user identifier to obtain a target user group corresponding to the target user identifier; the index data stores at least one optional user identifier corresponding to each of multiple optional user groups; the target user group is at least one of the optional user groups;
[0143] Process the page access data corresponding to the target user group according to the access path parameter to obtain a path analysis result corresponding to the target user identifier.
[0144] The operation process executed by the executable instructions stored in the computer-readable storage medium provided by the embodiments of the present invention is reasonably substantially the same as the signing method, and will not be elaborated here.
[0145] The executable instructions stored in the computer-readable storage medium provided by the embodiments of the present invention obtain a path query request; the path query request includes an access path parameter and a target user identifier; query in the index data according to the target user identifier to obtain a target user group identifier; the index data stores multiple optional user groups and the optional user identifiers corresponding to each of the optional user groups; by representing the user grouping relationship with the user group identifier and the user identifier corresponding to the user group and storing it in the index data, the storage space is effectively saved and the search efficiency is improved. Finally, process the page access data corresponding to the target user group according to the access path parameter to obtain a path analysis result corresponding to the target user identifier, so as to be able to provide in real time the access path analysis result of the user group where a specific user is located, and improve the efficiency of access path analysis and the user experience.
[0146] The embodiments of the present invention provide an access path analysis device for executing the above access path analysis method.
[0147] The embodiments of the present invention provide a computer program, which can be called by a processor to enable an access path analysis device to execute the access path analysis method in any of the above method embodiments.
[0148] The embodiments of the present invention provide a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions run on a computer, the computer executes the access path analysis method in any of the above method embodiments.
[0149] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems can also be used in conjunction with the teachings based herein. The structure required to construct such systems will be apparent from the above description. In addition, embodiments of the present invention are not directed to any particular programming language. It should be understood that the teachings of the present invention described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.
[0150] In the specification provided herein, a number of specific details are set forth. However, it will be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure an understanding of the present specification.
[0151] Similarly, it should be understood that, in order to streamline the present invention and assist in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of the present invention, the various features of the embodiments of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed methods should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim.
[0152] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0153] It should be noted that the above embodiments are illustrative of the present invention and not restrictive thereof, and alternative embodiments can be designed by those skilled in the art without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specifically stated, should not be construed as limiting the order of execution.
Claims
1. A method for analyzing access paths, characterized in that, The method includes: Obtaining a path query request; the path query request includes an access path parameter and a target user identifier; Querying in the index data according to the target user identifier to obtain a target user group corresponding to the target user identifier; at least one optional user identifier corresponding to each of a plurality of optional user groups is stored in the index data; the target user group is at least one of the optional user groups; Obtaining original page access data; sorting the original page access data according to the access time to obtain an access page identifier time sequence; filtering the access page identifier time sequence according to the access user identifier to obtain the page access data corresponding to each optional user identifier; Processing the page access data corresponding to the target user group according to the access path parameter to obtain a path analysis result corresponding to the target user identifier; wherein, the access path parameter includes a start page, an access path direction, and a path depth; the path analysis result includes an access page chain; the processing the page access data corresponding to the target user identifier according to the access path parameter to obtain the path analysis result corresponding to the target user identifier includes: searching in the access page identifier time sequence corresponding to the target user identifier according to the start page to obtain a target page node; starting from the target page node, searching for the number of page identifiers of the path depth in the access path direction of the access page identifier time sequence corresponding to the target user identifier to obtain the access page chain.
2. The method according to claim 1, wherein The index data includes compressed bitmap data; Before querying in the index data according to the target user identifier to obtain the target user group corresponding to the target user identifier, it includes: Obtaining user grouping information; the user grouping information includes the user identity identifiers of at least one optional user corresponding to each of the plurality of optional user groups; the optional user is a user within the optional user group; Performing format conversion on the user identity identifier to obtain a converted identifier; Compressively storing the converted identifiers corresponding to each of the optional user groups to obtain the compressed bitmap data.
3. The method according to claim 2, wherein The performing format conversion on the user identity identifier to obtain a converted identifier includes: Performing hash calculation on the user identity identifier to obtain the converted identifier with a target number of bits; wherein, the target number of bits is determined according to the data structure of the compressed bitmap data.
4. The method according to claim 2, characterized in that The querying in the index data according to the target user identifier to obtain the target user group corresponding to the target user identifier includes: Performing anti-timing processing on the compressed bitmap data to obtain the user identifiers corresponding to each optional user group identifier; Querying in the compressed bitmap data according to the target user identifier to obtain the target user group.
5. The method according to claim 1, wherein The searching for the number of page identifiers of the path depth in the access path direction of the access page identifier time sequence corresponding to the target user identifier starting from the target page node to obtain the access page chain includes: The page identifiers found are sequentially saved in a tree data structure, where the tree data structure includes multiple nodes. A corresponding node counter is established for each node in the access page chain, and the node counter is updated according to the number of times the page identifier corresponding to the node appears. Visualize the tree data structure to obtain the access page chain.
6. An access path analysis device, characterized in that, The device includes: An acquisition module, configured to acquire a path query request; the path query request includes an access path parameter and a target user identifier. A query module, configured to query in the index data according to the target user identifier to obtain a target user group corresponding to the target user identifier; the index data stores at least one optional user identifier corresponding to each of multiple optional user groups; the target user group is at least one of the optional user groups. A processing module, configured to obtain raw page access data; sort the raw page access data according to the access time to obtain an access page identifier time sequence; filter the access page identifier time sequence according to the access user identifier to obtain the page access data corresponding to each optional user identifier; process the page access data corresponding to the target user group according to the access path parameter to obtain a path analysis result corresponding to the target user identifier; where the access path parameter includes a start page, an access path direction, and a path depth; the path analysis result includes an access page chain; the processing the page access data corresponding to the target user identifier according to the access path parameter to obtain a path analysis result corresponding to the target user identifier includes: searching in the access page identifier time sequence corresponding to the target user identifier according to the start page to obtain a target page node; starting from the target page node, searching for the page identifier of the path depth number in the access path direction of the access page identifier time sequence corresponding to the target user identifier to obtain the access page chain.
7. An access path analysis device, characterized in that, It includes: A processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations of the access path analysis method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, At least one executable instruction is stored in the storage medium. When the executable instruction runs on the access path analysis device, it causes the access path analysis device to execute the operations of the access path analysis method according to any one of claims 1-5.
Citation Information
Patent Citations
Data query method and device based on access path, storage medium and processor
CN111125155A
Social relation chain establishing method and device, electronic equipment and computer readable storage medium
CN112149002A