Information mining method, device and equipment for applet
By filtering and normalizing the parameter of the mini program URL, the parameter mode of the target URL is quickly found, which solves the problem of low efficiency in mini program information mining and realizes efficient and accurate data analysis and mining.
Patent Information
- Application Number
- CN202510387716.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
The information mining efficiency of existing technology small and medium-sized programs is low, and it is difficult to efficiently obtain valuable business information when facing massive user URLs.
By obtaining the URL of the user accessing the applet, filtering parameter key-value pairs based on preset filtering rules, generating parameter modes, and normalizing the URL to obtain an accurate target page for information mining.
It improves the efficiency of information mining, from monthly to hourly level, reduces redundant parameter processing, improves the accuracy of data analysis, and supports the scalability of different mini program scenarios.
Smart Images

Figure CN120336650A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of Internet technologies, and in particular, to a method, apparatus, and device for mining details pages of mini programs. Background Art
[0002] With the development of computer and Internet technologies, many services can be carried out online through mini programs built into applications. For example, users can conduct services such as commodity purchases, hotel reservations, and fee payments in mini programs.
[0003] For an application, corresponding information can be mined through the behavior data of users within a mini program, so as to provide better services for users.
[0004] In traditional solutions, the specific information accessed by users can be obtained through a uniform resource locator (URL) for information mining. However, with the increase in the volume of mini programs, there are a vast number of user URLs, and a large number of them have complex structures.
[0005] Based on this, a more efficient information mining solution is required for mini programs. Summary of the Invention
[0006] One or more embodiments of this specification provide a method, apparatus, device, and storage medium for mining information of mini programs to solve the following technical problem: A more efficient information mining solution is required for mini programs.
[0007] To solve the above technical problem, one or more embodiments of this specification are implemented as follows:
[0008] A method provided by one or more embodiments of this specification includes:
[0009] Obtain a first URL corresponding to a user accessing a specified mini program;
[0010] Based on a preset filtering rule, filter the parameter key-value pairs included in the first URL to generate a corresponding parameter pattern;
[0011] Normalize the first URL according to the content included in the parameter pattern to obtain a second URL;
[0012] Access a corresponding first target page through the second URL for information mining.
[0013] An information mining apparatus for mini programs provided by one or more embodiments of this specification includes:
[0014] A URL acquisition module that acquires a first URL corresponding to a user accessing a specified applet;
[0015] A parameter mode generation module that filters the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter mode;
[0016] A URL normalization module that normalizes the first URL according to the content included in the parameter mode to obtain a second URL;
[0017] An information mining module that accesses a corresponding first target page through the second URL for information mining.
[0018] An information mining device for applets provided by one or more embodiments of this specification includes:
[0019] At least one processor; and,
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:
[0022] Acquire a first URL corresponding to a user accessing a specified applet;
[0023] Filter the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter mode;
[0024] Normalize the first URL according to the content included in the parameter mode to obtain a second URL;
[0025] Access a corresponding first target page through the second URL for information mining.
[0026] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to:
[0027] Acquire a first URL corresponding to a user accessing a specified applet;
[0028] Filter the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter mode;
[0029] Normalize the first URL according to the content included in the parameter mode to obtain a second URL;
[0030] Access the corresponding first target page through the second URL for information mining.
[0031] The above at least one technical solution adopted by one or more embodiments of this specification can achieve the following beneficial effects:
[0032] Through parameter filtering and URL normalization, quickly find the parameter pattern of the target URL, reduce the processing of redundant parameters, lower the computational complexity, speed up the analysis of a large number of URLs, and retrieve all the required target URLs accessed by the user at one time, improving the timeliness from monthly level to hourly level.
[0033] Retain key parameters through parameter pattern filtering, eliminate irrelevant interferences, and the data that meets the parameter pattern is basically accurate URLs, thereby improving the accuracy of data analysis and mining results.
[0034] The filtering rules can be dynamically adjusted to support the requirements of different applet scenarios and have strong scalability. Brief Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is a schematic flowchart of a method for information mining for applets provided by one or more embodiments of this specification;
[0037] Figure 2 It is a schematic diagram of a method for information mining for applets in an application scenario provided by one or more embodiments of this specification;
[0038] Figure 3 It is a schematic diagram of the generation of a fourth target page in an application scenario provided by one or more embodiments of this specification;
[0039] Figure 4 It is a schematic structural diagram of an information mining device for applets provided by one or more embodiments of this specification;
[0040] Figure 5 It is a schematic structural diagram of an information mining device for applets provided by one or more embodiments of this specification. Detailed Embodiments
[0041] An information mining method, apparatus, device, and storage medium for applets are provided in the embodiments of this specification.
[0042] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0043] Figure 1 A flowchart of an information mining method for applets provided in one or more embodiments of this specification. This method can be applied to different business fields, such as the applet business field, the Internet finance business field, the e-commerce business field, the instant messaging business field, the game business field, the official business field, etc. This process can be executed by a computing device in the corresponding field (such as the application program where the applet is located or the corresponding server). Some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.
[0044] Figure 1 The process in
[0045] S102: Obtain a first URL corresponding to a user's access to a specified applet.
[0046] For an application program, due to the characteristic that an applet can be used without installation, a large number of applets may be integrated inside many application programs. For example, for some payment platforms and social platforms, which have a large user base themselves, integrating applets can increase user convenience and at the same time increase user stickiness.
[0047] Figure 2 For one or more embodiments of this specification, a schematic diagram of an information mining method for applets in an application scenario is provided. The following is described in combination with Figure 1 and Figure 2 When a user accesses an applet inside an application program, the access address of the user will be recorded in the access log of the application program. The access address is usually represented in the form of a URL, and the first URL can be obtained by intercepting the log data within a corresponding time period. For example, intercept the log data within 30 days.
[0048] As the number of users of the application increases and the number of integrated mini-programs increases, the number of original URLs in the access logs will reach a huge amount. For example, for some applications that integrate millions of mini-programs internally, there may be tens of billions of original URLs in their access logs every day.
[0049] At this time, it is necessary to clarify the mini-program for which information mining is currently required, and this mini-program is called the specified mini-program. In the access log, extract the URL corresponding to the specified mini-program, and call it the first URL. Among them, when extracting the first URL, it can be confirmed according to the path pattern (PATH), sub-domain name, etc. corresponding to the specified mini-program.
[0050] S104: Filter the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter pattern.
[0051] For large-scale applications and sub-programs, the number of first URLs extracted is still large, so it is difficult to directly perform information mining.
[0052] In the embodiments of this specification, the main purpose is to obtain the detail page, so as to perform information mining based on the detail page and obtain corresponding business information. Among them, for different mini-programs, their business information can also be different. For example, the business information can include product information, travel information, payment information, etc. When the mini-program is an e-commerce platform mini-program, the required pages can include product detail pages, promotion activity detail pages, etc., for obtaining detailed product content or activity content. When the mini-program is a travel platform mini-program, the required pages can include hotel detail pages, flight activity detail pages, etc. When the mini-program is a business handling mini-program, the required pages can include specific business detail pages, etc.
[0053] Generally speaking, in the first URL generated by a user accessing a mini-program, it will include a path (PATH) and corresponding parameter key-value pairs. For example, taking an e-commerce platform mini-program as an example, when a user accesses a product detail page, a corresponding first URL is generated. An example of the first URL can be: https: / / h5.xxx.xx / goodsinfo / ?item_id=123&trace_id=abc. Here, in order to protect the privacy of the e-commerce platform, the path in the first URL is represented by "xxx" here, and the parameter values are represented by "123", "abc", etc. And for different applications and mini-programs, the beginning of their URLs can be https, http, or other strings set independently. The URLs in the embodiments of this specification are only exemplary descriptions, and the actual string content included can be modified as needed.
[0054] In this first URL, the part "https: / / h5.xxx.xx / goodsinfo / " before the "?" is the path, and the part after the "?" includes two parameter key-value pairs. Different parameter key-value pairs are separated by "&", and within a parameter key-value pair, the parameter name and the parameter value are connected by "=". For example, it contains two parameter key-value pairs: "item_id = 123" and "trace_id = abc". Inside "item_id = 123", the parameter name is "item_id", which can represent the product ID, and the parameter value is "123", representing the specific value of the product ID; inside "trace_id = abc", the parameter name is "trace_id", which can represent the user's access source, and the parameter value is "abc", representing the specific path of the access source.
[0055] Of course, the first URL here is just an example. In an actual URL, the specific representation of its path and parameter key-value pairs may be different accordingly. However, they can all generate parameter patterns through the information mining method for applets in the embodiments of this specification.
[0056] Based on the conventional programming settings for parameter key-value pairs in the URL and the actual product or business status in the applet, the rules for the URL of the details page are defined to facilitate setting corresponding filtering rules for the parameter key-value pairs. In the first URL, the parameter key-value pairs that obviously do not belong to those used to describe the product details page or business details page are filtered.
[0057] Among the remaining parameter key-value pairs after filtering, use their parameter names as the core parameter names. At this time, ignore their parameter values and only retain the core parameter names and the path to generate the corresponding parameter pattern.
[0058] At this time, changing the parameter value of the core parameter name often causes changes to the details page.
[0059] Still taking the first URL exemplified above as an example, after filtering, it is considered that both "item_id" and "trace_id" are core parameter names. Then its corresponding parameter pattern is PATH = https: / / h5.xxx.xx / goodsinfo / , and the core parameter set = [item_id, trace_id]. For the following two first URLs, "https: / / h5.xxx.xx / goodsinfo / ?item_id = 123&trace_id = abc" and "https: / / h5.xxx.xx / / goodsinfo / ?item_id = 321&trace_id = def", they correspond to two different product details pages, but their paths and core parameter names are the same, so their corresponding parameter patterns are also the same.
[0060] It should be noted that the parameter pattern here is only described by way of example and does not mean that the parameter pattern of this first URL must be the parameter pattern in the above example. In the following examples, in order to explain different situations and different embodiments, for the same first URL, there may be different descriptions of the parameter pattern. Therefore, the examples here are not fixed and do not conflict with the examples in the following text. For example, if it is considered that only "item_id" is the core parameter name after filtering, then in the final parameter pattern, only the path and this single core parameter name are included.
[0061] Specifically, for the URL rule of the detail page, it may include:
[0062] URL Rule 1: For detail pages such as goods and services, considering conventional programming settings, the ID usually appears in the URL, and it is the same set of URL parameter patterns, and the page switching is achieved by changing the parameter values. For example: "https: / / h5.xxx.xx / p / goodsinfo?item_id=101&...""https: / / h5.xxx.xx / p / goodsinfo?item_id=102&..." usually represents 2 different goods detail pages.
[0063] URL Rule 2: Generally speaking, still considering conventional programming settings, if a certain set of URL parameter patterns is determined to be a detail page, then other URLs that conform to this rule are generally of the same type of detail page. For example, the parameter pattern of a certain goods detail page is "the first layer PATH = https: / / h5.ele.me / p / goodsinfo, and its core parameter set = [biz_type, item_id, store_id]", then other URLs that meet this set of parameter patterns are very likely to be goods detail pages.
[0064] URL Rule 3: For parameters such as goods ID and service ID, considering conventional programming settings, the parameter values are generally pure numbers with a length of 5 to 30 digits or a combination with a small amount of English; and considering the actual state of goods or services, there are many variants of the parameter values, generally exceeding 1000 (excluding enumerated value parameters). For example: goods ID = 747633332547, service ID = XX202401231231231, and there are also some complex IDs, such as: a certain hotel ID = hz_1003 (representing hotel No. 1003 in the hz area).
[0065] Based on this, corresponding filtering rules are set to filter the parameter key-value pairs included in the first URL. After generating the parameter pattern, other URLs that meet the parameter pattern can be used as product detail pages for information mining after accessing these detail pages.
[0066] Furthermore, the filtering rules may include: a preset parameter name blacklist, a preset parameter name combination rule, a preset parameter value quantity limit, etc. In actual work, one or more filtering rules can be selected from these filtering rules to filter the parameter key-value pairs included in the first URL.
[0067] Regarding the parameter name blacklist, blacklists are set for some common parameter names that are considered not to belong to core parameter names for filtering. For example, for the parameter name "timestamp", its usual function is to represent a timestamp and is used to describe the access time of the URL, etc., which does not affect the generation of the detail page. Therefore, it can be added to the blacklist. Also, for the parameter name "order_id", its usual function is to represent the unique identifier corresponding to an order, transaction, business handling, etc., and is used to record and classify each transaction or business handling, and also does not affect the generation of the detail page. Therefore, it can be added to the blacklist.
[0068] Regarding the parameter name combination rule, considering the characteristics of product IDs, business IDs, etc., their parameter values are generally combinations of pure numbers with a length of 5 to 30 digits or with a small amount of English. Therefore, corresponding combination rules are set. For example, the combination rules set the length limit of the parameter name, the proportion of English characters in the parameter name, etc., thus forming the corresponding parameter name combination rule.
[0069] Regarding the preset parameter value quantity limit, for a mature small program, the number of products and the number of services in its detail page usually have a certain scale. Among them, the manifestation of the number of products may be more obvious. Therefore, for small programs such as those for collecting product detail pages, filtering rules corresponding to the parameter value quantity limit can be set, while for small programs used for performing services, this filtering rule can be not set. For small programs corresponding to the number of products, which can include e-commerce platforms, food delivery platforms, etc., corresponding quantity limits can be set according to the scale of different platforms. For example, for a large-scale e-commerce platform, it can be set that the parameter value quantity (that is, the variant of the parameter value) is not less than 10,000 types, while for a food delivery platform or a general-scale point-of-sale platform, it can be set that the parameter value quantity is not less than 1,000 types. Among them, if the method of random sampling is selected in the first URL of the specified small program for filtering parameter names, this quantity also needs to be reduced proportionally.
[0070] In the actual filtering process, if multiple filtering rules are selected for filtering, since the filtering rule for the parameter name blacklist is relatively simple, with low computing power consumption and fast processing speed, it can be used in the first round of filtering.
[0071] For the parameter name combination rule, its filtering rule is relatively more complex, but it is still limited to being able to filter the parameter name within a single parameter name in a single URL. Therefore, it is used in the second round of filtering. At this time, a large number of parameter names have been filtered out through the first round of filtering, which can effectively shorten the computing power and time consumed in the second round of the process.
[0072] Regarding the parameter value quantity limit, its filtering rule is even more complex. It is necessary to collect the variant quantities of the parameter values of this parameter name in multiple URLs. Filtering cannot be achieved through a single URL, and considering that perhaps not all applets require the filtering of this step, it is therefore placed in the last round of the filtering process.
[0073] Generate a parameter pattern based on the path and core parameter names included in the filtered first URL. In the filtered first URL, it contains a path and one or more parameter key-value pairs. In this parameter key-value pair, delete the parameter value and use the parameter name as the core parameter name. Finally, only retain the path and core parameter names to generate the parameter pattern.
[0074] Based on practical experience judgment, for a single applet, the number of parameter models obtained after three rounds of filtering is generally only 1 to 10.
[0075] Of course, in some cases, manual collection and statistics can also be carried out to generate the corresponding parameter pattern. For example, manually visit the required detail page URL, collect this URL in real time and perform manual processing, retain the PATH and core parameter names, generate the parameter pattern, and after confirming that the page is normal, organize to obtain the final parameter pattern.
[0076] S106: Normalize the first URL according to the content included in the parameter pattern to obtain a second URL.
[0077] Normalization can also be called forced normalization, which means that in the first URL, eliminate the invalid parameter key-value pairs outside the parameter pattern, thereby completing the refinement of the first URL.
[0078] At this time, the second URL only contains the path in the parameter pattern and the parameter key-value pairs corresponding to the core parameter names.
[0079] For example, assume that the extracted parameter pattern is: PATH = https: / / h5.xxx.xx / p / goodsinfo, and the core parameter name = [item_id]. At this time, there are two first URLs, namely Example 1 and Example 2. Among them, Example 1 is: https: / / h5.xxx.xx / p / goodsinfo?item_id = 11&biz_type = 0&store_id = 22, and Example 2 is: https: / / h5.xxx.xx / p / goodsinfo&item_id = 11&biz_type = 0&store_id = 22&comment_id = 33.
[0080] At this time, the paths PATH of these two first URLs are the same, and the parameter contains the core parameter name "item_id", but it is possible that Example 1 is the product details page and Example 2 is the product evaluation page.
[0081] After normalizing these two first URLs, the final obtained second URLs are both: https: / / h5.ele.me / p / goodsinfo?item_id = 11, that is, a product details page and a product evaluation page are both normalized to the product details page. Returning to the requirement goal, in this specification, in order to mine product information, only the product details page is required to complete the information mining. By this normalization method, more target pages may be obtained, which belongs to a valuable elimination and screening process, and is called normalization or forced normalization.
[0082] As Figure 2 shown, for the convenience of description, the set where the second URL is located is called the first result set. After obtaining the first result set, corresponding data statistics can be performed. For example, perform page view statistics (PV statistics), number of visiting clients statistics (UV statistics), visiting address statistics, etc., so as to facilitate information mining in the subsequent
[0083] S108: Access the corresponding first target page through the second URL to perform information mining.
[0084] After obtaining the second URL, the second URL can be directly accessed, and the page obtained by the access is called the first target page here.
[0085] Information mining can also be called page understanding. As Figure 2 shown, extract the basic information of the first target page, including page views, applet name, etc. Take a screenshot of the first target page and perform optical character recognition (OCR) to extract the corresponding information.
[0086] The process of extracting information can be to use a large model to extract the core elements in the page, such as page name, product name, product price, hotel name, hotel address, items required for the business, etc. At the same time, in the data statistics stage, information such as the number of visits, the number of access clients, and the distribution of access user addresses can also be obtained, and based on requirements, it can be unified with the core elements for analysis.
[0087] The extraction and selection of information can be set based on actual requirements. For example, the information mined can include product name, product price, business process, origin and destination of travel, etc. Subsequently, this information can be analyzed to implement functions such as more accurate big data product push and replenishment, and relevant business help information for users.
[0088] The process of extracting and understanding information can be achieved through a large model. For example, Figure 2 as shown, use the large model to judge whether the first target page is the required detail page, and through the understanding of the large model, obtain the required core elements.
[0089] At this time, secondary deduplication can also be performed based on requirements. When there is content in the extracted core elements that is the same as the previously extracted core elements, it is considered that the first target page may have been extracted before, and it is deduplicated a second time.
[0090] Through parameter filtering and URL normalization, quickly find the parameter pattern of the target URL, reduce the processing of redundant parameters, reduce the computational complexity, speed up the analysis of a large number of URLs, and retrieve all the required target URLs visited by users at one time, improving the timeliness from monthly to hourly.
[0091] Retain the key parameters through parameter pattern filtering, eliminate irrelevant interference, and the data that meets the parameter pattern is basically accurate URLs, thereby improving the accuracy of data analysis and mining results.
[0092] The filtering rules can be dynamically adjusted to support the requirements of different mini-program scenarios and have strong scalability.
[0093] Based on Figure 1 this method, this specification also provides some specific implementation schemes and extended schemes of this method, which will be continued to be described below.
[0094] In one or more embodiments of this specification, after obtaining the parameter pattern, it can be directly used for normalizing the first URL to obtain the second URL. Of course, in order to further increase the insurance and clarify the reliability of the parameter pattern, multiple example URLs can be generated according to this, and by accessing the pages of the example URLs, confirm whether they are the required target pages, so as to further confirm whether the parameter pattern is reliable.
[0095] Specifically, for the parameter mode, a number of example URLs are generated, and the corresponding second target pages are accessed through the example URLs. Here, the second target page refers to the page corresponding to the example URL.
[0096] The large model is used to mine information from the second target page to obtain the first business information, and the parameter modes whose first business information does not meet the requirements are deleted.
[0097] Taking the product detail page as an example, the first business information may include product name, price, specifications, express or delivery method, inventory, etc. If the second target page is a product detail page, it often contains this first business information and meets certain requirements. For example, there is usually only one product name, the express or delivery method contains a preset brand name, and the inventory quantity does not exceed a certain amount, etc.
[0098] If, through the analysis of the large model, the first business information of all or more than a certain proportion (such as 90%) of the example URLs meets the information in the product detail page, then this parameter mode can be considered reliable.
[0099] Of course, a cross-modal large model can also be used. The screenshot of the second target page is directly input into the cross-modal large model, and it analyzes whether the second target page is the required product detail page.
[0100] Furthermore, multiple channels can be used to generate example URLs to make this verification more reliable.
[0101] After determining the parameter mode, for all parameter modes, all core parameter names are extracted from them. For the filtered first URLs, the occurrence frequencies of the parameter values corresponding to the core parameter names are determined. Here, the occurrence frequency = the number of first URLs where the core parameter name is located / the total number of first URLs.
[0102] In the first method, stratified sampling is performed according to the occurrence frequency to obtain multiple parameter values, and multiple first example URLs are generated based on these multiple parameter values. For example, multiple segments are divided in order of the occurrence frequency from high to low. When performing stratified sampling, ensure that at least one filtered first URL is sampled in each segment. The sampling quantities in different segments can be the same or different. When they are different, a completely random method can be used, or the sampling quantity can be set according to the occurrence frequency.
[0103] For the first example URLs, since they are obtained from the first URLs, they can basically be accessed successfully. At this time, only the business information on their pages needs to be collected, and then it can be judged whether it is the required product detail page based on this page.
[0104] In the second method, multiple virtual values are added to the segments with the highest and / or lowest frequencies, and multiple second example URLs are generated based on these multiple virtual values.
[0105] For the segment with the highest frequency, the first URL in it may be the product detail page of some popular products or the page of some promotional activities, resulting in its high frequency. For the segment with the lowest frequency, the first URL in it may be the product detail page of some unpopular products, or the product favorites of a certain user, the product recommendation page of a certain user, etc., resulting in its low frequency.
[0106] Based on this, the pattern of the parameter values can be summarized based on the first URL in the highest and / or lowest segments. The pattern can include which interval the parameter values are concentrated in, or which parameter values have the highest frequency. Then, corresponding virtual values are generated according to this pattern. For example, assuming that in the highest segment, the parameter key-value pairs "tem_id = 11" and "tem_id = 15" corresponding to the core parameter name have the highest frequency, virtual values 12, 13, and 14 can be generated, combined with this core parameter name respectively to obtain new parameter key-value pairs "tem_id = 12", "tem_id = 13", "tem_id = 14", and combined with the path to generate three second example URLs. Of course, in addition to this method, other methods such as generating second example URLs near the parameter values of the parameter key-value pairs with the highest frequency can also be used.
[0107] For the second example URL, since it is generated by virtual values, there may be a situation where the product detail page does not exist, resulting in access failure, access exception, etc. At this time, the page may jump to the home page of the applet, or display descriptions such as "product does not exist", "access exception", etc. In this case, just ignore this second example URL.
[0108] Of course, if necessary, for the page judgment of the example URL, manual intervention can also be carried out, and the screenshot is manually judged whether it conforms to the required product detail page for the final evaluation.
[0109] In one or more embodiments of this specification, when normalizing the first URL, for a large number of URLs, directly expanding all URLs and performing equal-value matching of the parameter patterns will be very inefficient. Therefore, as Figure 2As shown, after obtaining the parameter pattern, data can be roughly filtered in the first URL by means of fuzzy matching, and the first URLs that do not conform to the path and the core parameter names can be screened out. For example, by means of like matching, the URLs with unmatched paths and parameter names that do not contain the core parameter names can be quickly filtered out, thereby reducing the computing power requirement and improving the matching efficiency. By matching in this way, generally more than 95% of the URLs can be filtered out.
[0110] In addition, for some URLs, there may be a situation of multiple encodings. At this time, as Figure 2 shown, if it is found that there is still a second nesting after the first decoding and expansion, the second decoding and expansion can be continued until the URL is fully expanded.
[0111] At this time, in the remaining first URLs, other content except the path and the parameter key-value pairs corresponding to the core parameter names is deleted to generate the second URL.
[0112] In one or more embodiments of this specification, after each information mining process is executed, the execution process of this time can be stored in the database. The execution process can include the corresponding parameter pattern, the second URL, etc. For the convenience of description here, the second URL that has been stored in the database after the previous information mining is called the historical URL.
[0113] Based on this, after obtaining the second URL, the historical URLs stored in the database are obtained. According to the historical URLs, the second URLs are de-duplicated to obtain the incremental URLs. It can be as Figure 2 shown, after the historical URLs are roughly filtered for data, the URLs are expanded, and the obtained set is called the second result set. For example, two second URLs are obtained this time, which are "https: / / h5.xxx.xx / p / goodsinfo?item_id=11" and "https: / / h5.xxx.xx / p / goodsinfo?item_id=12" respectively, which represent two different product detail pages. And in the historical URLs, "https: / / h5.xxx.xx / p / goodsinfo?item_id=12" already exists. Therefore, it is de-duplicated and deleted in the second URLs, and only "https: / / h5.xxx.xx / p / goodsinfo?item_id=11" is retained. For the convenience of description, the URL obtained after de-duplication is called the incremental URL.
[0114] Attempt to access the corresponding first target page according to the incremental URL. For an incremental URL, it can usually be accessed successfully. At this time, information mining is performed to obtain the second business information, and the corresponding required information can be obtained through the second business information, or further determine whether the first target page is the required page, etc.
[0115] Of course, there are also some cases where the access fails. For example, for some applets, their access logic is relatively strict, and there may be additional special considerations. The server will force all parameters to be passed back. Or, in addition to the core parameter names, some parameters are also required to be passed back. Once a parameter key-value pair is deleted from the URL, it may cause the access to fail. For example, for the URL "http: / / xxx.html?itemId=100&bizType=2", after judgment, in the parameter mode, the core parameter name is "itemId", and the parameter name "bizType" is mainly used to describe the business type and is not the core parameter name.
[0116] However, in the regulations of the access logic of some applets, when accessing this URL, if "bizType" is not carried, it may report an error, or the entered page does not meet the specifications. Such parameters can be called auxiliary parameters here.
[0117] At this time, the incremental URL can be directly abandoned and other URLs can be continued to be accessed, or it can be handed over to manual processing. However, for the same applet, the access logic set by its server is usually the same. Once the incremental URL is abandoned, it may cause all the incremental URLs in this batch to be abandoned. And handing it over to manual processing will increase the labor cost and be time-consuming and laborious.
[0118] Based on this, if the access fails, then according to the deleted content corresponding to the incremental URL, restore the incremental URL and update the parameter mode corresponding to the incremental URL. Among them, the deleted content mainly refers to the deleted parameter key-value pairs. By restoring the deleted parameter key-value pairs, it is determined which parameters the applet needs at least to access successfully.
[0119] Specifically, according to the preset parameter name restoration list and the positional relationship between each parameter key-value pair in the deleted content corresponding to the incremental URL and the core parameter name, determine the corresponding restoration priority.
[0120] Contrary to the parameter name blacklist, there is a parameter name restoration list (which can also be called a parameter name whitelist). In this parameter name restoration list, some common auxiliary parameter names are recorded. Generally speaking, the parameter names set in the parameter name restoration list can be considered to have the highest restoration priority.
[0121] For other parameters that are not recorded in the parameter name restoration list and appear in the incremental URL, it is generally considered that the closer they are to the core parameter name, the greater the possibility that they are auxiliary parameters. Therefore, the restoration priority can be determined according to the distance between them and the core parameter name. The closer the distance is to the core parameter name, the higher the restoration priority. Of course, the highest restoration priority will not exceed the parameters in the parameter name restoration list.
[0122] When there are multiple core parameter names in the incremental URL, the restoration priority of each parameter can be determined by the distance between each parameter and the nearest core parameter name.
[0123] According to the restoration priority, the parameter key-value pairs included in the deleted content of the incremental URL are restored in sequence. After each restoration, access to the first target page corresponding to the incremental URL is retried.
[0124] Among them, restoring in sequence means that, in the order from high to low of the restoration priority, a single parameter key-value pair is selected in sequence and restored to the corresponding position before deletion. If there are multiple parameter key-value pairs with the same restoration priority, the priority can be randomly selected.
[0125] Generally speaking, the number of auxiliary parameters is at most one or two. Therefore, if a certain parameter key-value pair is successfully accessed during the restoration process of a single parameter key-value pair, the parameter mode corresponding to the incremental URL is updated according to the parameter key-value pair corresponding to this restoration.
[0126] Of course, if all single parameter key-value pairs fail to be accessed, the combinations of all two, three, and up to all parameter keys are determined in ascending order of quantity, and then restored in sequence according to the restoration priority of this combination. Among them, the restoration priority of each combination is determined according to the parameter key-value pair with the highest restoration priority in this combination.
[0127] In this way, for the incremental URL, the auxiliary parameters required for successful access can be determined. If the auxiliary parameters of multiple incremental URLs are the same, the parameter name of this auxiliary parameter can be added to the parameter mode as the core parameter name.
[0128] Furthermore, for auxiliary parameters, which are different from the already determined core parameter names, not only do their parameter names have certain meanings, but their parameter values may also have certain meanings. For example, for the auxiliary parameter "bizType", when bizType = 2, it indicates that it belongs to the specified second business type. At this time, this business type may be a product access service, and the product details page can be accessed based on this business type. However, if the parameter value of "bizType" is changed to another value, for example, changed to bizType = 3, at this time, this business type may be converted to the payment service for this product, and at this time, it is no longer possible to obtain the product details page.
[0129] Based on this, determine the first target page accessed according to the parameter key-value pairs restored this time, that is, after restoration, determine the first target page to be accessed.
[0130] In the parameter key-value pairs, replace the parameter values and access the corresponding third target page according to the incremental URL after replacing the parameter key-value pairs, that is, after changing the parameter value of the auxiliary parameter, re-access the corresponding URL, and call this page the third target page.
[0131] If the format difference degree between the business information of the third target page and the first target page is lower than the preset degree, update the parameter mode corresponding to the incremental URL according to the parameter names in the parameter key-value pairs restored this time.
[0132] The format difference degree here mainly refers to whether the two contain the same format content. For example, assume that both contain content such as product name and product price. Even if the specific values of the product name, price, etc. between the two are different, it can be considered that the format difference degree between the two is lower than the preset degree.
[0133] If the format difference degree is lower than the preset degree, it can be considered that even if the parameter value of the auxiliary parameter is changed, only other product details pages are obtained, or the current product details page is still stayed at. At this time, only update the parameter mode according to the parameter name of the auxiliary parameter.
[0134] If other parameter value attempts are made and most of the parameter values of other parameter values result in access failures, or their format difference degree is relatively large, it is considered that this auxiliary parameter needs to have a fixed parameter value to ensure that the accessed page is the required product details page. At this time, when updating the parameter mode, the parameter name and parameter value of this auxiliary parameter need to be updated in the parameter mode together.
[0135] Among them, in order to ensure accuracy, the parameter values of auxiliary parameters can be replaced for multiple incremental URLs. When multiple incremental URLs all meet the requirement of fixed parameter values, or do not require fixed parameter values, the parameter mode is updated.
[0136] In one or more embodiments of this specification, when deduplicating the second URL according to the historical URL, the same URLs are deleted from the second URL according to the historical URL. This step has been described above and will not be elaborated here.
[0137] However, the deduplication corresponding to this process can only remove exactly the same second URLs. In fact, there are still some second URLs that are seemingly different but actually the same and have not been removed.
[0138] For example, when the applet is upgraded or the interface is updated, based on the original core parameter name "itemId", a new parameter name "itemId2" is added or replaced. Essentially, the effects and the pages they connect to are the same. It is difficult to perform deduplication only through the above method.
[0139] Based on this, for the core parameter name in the parameter mode corresponding to the second URL, it is determined whether there is a corresponding standardization relationship with the historical parameter name in the parameter mode corresponding to the historical URL. For the convenience of description, the parameter names included in the historical URL are referred to as historical parameter names here.
[0140] The standardization relationship can be preset. Specifically, it can be set as follows: If between two parameter names, the core parameter name has a few more characters (including letters, numbers, special symbols, etc.) compared to the historical parameter name, it can be considered that there is a standardization relationship between the two. For example, assuming the historical parameter name is "itemId", then parameter names such as "itemId2", "itemIdV2", "itemId_v2)" can all be considered to have a standardization relationship with this historical parameter name. Or, if between two parameter names, most of the characters are the same, only one or two characters are different, and there is a numerically continuous relationship between the different characters, it can also be considered that there is a standardization relationship between the two. For example, assuming the historical parameter name is "itemId1", then "itemId2" can be considered to have a standardization relationship with this historical parameter name. Or, for the words included in the historical parameter name and the core parameter name, although their representations in characters are different, but after being translated into the language of the region where the applet is located, the translation results are basically the same, it can also be considered that there is a standardization relationship between the two. For example, assuming the historical parameter name is "itemId", then "objectId" can be considered to have a standardization relationship with this historical parameter name.
[0141] The specific settings included in the normalization relationship can be set based on actual requirements.
[0142] Here, it is generally considered that for the same product detail page, in the absence of special events (such as promotional activities, the product becoming extremely popular, etc.), its access frequency is usually relatively stable. Therefore, when a new core parameter name corresponds to a new interface to access this product detail page, the access frequency of the old interface corresponding to the historical parameter name should decrease.
[0143] Therefore, if there is a core parameter name with a normalization relationship, based on the access frequency relationship between this core parameter name and the parameter key-value pair where the historical parameter name is located, it can be determined whether there is an update relationship between the core parameter name and the historical parameter name.
[0144] If there is an update relationship between the two, it is considered that even though their core parameter names are different, as long as the parameter values are the same, the product detail pages corresponding to their URLs should also be the same. Therefore, in the second URL, the second URL with the parameter value corresponding to the core parameter name being the same as the parameter value corresponding to the historical parameter name in the historical URL can be deleted. For example, for the same product detail page, its second URL is "http: / / xxx?itemId2 = 123", and its historical URL is "http: / / xxx?itemId = 123". If it is determined that there is an update relationship between "itemId" and "itemId2" and their parameter values are the same, it can be considered that the two point to the same product detail page, and this second URL can be deleted as a duplicate URL.
[0145] Of course, if you are not confident, you can also select several second URLs for access to determine whether they are really duplicate pages.
[0146] Furthermore, when determining whether there is an update relationship between the core parameter name and the historical parameter name, screening can be performed based on the parameter values. For this core parameter name and this historical parameter name, determine the first access frequency and the second access frequency corresponding to the parameter key-value pairs with the same parameter values respectively. For example, the second URL is "http: / / xxx?itemId2 = 123", and the historical URL is "http: / / xxx?itemId = 123". There is a normalization relationship between the parameter names of the two and the parameter values are the same. Then the access frequencies of the two parameter key-value pairs can be obtained respectively. Here, the access frequency corresponding to the key-value pair "itemId2 = 123" where the core parameter name is located is called the first access frequency, and the access frequency corresponding to the key-value pair "itemId = 123" where the historical parameter name is located is called the second access frequency.
[0147] Among them, the access frequency can be obtained by querying the proportion in all the corresponding URLs. For example, for the key-value pair "itemId2 = 123" where the core parameter name is located, search for the URLs in which this parameter key-value pair appears in all the first URLs, so as to obtain its access frequency.
[0148] Generally speaking, in the process of information mining, for the product detail page, it is necessary to obtain the user access frequency of each product in order to facilitate accurate delivery and big data analysis. Therefore, for the key-value pairs where the historical parameter names are located, they can be obtained from the database or the analysis results of information mining.
[0149] At this time, overall, if within a certain time range, the sum of the first access frequency and the second access frequency is in a stable range, and the first access frequency rises while the second access frequency drops, it can be considered that the user gradually starts to use the upgraded mini-program of the system to access this product detail page. Therefore, in fact, the two correspond to the same product detail page. The rising and falling speeds at this time may be determined by factors such as the update frequency of the application where the mini-program is located for different users and the delivery area of the mini-program system upgrade.
[0150] Or, from an individual perspective, there are multiple users. Within a specified time period, if the first access frequency suddenly rises and the second access frequency suddenly drops (usually drops to 0), it is considered that this user has used the interface after the mini-program upgrade and will naturally no longer use the previous interface to access. Therefore, the access frequency corresponding to its parameter key-value pair will mutate.
[0151] If either of the above two situations occurs, it can be determined that there is an update relationship between the core parameter name and the historical parameter name.
[0152] In one or more embodiments of this specification, the main purposes of page mining include making more intelligent recommendations to users. For some products, when they may include multiple product detail pages, here a URL and a target page can be generated through the parameter mode so that the user can view all the product detail pages at one time, thereby realizing convenient browsing for the user.
[0153] Specifically, if there are multiple parameter modes, in the multiple parameter modes, extract the specified parameter modes with the same path. Generally speaking, for the same mini-program, the paths of the parameter modes are generally the same. Of course, there are also some mini-programs that have made detailed path divisions for their different modules. At this time, for each path, select the corresponding specified parameter mode for itself.
[0154] In the specified parameter mode, some or all of the core parameter names are extracted. According to the same path and the extracted some or all of the core parameter names, a combined parameter mode is generated. Generally speaking, all of the core parameter names can be directly extracted. However, when there are too many core parameter names, only some of the core parameter names can be selected. For example, there are two parameter modes in total, which are: PATH = https: / / h5.xxx.xx / goodsinfo / , core parameter name = item_id, and PATH = https: / / h5.xxx.xx / goodsinfo / , core parameter name = trace_id. At this time, all of the core parameter names "item_id" and "trace_id" under this path are selected. And the generated combined parameter mode is: PATH = https: / / h5.xxx.xx / goodsinfo / , core parameter name = item_id&trace_id.
[0155] In this way, a new URL can be generated according to the combined parameter mode, and the fourth target page corresponding to the new URL can be recommended to the user. In this fourth target page, there are the target pages corresponding to "item_id" and the target pages corresponding to "trace_id", enabling the user to complete the detailed browsing of a product through one address (i.e., the new URL) and one page (i.e., the fourth target page).
[0156] Among them, when generating the new URL, it is necessary to clarify that it is the address for the same product and then perform splicing, otherwise it will have the opposite effect of misleading the user. Therefore, in the second URL, the parameter value corresponding to the core parameter name is determined. This parameter value can be any parameter value included in the second URL. Of course, it can also be further strictly required that this parameter value needs to generate parameter key-value pairs for all of the core parameter names in the combined parameter mode in the second URL, so as to further ensure that the product has corresponding target pages for each core parameter name.
[0157] At this time, according to the combined parameter mode and this parameter value, a new URL is generated, and an attempt is made to access the fourth target page corresponding to the new URL. At this time, the new URL can be generated in a preset format. For example, the generated new URL can be https: / / h5.xxx.xx / goodsinfo / ?item_id = 11&trace_id = 11.
[0158] Since this new URL is not the address format originally set for the specified applet, an attempt to access it is required.
[0159] If the access is successful, when the user visits the product next time, the fourth target page corresponding to the newly added URL can be directly displayed to the user for the user to browse.
[0160] If the access fails, it means that the specified applet does not support this address format. At this time, if a new address can be generated through the server of the specified applet, the address of the newly added URL can be generated so that the fourth target page contains all the content of the original two target pages.
[0161] If it is difficult to perform corresponding processing through the server of the specified applet, a browser entry can be created in the application where the specified applet is located to access the target pages corresponding to the parameter key-value pairs of the newly added URL respectively through the application and splice them to obtain the fourth target page.
[0162] Figure 3 It is a schematic diagram of the generation of the fourth target page in an application scenario provided by one or more embodiments of this specification. When the user visits the product next time, with the permission of the specified applet, the specified applet feeds back the access link to the application. The application can call the internal browser of itself, create an entry through this browser to enter the address of the newly added URL, and then the target pages corresponding to the parameter key-value pairs of the application respectively. After splicing the target pages in the corresponding format, the fourth target page is obtained and displayed to the user.
[0163] Based on the same idea, one or more embodiments of this specification also provide a device and equipment corresponding to the above method, as Figure 4 、 Figure 5 shown.
[0164] Figure 4 It is a schematic diagram of the structure of an information mining device for applets provided by one or more embodiments of this specification. The device includes:
[0165] A URL acquisition module 402 that acquires the first URL corresponding to the user's access to the specified applet;
[0166] A parameter pattern generation module 404 that filters the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter pattern;
[0167] A URL normalization module 406 that normalizes the first URL according to the content included in the parameter pattern to obtain a second URL;
[0168] An information mining module 408 that accesses the corresponding first target page through the second URL for information mining.
[0169] Optionally, the parameter pattern generation module 404 filters the parameter key-value pairs included in the first URL based on a preset parameter name blacklist, and / or a preset parameter name combination rule, and / or a preset parameter value quantity limit;
[0170] Based on the path and core parameter names included in the filtered first URL, a parameter pattern is generated.
[0171] Optionally, it further includes an example verification module 410;
[0172] The example verification module generates a number of example URLs for the parameter pattern;
[0173] Access the corresponding second target page through the example URL;
[0174] Through information mining on the second target page by a large model, first service information is obtained, and the parameter patterns for which the first service information does not meet the requirements are deleted.
[0175] Optionally, the example verification module 410 determines the occurrence frequency of each parameter value corresponding to the core parameter name for the filtered first URL;
[0176] Based on the occurrence frequency, segmented sampling is performed to obtain multiple parameter values, and multiple first example URLs are generated according to the multiple parameter values;
[0177] In the segments with the highest and / or lowest occurrence frequencies, multiple virtual values are added, and multiple second example URLs are generated according to the multiple virtual values.
[0178] Optionally, the URL normalization module 406 screens out the first URLs that do not conform to the path and the core parameter names in the first URL by means of fuzzy matching;
[0179] In the remaining first URLs, other content except the path and the parameter key-value pairs corresponding to the core parameter names is deleted to generate a second URL.
[0180] Optionally, the information mining module 408 obtains the historical URLs stored in the database;
[0181] Based on the historical URLs, the second URL is de-duplicated to obtain an incremental URL;
[0182] Attempt to access the corresponding first target page according to the incremental URL;
[0183] If the access is successful, information mining is performed to obtain second service information;
[0184] If the access fails, restore the incremental URL according to the deleted content corresponding to the incremental URL, and update the parameter mode corresponding to the incremental URL.
[0185] Optionally, the information mining module 408 determines the corresponding restoration priority according to the preset parameter name restoration list and the positional relationship between each parameter key-value pair and the core parameter name in the deleted content corresponding to the incremental URL;
[0186] Restore the parameter key-value pairs included in the deleted content of the incremental URL in sequence according to the restoration priority, and after each restoration, retry accessing the first target page corresponding to the incremental URL;
[0187] If the access is successful, update the parameter mode corresponding to the incremental URL according to the parameter key-value pair corresponding to this restoration.
[0188] Optionally, the information mining module 408 determines the first target page accessed according to the parameter key-value pair of this restoration;
[0189] In the parameter key-value pair, replace the parameter value, and access the corresponding third target page according to the incremental URL after replacing the parameter key-value pair;
[0190] If the format difference degree of the service information between the third target page and the first target page is lower than the preset degree, update the parameter mode corresponding to the incremental URL according to the parameter name in the parameter key-value pair corresponding to this restoration.
[0191] Optionally, the information mining module 408 deletes the same URL in the second URL according to the historical URL;
[0192] For the core parameter name in the parameter mode corresponding to the second URL, determine whether there is a corresponding standardization relationship with the historical parameter name in the parameter mode corresponding to the historical URL;
[0193] If there is, determine whether there is an update relationship between the core parameter name and the historical parameter name based on the access frequency relationship between the parameter key-value pairs where the core parameter name and the historical parameter name are located;
[0194] If so, in the second URL, delete the second URL where the parameter value corresponding to the core parameter name is the same as the parameter value corresponding to the historical parameter name in the historical URL.
[0195] Optionally, the information mining module 408 determines the first access frequency and the second access frequency corresponding to the parameter key-value pairs with the same parameter values for the core parameter name and the historical parameter name;
[0196] If the sum of the first access frequency and the second access frequency is within a stable range, and the first access frequency increases while the second access frequency decreases; and / or, there are multiple users, and within a specified time period, the first access frequency suddenly increases and the second access frequency suddenly decreases;
[0197] Then it is determined that there is an update relationship between the core parameter name and the historical parameter name.
[0198] Optionally, it further includes a user recommendation module 412;
[0199] For the user recommendation module, if there are multiple parameter patterns, among the multiple parameter patterns, extract the specified parameter pattern with the same path;
[0200] In the specified parameter pattern, extract some or all of the core parameter names;
[0201] According to the same path and the extracted some or all of the core parameter names, generate a combined parameter pattern;
[0202] Generate a new URL according to the combined parameter pattern, and recommend the fourth target page corresponding to the new URL to the user.
[0203] Optionally, the user recommendation module 412 determines the parameter value corresponding to the core parameter name in the second URL;
[0204] According to the combined parameter pattern and the parameter value, generate a new URL, and attempt to access the fourth target page corresponding to the new URL;
[0205] If the access fails, create a browser entry in the application where the specified applet is located, so as to access the target pages corresponding to each parameter key-value pair of the new URL respectively through the application, and splice them to obtain the fourth target page.
[0206] Figure 5 It is a schematic structural diagram of an information mining device for applets provided by one or more embodiments of this specification. The device includes:
[0207] At least one processor; and,
[0208] A memory communicatively connected to the at least one processor; wherein,
[0209] The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor, so that the at least one processor can:
[0210] Obtain the first URL corresponding to the user's access to a specified applet;
[0211] Based on a preset filtering rule, filter the parameter key-value pairs included in the first URL to generate a corresponding parameter pattern;
[0212] Normalize the first URL according to the content included in the parameter pattern to obtain a second URL;
[0213] Access the corresponding first target page through the second URL for information mining.
[0214] Based on the same idea, one or more embodiments of this specification also provide a non-volatile computer storage medium corresponding to the above method, storing computer-executable instructions, and the computer-executable instructions are set as:
[0215] Obtain the first URL corresponding to the user's access to a specified applet;
[0216] Based on a preset filtering rule, filter the parameter key-value pairs included in the first URL to generate a corresponding parameter pattern;
[0217] Normalize the first URL according to the content included in the parameter pattern to obtain a second URL;
[0218] Access the corresponding first target page through the second URL for information mining.
[0219] In the 1990s, it was obvious to distinguish whether an improvement in a technology was an improvement in hardware (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement in method flows). However, with the development of technology, many improvements in method flows today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit, whose logical function is determined by the user programming the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without asking the chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply making a little logical programming of the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit implementing the logical method flow.
[0220] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0221] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0222] For the convenience of description, the above devices are described by dividing them into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0223] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0224] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0225] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0226] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0227] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0228] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0229] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0230] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0231] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0232] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the description of the method embodiments.
[0233] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0234] The above is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. An information mining method for applets, comprising: Obtaining a first URL corresponding to a user accessing a specified applet; Filtering the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter pattern; Normalizing the first URL according to the content included in the parameter pattern to obtain a second URL; Accessing a corresponding first target page through the second URL to perform information mining.
2. The method according to claim 1, filtering the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter pattern, specifically comprising: Filtering the parameter key-value pairs included in the first URL based on a preset parameter name blacklist, and / or, a preset parameter name combination rule, and / or, a preset parameter value quantity limit; Generating a parameter pattern according to the path and core parameter names included in the filtered first URL.
3. The method according to claim 2, the method further comprising: Generating a plurality of example URLs for the parameter pattern; Accessing a corresponding second target page through the example URLs; Performing information mining on the second target page through a large model to obtain first business information, and deleting the parameter patterns for which the first business information does not meet the requirements.
4. The method according to claim 3, generating a plurality of example URLs for the parameter pattern, specifically comprising: Determining the occurrence frequency of each parameter value corresponding to the core parameter name for the filtered first URL; Performing segmented sampling according to the occurrence frequency to obtain a plurality of parameter values, and generating a plurality of first example URLs according to the plurality of parameter values; Adding a plurality of virtual values in the segment with the highest and / or lowest occurrence frequency, and generating a plurality of second example URLs according to the plurality of virtual values.
5. The method according to claim 2, normalizing the first URL according to the content included in the parameter pattern to obtain a second URL, specifically comprising: In the first URL, screening out the first URLs that do not conform to the path and the core parameter name by means of fuzzy matching; In the remaining first URLs, deleting other content except the path and the parameter key-value pairs corresponding to the core parameter name to generate a second URL.
6. The method according to claim 5, accessing a corresponding first target page through the second URL to perform information mining, specifically comprising: Obtaining the historical URLs stored in the database; Removing duplicates from the second URL according to the historical URLs to obtain an incremental URL; Attempting to access a corresponding first target page according to the incremental URL; If the access is successful, performing information mining to obtain second business information; If the access fails, restoring the incremental URL according to the deletion content corresponding to the incremental URL, and updating the parameter pattern corresponding to the incremental URL.
7. The method according to claim 6, restoring the delta URL according to the deletion content corresponding to the delta URL and updating the parameter pattern corresponding to the delta URL, specifically including: Determining the corresponding restoration priority according to the preset parameter name restoration list and the positional relationship between each parameter key-value pair and the core parameter name in the deletion content corresponding to the delta URL; Restoring the parameter key-value pairs included in the deletion content of the delta URL in sequence according to the restoration priority, and after each restoration, re-attempting to access the first target page corresponding to the delta URL; If the access is successful, updating the parameter pattern corresponding to the delta URL according to the parameter key-value pair corresponding to this restoration.
8. The method according to claim 7, updating the parameter pattern corresponding to the delta URL according to the parameter key-value pair corresponding to this restoration, specifically including: Determining the first target page accessed according to the parameter key-value pair restored this time; In the parameter key-value pair, replacing the parameter value and accessing the corresponding third target page according to the delta URL after replacing the parameter key-value pair; If the format difference degree of the service information between the third target page and the first target page is lower than the preset degree, updating the parameter pattern corresponding to the delta URL according to the parameter name in the parameter key-value pair corresponding to this restoration.
9. The method according to claim 6, de-duplicating the second URL according to the historical URL to obtain a delta URL, specifically including: Deleting the same URLs in the second URL according to the historical URL; For the core parameter names in the parameter pattern corresponding to the second URL, determining whether there is a corresponding standardization relationship with the historical parameter names in the parameter pattern corresponding to the historical URL; If so, determining whether there is an update relationship between the core parameter name and the historical parameter name based on the access frequency relationship between the parameter key-value pairs where the core parameter name and the historical parameter name are located; If so, deleting the second URLs in which the parameter values corresponding to the core parameter names are the same as the parameter values corresponding to the historical parameter names in the historical URL.
10. The method according to claim 9, determining whether there is an update relationship between the core parameter name and the historical parameter name based on the access frequency relationship between the parameter key-value pairs where the core parameter name and the historical parameter name are located, specifically including: For the core parameter name and the historical parameter name, determining the first access frequency and the second access frequency corresponding to the parameter key-value pairs with the same parameter values; If the sum of the first access frequency and the second access frequency is in a stable range, and the first access frequency increases and the second access frequency decreases; and / or, there are multiple users, and within a specified time period, the first access frequency suddenly increases and the second access frequency suddenly decreases; Then determining that there is an update relationship between the core parameter name and the historical parameter name.
11. The method according to claim 2, the method further includes: If there are multiple parameter patterns, among the multiple parameter patterns, extract the specified parameter pattern with the same path; In the specified parameter pattern, extract some or all of the core parameter names; Generate a combined parameter pattern according to the same path and the extracted some or all of the core parameter names; Generate a new URL according to the combined parameter pattern, and recommend the fourth target page corresponding to the new URL to the user.
12. The method according to claim 11, generating a new URL according to the combined parameter pattern, specifically including: In the second URL, determine the parameter value corresponding to the core parameter name; Generate a new URL according to the combined parameter pattern and the parameter value, and attempt to access the fourth target page corresponding to the new URL; If the access fails, create a browser entry in the application where the specified applet is located, so as to access the target pages corresponding to each parameter key-value pair of the new URL respectively through the application, and splice them to obtain the fourth target page.
13. An information mining device for an applet, including: A URL acquisition module that acquires the first URL corresponding to a user accessing a specified applet; A parameter pattern generation module that filters the parameter key-value pairs included in the first URL based on a preset filtering rule to generate a corresponding parameter pattern; A URL normalization module that normalizes the first URL according to the content included in the parameter pattern to obtain a second URL; An information mining module that accesses the corresponding first target page through the second URL to perform information mining.
14. The device according to claim 13, wherein the parameter pattern generation module filters the parameter key-value pairs included in the first URL based on a preset parameter name blacklist, and / or a preset parameter name combination rule, and / or a preset parameter value quantity limit; Generate a parameter pattern according to the path and the core parameter name included in the filtered first URL.
15. The device according to claim 14, further comprising an example verification module; The example verification module generates a plurality of example URLs for the parameter pattern; Access the corresponding second target page through the example URL; Perform information mining on the second target page through a large model to obtain first service information, and delete the parameter patterns for which the first service information does not meet the requirements.
16. The device according to claim 15, wherein the example verification module determines the occurrence frequency of each parameter value corresponding to the core parameter name for the filtered first URL; Perform segmented sampling according to the occurrence frequency to obtain a plurality of parameter values, and generate a plurality of first example URLs according to the plurality of parameter values; Add a plurality of virtual values in the segment with the highest and / or lowest occurrence frequency, and generate a plurality of second example URLs according to the plurality of virtual values.
17. The device according to claim 14, wherein the URL normalization module screens out the first URLs that do not conform to the path and the core parameter name in the first URL by means of fuzzy matching; In the remaining first URLs, delete the content other than the path and the parameter key-value pairs corresponding to the core parameter names to generate second URLs.
18. The device according to claim 17, wherein the information mining module obtains the historical URLs stored in the database; Deduplicate the second URLs according to the historical URLs to obtain incremental URLs; Attempt to access the corresponding first target pages according to the incremental URLs; If the access is successful, perform information mining to obtain second service information; If the access fails, restore the incremental URLs according to the deleted content corresponding to the incremental URLs, and update the parameter patterns corresponding to the incremental URLs.
19. The device according to claim 18, wherein the information mining module determines the corresponding restoration priorities according to the preset parameter name restoration list and the positional relationship between each parameter key-value pair and the core parameter name in the deleted content corresponding to the incremental URLs; Restore the parameter key-value pairs included in the deleted content of the incremental URLs in sequence according to the restoration priorities, and after each restoration, re-attempt to access the first target pages corresponding to the incremental URLs; If the access is successful, update the parameter patterns corresponding to the incremental URLs according to the parameter key-value pairs corresponding to the current restoration.
20. The device according to claim 19, wherein the information mining module determines the first target pages accessed according to the parameter key-value pairs restored this time; In the parameter key-value pairs, replace the parameter values, and access the corresponding third target pages according to the incremental URLs after replacing the parameter key-value pairs; If the format difference degree of the service information between the third target pages and the first target pages is lower than the preset degree, update the parameter patterns corresponding to the incremental URLs according to the parameter names in the parameter key-value pairs corresponding to the current restoration.
21. The device according to claim 18, wherein the information mining module deletes the same URLs in the second URLs according to the historical URLs; For the core parameter names in the parameter patterns corresponding to the second URLs, determine whether there is a corresponding standardization relationship with the historical parameter names in the parameter patterns corresponding to the historical URLs; If so, determine whether there is an update relationship between the core parameter name and the historical parameter name based on the access frequency relationship between the parameter key-value pairs where the core parameter name and the historical parameter name are located; If so, in the second URLs, delete the second URLs where the parameter values corresponding to the core parameter names are the same as the parameter values corresponding to the historical parameter names in the historical URLs.
22. The device according to claim 21, wherein the information mining module determines the first access frequency and the second access frequency corresponding to the parameter key-value pairs with the same parameter values for the core parameter name and the historical parameter name respectively; If the sum of the first access frequency and the second access frequency is within a stable range, and the first access frequency increases while the second access frequency decreases; and / or, there are multiple users, and within a specified time period, the first access frequency suddenly increases and the second access frequency suddenly decreases; Then it is determined that there is an update relationship between the core parameter name and the historical parameter name.
23. The device according to claim 14, further comprising a user recommendation module; The user recommendation module, if there are multiple parameter patterns, extracts a specified parameter pattern with the same path from the multiple parameter patterns; Extracts some or all of the core parameter names from the specified parameter pattern; Generates a combined parameter pattern according to the same path and the extracted some or all of the core parameter names; Generates a new URL according to the combined parameter pattern, and recommends the fourth target page corresponding to the new URL to the user.
24. The device according to claim 23, the user recommendation module determines the parameter value corresponding to the core parameter name in the second URL; Generates a new URL according to the combined parameter pattern and the parameter value, and attempts to access the fourth target page corresponding to the new URL; If the access fails, a browser entry is created in the application where the specified applet is located, so as to access the target pages corresponding to each parameter key-value pair of the new URL respectively through the application, and the fourth target page is obtained by splicing.
25. An information mining device for an applet, comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Obtain the first URL corresponding to the user's access to the specified applet; Based on a preset filtering rule, filter the parameter key-value pairs included in the first URL to generate a corresponding parameter pattern; Normalize the first URL according to the content included in the parameter pattern to obtain a second URL; Access the corresponding first target page through the second URL for information mining.