Page recognition method, device, computer equipment and storage medium

By querying the characteristics of the target page, similar pages in massive pages can be quickly filtered out, and through precise matching confirmation, the problem of low efficiency of traditional sequential matching is solved, and efficient and accurate page recognition is achieved.

CN113946365BActive Publication Date: 2025-07-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010689776.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-17
Publication Date
2025-07-01
Estimated Expiration
2040-07-17

AI Technical Summary

Technical Problem

In traditional methods, identifying similar pages through sequential matching is inefficient, especially in a massive page code base, which is a long time to match and consumes resources.

Method used

By obtaining the original features of the target page and performing feature dimensionality reduction processing, the page features after dimensionality reduction are obtained, these features are used for inverted index query, and the identifiers of suspected similar pages are quickly filtered out, and similar pages are confirmed through the precise matching of the original features.

Benefits of technology

It realizes the rapid identification of pages similar to the target page in a large number of pages, improving efficiency and accuracy, and reducing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113946365B_ABST
    Figure CN113946365B_ABST
Patent Text Reader

Abstract

This application relates to a page recognition method, device, computer device, and storage medium. The method includes: obtaining original target page features extracted from the page code of a target page; performing feature dimensionality reduction processing on the original target page features to obtain page features after dimensionality reduction; querying page identifiers corresponding to the page features to obtain suspected similar page identifiers; obtaining the original page features corresponding to each of the suspected similar page identifiers; respectively matching each of the original page features with the original target page features; and determining the page corresponding to the original page feature that passes the match as a page similar to the target page. Using this method can improve the efficiency of page recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a page recognition method, apparatus, computer device, and storage medium. Background Art

[0002] With the rapid development of front-end technologies, pages under various platforms are diverse. For example, there are a large number of mini-program pages in the mini-program platform. In many scenarios, it is necessary to screen out pages similar to a target page from a large number of pages. For example, among a large number of pages, there are usually some illegally plagiarized pages, and the platform usually needs to identify similar plagiarized pages to the target page from a large number of pages.

[0003] In traditional methods, sequential matching is used to identify similar pages, that is, the page code of a target page is matched one by one with the page codes of a large number of pages in the code library. The average time consumption of this matching process is relatively high each time. Coupled with the extremely large number of page codes in the code library, usually at the billion level, therefore, the traditional method of page recognition through sequential matching has very low efficiency. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a page recognition method, apparatus, computer device, and storage medium that can improve efficiency.

[0005] A page recognition method, characterized in that the method includes:

[0006] Obtain the original features of the target page extracted from the page code of the target page;

[0007] Perform feature dimensionality reduction processing on the original features of the target page to obtain the page features after dimensionality reduction;

[0008] Query the page identifier corresponding to the page features to obtain the suspected similar page identifier;

[0009] Obtain the original features of the pages corresponding to each of the suspected similar page identifiers;

[0010] Respectively match each of the original features of the pages with the original features of the target page;

[0011] Determine the page corresponding to the original features of the page that passes the match as the page similar to the target page.

[0012] A page recognition apparatus, the apparatus includes:

[0013] A feature acquisition module, configured to obtain the original features of the target page extracted from the page code of the target page;

[0014] A feature dimensionality reduction module, which is used to perform feature dimensionality reduction processing on the original features of the target page to obtain the page features after dimensionality reduction;

[0015] An index query module, which is used to query the page identifier corresponding to the page feature to obtain a suspected similar page identifier;

[0016] A matching module, which is used to obtain the original page features corresponding to each of the suspected similar page identifiers; respectively match each of the original page features with the original features of the target page; and determine the page corresponding to the original page feature that passes the match as the page similar to the target page.

[0017] In one embodiment, the index query module is further used to use the page feature as an index item, and according to the pre-set mapping relationship between the page feature and the page identifier, query at least one page identifier corresponding to the page feature after dimensionality reduction to obtain a suspected similar page identifier.

[0018] In one embodiment, the index query module is further used to locate the position corresponding to the page feature in a pre-set bitmap; wherein, each position in the bitmap uniquely records a page feature; based on the pre-built index, perform an inverted index query with the located position as the index item to obtain at least one page identifier corresponding to the located position, and obtain a suspected similar page identifier; wherein, the index includes the mapping relationship between the position in the bitmap and the page identifier; the page identifier having a mapping relationship with the position refers to the page identifier corresponding to the page code having the page feature recorded at the position.

[0019] In one embodiment, the device further includes:

[0020] An index construction module, which is used to extract features from the page codes in the page code library; perform feature dimensionality reduction processing on the extracted features of each page code; record the page features after dimensionality reduction at the corresponding positions in the bitmap; for the positions in the bitmap, determine the page codes having the page features recorded at the positions, and establish the mapping relationship between the positions and the page identifiers corresponding to the determined page codes to generate an index.

[0021] In one embodiment, the index construction module is further used to divide the page identifiers corresponding to the determined page codes into blocks to obtain a plurality of page identifier blocks; connect the plurality of page identifier blocks to generate a page identifier chain; and establish the mapping relationship between the positions and the page identifier chain to generate an index.

[0022] In one embodiment, the device further includes:

[0023] An update module, which is used to extract features and perform feature dimensionality reduction on the updated page code when the page code in the page code library is updated; update the page features after dimensionality reduction of the updated page code to the corresponding positions in the bitmap; in the index, establish a mapping relationship between the updated positions and the page identifiers corresponding to the updated page code.

[0024] In one embodiment, the index query module is further configured to perform an inverted index query based on a pre-constructed index, using the located position as an index entry, to determine a page identifier chain corresponding to the located position; starting from the first page identifier block in the page identifier chain, read the page identifier in the page identifier block, and for the next page identifier block on the page identifier chain, iteratively execute the step of reading the page identifier in the page identifier block until the iteration stops after reading the page identifier in the last page identifier block on the page identifier chain; wherein, the page identifiers in each of the page identifier blocks on the page identifier chain are the page identifiers corresponding to the page code having the page features recorded at the located position.

[0025] In one embodiment, the feature dimensionality reduction module is further configured to respectively convert the feature dimension and the range of the feature values of the feature values to a preset dimension range and a preset feature value range to obtain the page features after dimensionality reduction.

[0026] In one embodiment, the feature dimensionality reduction module is further configured to respectively perform a hash operation on each feature value in the original features of the target page according to at least one hash function to obtain the hash values of the feature values in each hash operation; respectively select the feature values corresponding to the minimum hash value obtained in each hash operation from the feature values in multiple dimensions; and determine the page features after dimensionality reduction of the original features of the target page according to the selected feature values.

[0027] In one embodiment, the target page is a sub-application page; the page identifier is the identifier of the sub-application page; the sub-application page is a page provided by the sub-application; the sub-application is a lightweight application running in the environment provided by the original parent application.

[0028] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0029] Obtain the original features of the target page extracted from the page code of the target page;

[0030] Perform feature dimensionality reduction processing on the original features of the target page to obtain the page features after dimensionality reduction;

[0031] Query the page identifier corresponding to the page feature to obtain a suspected similar page identifier;

[0032] Obtain the original page features corresponding to each of the suspected similar page identifiers;

[0033] Match each of the original page features with the original target page feature respectively;

[0034] Determine the page corresponding to the original page feature that passes the match as the page similar to the target page.

[0035] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0036] Obtain the original target page feature extracted from the page code of the target page;

[0037] Perform feature dimensionality reduction processing on the original target page feature to obtain the dimensionality-reduced page feature;

[0038] Query the page identifier corresponding to the page feature to obtain a suspected similar page identifier;

[0039] Obtain the original page features corresponding to each of the suspected similar page identifiers;

[0040] Match each of the original page features with the original target page feature respectively;

[0041] Determine the page corresponding to the original page feature that passes the match as the page similar to the target page.

[0042] For the above page recognition method, device, computer device and storage medium, perform feature dimensionality reduction processing on the original target page feature of the target page, and query the suspected similar page identifier according to the dimensionality-reduced page feature. Since the dimensionality-reduced page feature greatly reduces the complexity of the page feature, directly performing an inverted index according to the dimensionality-reduced page feature can quickly query the suspected similar page identifier to achieve rough matching and screening. Then, match the original page features corresponding to each suspected similar page identifier with the original target page feature, that is, through the fine matching process of the original features, perform higher-precision matching and screening on the suspected similar pages screened by the rough matching, and determine the page corresponding to the original page feature that passes the match as the page similar to the target page. Through the rough matching of dimensionality reduction and inverted index query, combined with the precision matching of the original page features, it is possible to quickly and accurately identify the page similar to the target page from a large number of pages. Brief Description of the Drawings

[0043] Figure 1It is an application environment diagram of the page recognition method in an embodiment;

[0044] Figure 2 It is a schematic flowchart of the page recognition method in an embodiment;

[0045] Figure 3 It is a schematic diagram of feature dimensionality reduction in an embodiment;

[0046] Figure 4 It is a schematic diagram of building an index in an embodiment;

[0047] Figure 5 It is a system architecture diagram in an embodiment;

[0048] Figure 6 It is a simple schematic diagram of the feature update and matching process in an embodiment;

[0049] Figure 7 It is a block diagram of the page recognition device in an embodiment;

[0050] Figure 8 It is a block diagram of the page recognition device in another embodiment;

[0051] Figure 9 It is a block diagram of the page recognition device in yet another embodiment;

[0052] Figure 10 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] The page recognition method provided by the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0055] A technician can select a target page through the terminal 102, and the terminal 102 notifies the server 104 of the selected target page. The server 104 can then obtain the original target page features extracted from the page code of the target page; perform feature dimensionality reduction processing on the original target page features to obtain the dimensionality-reduced page features; query the page identifiers corresponding to the page features to obtain suspected similar page identifiers; obtain the original page features corresponding to each of the suspected similar page identifiers; respectively match each of the original page features with the original target page features; and determine the page corresponding to the original page feature that passes the match as a page similar to the target page. Further, the server 104 can notify the terminal 102 of the identified page.

[0056] In one embodiment, as Figure 2 shown, a page recognition method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:

[0057] Step 202, obtain the original target page features extracted from the page code of the target page.

[0058] Among them, the target page is the reference page for determining whether there are similar pages. That is, based on the target page, similar pages to this target page are queried from the page code library.

[0059] In one embodiment, the target page can be the page to be determined whether it is plagiarized. It can be understood that in the application scenario of identifying plagiarized pages, if there are similar pages to the target page, it means that the target page has been plagiarized. The page similar to the target page is the abnormal page suspected of plagiarism.

[0060] In one embodiment, the target page can be a sub-application page. A sub-application page is a page provided by a sub-application. A sub-application is a lightweight application that can be implemented in the environment provided by the parent application and can be used without downloading and installing. The parent application is the application program that hosts the sub-application and provides an environment for the implementation of the sub-application. The parent application is a native application program. A native application program is an application program that can be directly run on the operating system.

[0061] The original target page features are the original page features of the target page. The page original features refer to the page features directly obtained after feature extraction of the page code and without feature dimensionality reduction processing. It can be understood that the original target page features are the page features extracted from the page code of the target page and without feature dimensionality reduction processing.

[0062] In one embodiment, the server can directly obtain the original features of the target page from the memory or the storage system. Specifically, the server can pre-process the feature extraction of each page code in the code library to obtain the original features of each page code, and store the original features of each page code in the local memory or send them to the storage system for storage. In this way, the server can directly obtain the original features of the target page corresponding to the target page from the memory or the storage system.

[0063] In one embodiment, the server can also perform feature extraction processing on the page code of the target page to extract the original features of the target page.

[0064] Step 204, perform feature dimensionality reduction processing on the original features of the target page to obtain the page features after dimensionality reduction.

[0065] Among them, the page features after dimensionality reduction are the page features after reducing the feature dimension of the original features of the target page to a preset dimension range.

[0066] In one embodiment, the page features after dimensionality reduction are the page features after reducing the feature dimension of the original features of the target page to a preset dimension range and shrinking the magnitude of the feature values to a preset feature value range. That is, compared with the original features of the target page, the page features after dimensionality reduction have a reduced feature dimension and a reduced range of feature values.

[0067] In one embodiment, each of the page features after dimensionality reduction is a fixed-length page feature. That is, the page features after dimensionality reduction meet the preset length.

[0068] Specifically, the server can use a linear or non-linear function to perform feature dimensionality reduction processing on the original features of the target page.

[0069] In one embodiment, the original features of the target page have feature values with multiple feature dimensions. In this embodiment, performing feature dimensionality reduction processing on the original features of the target page to obtain the page features after dimensionality reduction includes: respectively converting the feature dimension and the range of the feature values into a preset dimension range and a preset feature value range to obtain the page features after dimensionality reduction.

[0070] In one embodiment, the server can use the minhash algorithm to perform feature dimensionality reduction processing on the original features of the target page.

[0071] In one embodiment, the feature dimension and the range of the feature values of the feature values are respectively converted into a preset dimension range and a preset feature value range, and the page features after dimensionality reduction are obtained, including: respectively performing a hash operation on each feature value in the original features of the target page according to at least one hash function to obtain the hash value of each feature value in each hash operation; respectively selecting the feature values corresponding to the minimum hash value obtained in each hash operation from the feature values of multiple dimensions; and determining the page features after dimensionality reduction of the original features of the target page according to the selected feature values.

[0072] It can be understood that when there are multiple hash functions, the feature values corresponding to the minimum hash value obtained by performing a hash operation using each hash function are used as the page features after dimensionality reduction of the original features of the target page.

[0073] When there is a single hash function, a hash operation is performed on each feature value in the original features of the target page using this hash function to obtain hash values, and a preset number of feature values are selected in ascending order of the hash values as the page features after dimensionality reduction of the original features of the target page.

[0074] Step 206: Query the page identifier corresponding to the page features to obtain the suspected similar page identifier.

[0075] In one embodiment, the server may use the page features after dimensionality reduction as an index item to query the page identifier corresponding to the page features after dimensionality reduction to obtain the suspected similar page identifier.

[0076] In one embodiment, step 206 includes: using the page features as an index item, and querying at least one page identifier corresponding to the page features after dimensionality reduction according to the pre-set mapping relationship between the page features and the page identifiers to obtain the suspected similar page identifier.

[0077] Among them, the suspected similar page identifier is the identifier of the suspected similar page. It can be understood that the suspected similar page is the candidate page, and subsequently, the final similar page needs to be identified from the candidate pages through steps 208 to 212.

[0078] In one embodiment, the page identifier may be the identifier of the sub-application page. The suspected similar page identifier may be the identifier of the suspected similar sub-application page.

[0079] Specifically, the server pre-builds a mapping relationship between the reduced-dimensional page features and the page identifiers. Based on the mapping relationship, the server can use the reduced-dimensional page features as index items to query the page identifiers corresponding to the reduced-dimensional page features to obtain suspected similar page identifiers. It should be noted that the mapping relationship between the page features and the page identifiers can be direct or indirect (for example, a mapping relationship is established between the position of the page features in the bitmap and the page identifier).

[0080] The same reduced-dimensional page feature has a mapping relationship with one or more page identifiers. It can be understood that the original page features of the same or similar page codes will be the same or similar, and the reduced-dimensional page features will also be the same or similar. Therefore, the page codes corresponding to the page identifiers that have a mapping relationship with the same page feature are similar to a large extent.

[0081] It is understandable that, since the content length of each page code may be different (for example, the content length of the mini-program page code is basically different), the number of feature dimensions and the range of feature values ​​in the extracted original features of the page are uncontrollable. Such actual data is very unfavorable for query (for example, it is not conducive to index construction) and will consume performance and a lot of query time. Therefore, the server can perform feature dimensionality reduction processing on the original features of each page code in the code library, so as to convert the number of feature dimensions and the range of feature values ​​into a controllable range (that is, within the preset dimension range and the preset feature value range). Then, a mapping relationship is established based on the reduced-dimensional page features and page identifiers corresponding to each page code in the code library. Subsequently, based on the mapping relationship, the search complexity can be reduced and the search efficiency can be improved.

[0082] In other embodiments, the server may also match the reduced-dimensionality page features corresponding to the target page with the reduced-dimensionality page features corresponding to each page in the page code library, and determine the page identifier corresponding to the matched reduced-dimensionality page features as the suspected similar page identifier. It can be understood that since the feature dimension and feature value range of the reduced-dimensionality page features are in a controllable and unified range, matching the reduced-dimensionality page features can also improve the matching efficiency, thereby improving the page recognition efficiency, compared to the traditional method of directly matching according to the original features of different pages of unequal lengths.

[0083] Step 208: Obtain the original features of the pages corresponding to the suspected similar page identifiers.

[0084] It can be understood that the original features of the page corresponding to the suspected similar page identifier refer to the original features of the page code corresponding to the suspected similar page identifier.

[0085] In one embodiment, the server can directly obtain the original page features corresponding to each suspected similar page identifier from the memory. It can be understood that the server first extracts the original page features of each page code in the code library, and then performs feature dimensionality reduction processing on each original page feature, so as to construct a mapping relationship between the dimensionality-reduced page features and the page identifiers. Therefore, during the process of constructing the mapping relationship between the dimensionality-reduced page features and the page identifiers, the server has already stored the original page features of each page code in the memory, and thus can directly obtain the original page features corresponding to each suspected similar page identifier from the memory.

[0086] In another embodiment, the server can obtain the original page features corresponding to each suspected similar page identifier from the storage system. It can be understood that after the server extracts the original page features of each page code, it can store each original page feature in the storage system. In this case, the server can also obtain the original page features corresponding to each suspected similar page identifier from the storage system.

[0087] In other embodiments, the server can also perform feature extraction processing on the page codes corresponding to each suspected similar page identifier to extract the original page features corresponding to each suspected similar page identifier. The method for obtaining the original page features corresponding to the suspected similar page identifiers is not limited here.

[0088] Step 210: Match each original page feature with the original target page feature respectively.

[0089] It should be noted that each original page feature in step 210 refers to the original page feature corresponding to the suspected similar page identifier screened and queried.

[0090] Compared with the dimensionality-reduced page features, the original page features have more dimensions and a wider range of feature information, that is, they have more comprehensive feature information. Therefore, matching the original page features corresponding to the suspected similar page identifiers with the original target page features is equivalent to a matching process with higher precision and belongs to a fine matching process. Fine matching means matching between all and complete page features.

[0091] It can be understood that finding the suspected similar page identifiers according to the dimensionality-reduced page features belongs to a rough matching process. It is equivalent to first performing rough matching through steps 202 to 206 to initially screen out the suspected similar pages, and then performing fine matching through steps 208 to 212. Through fine matching, the final pages similar to the target page are screened out from the suspected similar pages.

[0092] Specifically, the server can match the original page features corresponding to each suspected similar page identifier with the original target page features one by one. It can be understood that since the suspected similar page identifiers are obtained through preliminary screening queries, which is equivalent to having gone through a round of screening by rough matching, then, performing a one-to-one sequential matching of the original page features corresponding to the suspected similar page identifiers with the original target page features will not result in problems of excessive data volume and excessive time consumption.

[0093] Step 212: Determine the page corresponding to the original page feature that passes the matching as the page similar to the target page.

[0094] Specifically, in step 206, the matching results between the original page features corresponding to each suspected similar page identifier and the original target page features can include two results: passing the matching and not passing the matching. Therefore, when the original page feature passes the matching with the original target page feature, it indicates that the page corresponding to the original page feature is similar to the target page, and then this page is determined as the page similar to the target page. When the matching fails, it indicates that the page corresponding to the original page feature is not similar enough to the target page, and then it can be determined that it is not the page similar to the target page.

[0095] The above page recognition method performs dimensionality reduction processing on the original target page features of the target page, and queries the suspected similar page identifiers according to the dimensionality-reduced page features. Since the dimensionality-reduced page features greatly reduce the complexity of the page features, directly performing an inverted index based on the dimensionality-reduced page features can quickly query the suspected similar page identifiers to achieve rough matching screening. Then, match the original page features corresponding to each suspected similar page identifier with the original target page features, that is, through the fine matching process of the original features, perform a higher-precision matching screening on the suspected similar pages screened by the rough matching, and determine the page corresponding to the original page feature that passes the matching as the page similar to the target page. Through the rough matching of dimensionality reduction and inverted index query, combined with the precision matching of the original page features, it is possible to quickly and accurately identify the page similar to the target page from a large number of pages.

[0096] In one embodiment, taking the page feature as the index item, according to the pre-set mapping relationship between the page feature and the page identifier, query at least one page identifier corresponding to the dimensionality-reduced page feature, and the obtained suspected similar page identifier includes: locating the position corresponding to the page feature in the pre-set bitmap; where each position in the bitmap uniquely records a page feature; based on the pre-constructed index, perform an inverted index query with the located position as the index item to obtain at least one page identifier corresponding to the located position, and obtain the suspected similar page identifier.

[0097] Among them, a bitmap, that is, Bitmap, is a data structure representing a dense set in a finite field. Each element appears at least once, and no other data is associated with the elements. It can be understood that in the embodiments of the present application, each position in the bitmap uniquely records the dimension-reduced page features corresponding to each page feature in the code library. Different dimension-reduced page features are recorded at different positions in the bitmap.

[0098] It can be understood that the server has pre-established an index, which includes the mapping relationship between the positions in the bitmap and the page identifiers. It can be understood that the page identifier having a mapping relationship with the position refers to the page identifier corresponding to the page code having the page feature recorded at that position. That is, there is a mapping relationship between the page identifiers corresponding to the page codes recording the same page feature and the position where the page feature is located in the bitmap.

[0099] Then, the server can locate the position corresponding to the dimension-reduced page feature in the pre-established bitmap. The server can use the located position as an index item to perform an inverted index query in the pre-constructed index to find at least one page identifier corresponding to the located position, and obtain suspected similar page identifiers.

[0100] In one embodiment, the server can compare the page features recorded at each position in the bitmap with the dimension-reduced page features corresponding to the target page, determine the collision rate between the page features recorded at each position and the dimension-reduced page features corresponding to the target page, and select the position corresponding to the page feature with the maximum collision rate as the position corresponding to the dimension-reduced page features of the target page in the bitmap.

[0101] Now, in combination with Figure 3 an example will be given for illustration. Figure 3 Taking the dimension reduction by the min-hash algorithm as an example for illustration. Referring to Figure 3 , what is shown in region 302 is the traditional method of directly comparing the similarity of page original features. It can be seen from 302 that the page original features in the traditional method are all variable-length fingerprints. To compare the similarity between page A and page B, it is necessary to calculate the similarity between the variable-length fingerprint a of page A and the variable-length fingerprint b of page B through the Jaccard similarity algorithm. However, in the solution of the present application, after the page original features are dimension-reduced by the min-hash algorithm, the collision rate of the dimension-reduced page features is calculated. Referring to Figure 3, the fixed-length minimum hash fingerprint a’ is the page feature after dimensionality reduction of the variable-length fingerprint a, and the fixed-length minimum hash fingerprint b’ is the page feature after dimensionality reduction of the variable-length fingerprint b. Assume that the target page is page A, and the bitmap records the page feature after dimensionality reduction of page B. Then, the fixed-length minimum hash fingerprint a’ corresponding to page A after dimensionality reduction (i.e., the page feature after dimensionality reduction corresponding to page A) can be compared with the fixed-length minimum hash fingerprint b’ corresponding to page B after dimensionality reduction recorded in the bitmap (i.e., the page feature after dimensionality reduction corresponding to page B), and the collision rate between the two can be calculated. In this way, after calculating the collision rate between the fixed-length minimum hash fingerprint a’ and the fixed-length minimum hash fingerprint recorded in the bitmap, the position corresponding to the page feature after dimensionality reduction of the target page A in the bitmap can be determined. That is, the minimum hash fingerprint with the highest collision rate is the most similar to the fixed-length minimum hash fingerprint a’. Therefore, the position where the minimum hash fingerprint with the highest collision rate is located is the position corresponding to the page feature after dimensionality reduction of the target page A in the bitmap.

[0102] In the above embodiment, by searching through the bitmap structure, the search time complexity can be stabilized at O(1), thereby improving the search query efficiency.

[0103] In one embodiment, the method further includes the following steps: extracting features from the page codes in the page code library; performing dimensionality reduction processing on the features of each extracted page code; recording each page feature after dimensionality reduction at the corresponding position in the bitmap; for the position in the bitmap, determining the page code with the page feature recorded at this position, and establishing a mapping relationship between this position and the page identifier corresponding to the determined page code to generate an index.

[0104] Among them, the page identifier corresponding to the page code refers to the unique identifier of the page represented by this page code.

[0105] Specifically, the server can determine, for each position in the bitmap, the page code with the page feature after dimensionality reduction recorded at this position, that is, determine the page code from which the original page feature before dimensionality reduction of the page feature recorded at this position is extracted. It can be understood that there can be multiple groups of page codes with the page feature recorded at this position. A group of page codes represents a page. The server can establish a mapping relationship between this position and the page identifier corresponding to the determined page code to generate an index. That is, the same position can correspond to one or more (more than one means at least two) page identifiers.

[0106] It should be noted that the page identifiers corresponding to the same position in the bitmap can be located in the same block or in different blocks, and a chained structure is formed between different blocks. That is, the same position in the bitmap can correspond to a block with all corresponding page identifiers. The same position in the bitmap can also correspond to a page identifier chain. Each block on the page identifier chain has a page identifier corresponding to this position, and the set of page identifiers in each block on the page identifier chain is all the page identifiers corresponding to this position.

[0107] In the above embodiment, the dimension-reduced page features corresponding to each page code in the page code library are recorded through the bitmap structure, and subsequent fast search can be performed based on this bitmap structure, thereby improving the search query efficiency.

[0108] In one embodiment, establishing a mapping relationship between the position and the page identifier corresponding to the determined page code, generating an index includes: dividing the page identifiers corresponding to the determined page code into blocks to obtain multiple page identifier blocks; connecting the multiple page identifier blocks to generate a page identifier chain; establishing a mapping relationship between the position and the page identifier chain to generate an index.

[0109] Among them, a page identifier block is a block structure for storing page identifiers. A page identifier chain is a chain structure including multiple page identifier blocks. The sum of the page identifiers in each page identifier block on the page identifier chain is all the page identifiers corresponding to this position. The page identifier corresponding to this position refers to the page identifier corresponding to the page code with the page features recorded at this position. That is, the page identifiers in each page identifier block are the page identifiers corresponding to the page code with the page features located at the located position.

[0110] In one embodiment, the same page identifier block has the same or different numbers of page identifiers.

[0111] In one embodiment, the server can establish a mapping relationship between the information headers corresponding to each position in the bitmap and the corresponding page identifier chain to generate an index.

[0112] Figure 4 It is a schematic diagram for constructing an index in one embodiment. Refer to Figure 4 , different dimension-reduced page features are recorded at different positions in the bitmap. The information headers of each position respectively have corresponding page identifier chains. Each page identifier chain has multiple page identifier blocks. Each page identifier block has multiple page identifiers. Taking the first position 402 as an example, assume that the first position corresponds to 100 page identifiers, which are divided into 2 page identifier blocks, and each page identifier block includes 50 page identifiers. It should be noted that Figure 4Different information headers correspond to different page identifiers. The numbers in pages 1... page N+50 in the figure are only used to represent the number of pages. For example, page 1 corresponding to the information headers of 8 and 6 in the bitmap does not represent the same page.

[0113] In one embodiment, based on the pre-constructed index, an inverted index query is performed with the located position as the index item to obtain at least one page identifier corresponding to the located position. Obtaining suspected similar page identifiers includes: based on the pre-constructed index, performing an inverted index query with the located position as the index item to determine a page identifier chain corresponding to the located position; starting from the first page identifier block in the page identifier chain, reading the page identifiers in the page identifier block, and for the next page identifier block on the page identifier chain, iteratively execute the step of reading the page identifiers in the page identifier block until the page identifiers in the last page identifier block on the page identifier chain are read and then stop the iteration.

[0114] Specifically, the server can perform an inverted index query based on the pre-constructed index with the located position as the index item to determine a page identifier chain corresponding to the located position.

[0115] In one embodiment, the server can query the information header corresponding to the located position based on the pre-constructed index and determine the information body corresponding to the information header. Among them, the information body includes a page identifier chain.

[0116] The server can start from the first page identifier block in the page identifier chain, read the page identifiers in the page identifier block, and after reading the page identifiers in a page identifier block, continue to read the page identifiers in the next page identifier block on the page identifier chain, and iteratively read the page identifiers in each page identifier block on the page identifier chain in this way, so as to obtain all the page identifiers corresponding to this position.

[0117] Similarly combined with Figure 4 for illustration. Referring to Figure 4 , assuming that the located position is 402, then, the information header corresponding to position 402 can be queried first, and then the information body corresponding to the information header can be determined, and then the page identifier chain L corresponding to position 402 can be determined. Then, 50 page identifiers in the first page identifier block on the page identifier chain L can be read first, and then 50 page identifiers in the second page identifier block can be read.

[0118] In the above embodiment, the constructed index structure is a block-chain structure. Therefore, excessive jumps are avoided during querying, so that continuous data copying can be performed more efficiently, that is, page identifiers can be obtained more efficiently.

[0119] In one embodiment, the method further includes: when the page code in the page code library is updated, feature extraction and feature dimensionality reduction processing are performed on the updated page code; the page features obtained by dimensionality reduction of the updated page code are updated to the corresponding positions in the bitmap; in the index, a mapping relationship is established between the updated positions and the page identifiers corresponding to the updated page code.

[0120] It can be understood that updating to the corresponding positions in the bitmap means recording the page features obtained by dimensionality reduction of the updated page code in the bitmap. It can be understood that when the page features obtained by dimensionality reduction change, the corresponding positions in the bitmap also change.

[0121] Now in combination with Figure 5 the architecture diagram therein for illustration. Referring to Figure 5 , the server includes a fine matching unit and a coarse matching unit. Among them, the fine matching unit is used for feature extraction (i.e., extracting the original page features), feature storage (i.e., storing the original page features in the memory), and fine matching (i.e., matching between the original page features). The coarse matching unit is used for dimensionality reduction of features and index construction based on the page features obtained by dimensionality reduction (i.e., constructing an index between the page features obtained by dimensionality reduction and the page identifiers). The coarse matching unit is also used for feature update (i.e., updating the page features obtained by dimensionality reduction) and coarse matching processing (i.e., the processing of finding suspected similar page identifiers according to the page features obtained by dimensionality reduction). The fine matching unit and the coarse matching unit in the server can respectively interact with the storage system.

[0122] Now in combination with Figure 5 the structure diagram of to briefly describe the entire page recognition and processing process. The fine matching unit can perform feature extraction on each page code in the page code library and store the extracted original page features in its memory. The fine matching unit can input the extracted original page features into the coarse matching unit, so that the coarse matching unit can perform dimensionality reduction processing on each original page feature and construct an index according to the page features obtained by dimensionality reduction to establish an index mapping relationship between the page features obtained by dimensionality reduction and the corresponding page identifiers. In addition, the fine matching unit in the server can also send the extracted original page features to the storage system, so that the storage system performs permanent storage on the original page features.

[0123] When the page code in the page code library of the fine matching unit is updated, in addition to updating the page original features of the updated page code in its own memory and sending the updated page original features to the storage system, the fine matching unit can also send an update request to the coarse matching unit, so that the coarse matching unit can reduce the dimension of the updated page original features, thereby obtaining the updated dimensionality-reduced page features. The coarse matching unit can also update the constructed index according to the updated dimensionality-reduced page features (for example, re-determine the position corresponding to the updated dimensionality-reduced page features in the bitmap, and establish the mapping relationship between the re-determined position and the page identifier to achieve the purpose of updating the index).

[0124] When the terminal notifies the target page, the fine matching unit in the server can obtain the target page original features of the target page from its own memory and send a matching request to the coarse matching unit, so that after the coarse matching unit reduces the dimension of the target page original features, it obtains the dimensionality-reduced page features corresponding to the target page. The coarse matching unit can perform coarse matching processing according to the dimensionality-reduced page features corresponding to the target page, that is, perform an inverted index query based on the constructed index, and query the suspected similar page identifiers corresponding to the dimensionality-reduced page features. The coarse matching unit can pass the suspected similar page identifiers found by the coarse matching to the fine matching unit, so that the fine matching unit can find the page original features corresponding to the suspected similar page identifiers in the memory, and sequentially match the found page original features with the target page original features one by one to achieve fine matching processing. Furthermore, according to the fine matching result, the pages finally similar to the target page are screened out from the suspected similar pages corresponding to the suspected similar page identifiers.

[0125] It can be understood that in the case of the entire server restart or a large-scale change in the version that requires re-constructing the index, the coarse matching unit can pull the stored page original features from the storage system and perform dimensionality reduction processing on them to reconstruct the index. When the data does not change significantly, the coarse matching unit can update the index only according to the update request of the fine matching unit.

[0126] In the above embodiments, the index structure is updated according to the feature update of the page code, which improves the accuracy of the constructed index, and the subsequent query processing based on this index is also more accurate.

[0127] Figure 6 It is a simple schematic diagram of the feature update and matching process in an embodiment. Refer to Figure 6, during the update process, the fine matching unit is responsible for extracting the original page features from the page codes with updates in the page code library to implement feature extraction processing, and updating and storing the original page features to implement feature storage processing. The fine matching unit passes the extracted original page features to the rough matching unit, so that after the rough matching unit reduces their dimensions, it updates the index and stores them. During the matching process, the fine matching unit extracts features from the page code of the target page, and passes the extracted original features of the target page to the rough matching unit. After the rough matching unit reduces their dimensions, it performs inverted index matching (i.e., implements rough matching) to find out the suspected similar page identifiers corresponding to the page features after dimension reduction. The rough matching unit passes the found suspected similar page identifiers to the fine matching unit, so that the fine matching unit pulls the original page features corresponding to the found suspected similar page identifiers from the memory, and matches the found original page features with the original features of the target page of the target page, that is, implements full feature matching (which also belongs to fine matching), so as to identify the page finally similar to the target page according to the matching result.

[0128] This application also provides an application scenario, which applies the above page recognition method. Specifically, the application of the page recognition method in this application scenario is as follows:

[0129] The server has pre-extracted features for the applet pages in the applet code library, and performed feature dimension reduction processing on the extracted original features of the applet pages to generate the applet page features after dimension reduction. The server records the applet page features after dimension reduction at the corresponding positions in the bitmap, and establishes a mapping relationship between each position in the bitmap and the applet page identifier, thereby constructing an index. The applet page identifier that has a mapping relationship with the position is the applet page identifier corresponding to the applet page code with the applet page features recorded at that position.

[0130] The technician selects a target mini-program page (the mini-program page is the sub-application page) to be identified for plagiarism, and the terminal notifies the server of the selected target mini-program page. The server extracts features from the target mini-program page to obtain the original features of the target page, and performs dimensionality reduction processing on the original features of the target page. The server can determine the position corresponding to the dimensionality-reduced mini-program page features in the bitmap, and then use the determined position as an index term to query the mini-program page identifier corresponding to this position from the pre-constructed index, obtaining a suspected similar mini-program page identifier. The server can obtain the original page features corresponding to each suspected similar mini-program page identifier, and match the obtained original page features with the original features of the target page of the target mini-program page. The server can determine the mini-program page corresponding to the original page features that pass the match as a mini-program page similar to the target mini-program page. Thus, through rough matching (i.e., dimensionality reduction and inverted index query) and fine matching (i.e., matching between the original page features), the abnormal mini-program pages (i.e., the mini-program pages similar to the target mini-program page) suspected of plagiarizing the target mini-program page can be quickly and accurately identified from the mini-program code library.

[0131] In the above embodiment, since the number of mini-program pages in the mini-program page code library is in the hundreds of millions, and moreover, the content lengths of the codes of each mini-program page are different, so using the traditional method to perform one-to-one matching of the mini-program codes in the library is very time-consuming, with low efficiency, and a very large consumption of system resources. However, the page recognition method in the embodiments of the present application can quickly screen out suspected similar mini-program pages through dimensionality reduction and inverted index query. Furthermore, the original page features of the suspected similar mini-program pages are finely matched with the original features of the target page of the target mini-program page, so that the final abnormal mini-program pages suspected of plagiarism can be quickly and accurately screened out, greatly saving the consumption of system resources.

[0132] It can be understood that the page recognition method in the embodiments of the present application can also be applied in scenarios other than mini-program pages. That is, it can be applicable to any application scenario for identifying similar pages. For example, identifying similar pages on the web side from the code library.

[0133] It should be understood that although the steps in the flowchart of the present application are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0134] In one embodiment, as Figure 7 shown, a page recognition device is provided. This device can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the device includes: a feature acquisition module 702, a feature dimensionality reduction module 704, an index query module 706, and a matching module 708, where:

[0135] The feature acquisition module 702 is configured to acquire the original features of the target page extracted from the page code of the target page.

[0136] The feature dimensionality reduction module 704 is configured to perform feature dimensionality reduction processing on the original features of the target page to obtain the dimensionality-reduced page features.

[0137] The index query module 706 is configured to query the page identifier corresponding to the page features to obtain the suspected similar page identifiers.

[0138] The matching module 708 is configured to acquire the original features of the pages corresponding to each suspected similar page identifier; respectively match each original page feature with the original features of the target page; and determine the page corresponding to the original page feature that passes the match as the page similar to the target page.

[0139] In one embodiment, the index query module 706 is further configured to use the page features as index items, and query at least one page identifier corresponding to the dimensionality-reduced page features according to the pre-set mapping relationship between the page features and the page identifiers to obtain the suspected similar page identifiers.

[0140] In one embodiment, the index query module 706 is further configured to locate the position corresponding to the page feature in a preset bitmap; wherein, each position in the bitmap uniquely records a page feature; based on the pre-constructed index, perform an inverted index query with the located position as the index item to obtain at least one page identifier corresponding to the located position, and obtain suspected similar page identifiers; wherein, the index includes the mapping relationship between the position in the bitmap and the page identifier; the page identifier having a mapping relationship with the position refers to the page identifier corresponding to the page code having the page feature recorded by the position.

[0141] As Figure 8 shown, in one embodiment, the apparatus further includes:

[0142] An index construction module 701, configured to extract features from the page codes in the page code library; perform feature dimensionality reduction processing on the extracted features of each page code; record the dimensionality-reduced each page feature at the corresponding position in the bitmap; for the position in the bitmap, determine the page code having the page feature recorded by the position, and establish the mapping relationship between the position and the page identifier corresponding to the determined page code, and generate an index.

[0143] In one embodiment, the index construction module 701 is further configured to block the page identifiers corresponding to the determined page codes to obtain a plurality of page identifier blocks; connect the plurality of page identifier blocks to generate a page identifier chain; establish the mapping relationship between the position and the page identifier chain, and generate an index.

[0144] As Figure 9 shown, in one embodiment, the apparatus further includes:

[0145] An update module 710, configured to, when the page codes in the page code library are updated, perform feature extraction and feature dimensionality reduction processing on the updated page codes; update the dimensionality-reduced page features of the updated page codes to the corresponding positions in the bitmap; in the index, establish the mapping relationship between the updated positions and the page identifiers corresponding to the updated page codes.

[0146] In one embodiment, the index query module 706 is further configured to, based on the pre-constructed index, perform an inverted index query with the located position as the index item to determine the page identifier chain corresponding to the located position; starting from the first page identifier block in the page identifier chain, read the page identifiers in the page identifier block, and for the next page identifier block on the page identifier chain, iteratively execute the step of reading the page identifiers in the page identifier block until the iteration stops after reading the page identifiers in the last page identifier block on the page identifier chain; wherein, the page identifiers in each page identifier block on the page identifier chain are the page identifiers corresponding to the page codes having the page feature recorded by the located position.

[0147] In one embodiment, the feature dimensionality reduction module 704 is further configured to convert the feature dimension and the range of the feature values of the feature values into a preset dimension range and a preset feature value range respectively, so as to obtain the page features after dimensionality reduction.

[0148] In one embodiment, the feature dimensionality reduction module 704 is further configured to perform a hashing operation on each feature value in the original features of the target page according to at least one hash function respectively, so as to obtain the hash value of each feature value in each hashing operation; select the feature values corresponding to the minimum hash value obtained by each hashing operation respectively from the feature values in multiple dimensions; and determine the page features after dimensionality reduction of the original features of the target page according to the selected feature values.

[0149] In one embodiment, the target page is a sub-application page; the page identifier is the identifier of the sub-application page; the sub-application page is a page provided by the sub-application; and the sub-application is a lightweight application running in the environment provided by the original parent application.

[0150] For the specific limitations on the page recognition device, reference may be made to the limitations on the page recognition method in the foregoing text, which will not be elaborated here. Each module in the foregoing page recognition device may be implemented in whole or in part by software, hardware, and their combination. The foregoing modules may be embedded in the processor in the computer device in the form of hardware or independent of the processor, or may be stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the foregoing modules.

[0151] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 10 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store page recognition data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a page recognition method is implemented.

[0152] Those skilled in the art can understand that Figure 10 the structure shown in is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0153] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0154] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0155] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0156] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0158] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A page recognition method, characterized in that, The method includes: Obtaining the original features of the target page extracted from the page code of the target page; Performing a hashing operation on each feature value in the original features of the target page using at least one hashing function to obtain the hash value of each feature value in each hashing operation. When the at least one hashing function is multiple hashing functions, the feature value corresponding to the minimum hash value obtained by performing the hashing operation using each hashing function is used as the page feature after dimensionality reduction. When the at least one hashing function is a single hashing function, a preset number of feature values are selected in ascending order of the hash values as the page feature after dimensionality reduction; Locating the position corresponding to the page feature in a preset bitmap, where each position in the bitmap uniquely records a page feature; Based on a pre-constructed index, performing an inverted index query using the located position as the index item to obtain a page identifier chain corresponding to the located position, and reading the page identifiers in each page identifier block in the page identifier chain to obtain each suspected similar page identifier. Each page identifier chain has multiple page identifier blocks, and each page identifier block has multiple page identifiers. The index includes the mapping relationship between the position in the bitmap and the page identifier. The page identifier having a mapping relationship with the position refers to the page identifier corresponding to the page code having the page feature recorded at the position; Obtaining the original features of the pages corresponding to each of the suspected similar page identifiers; Respectively matching each of the original page features with the original features of the target page; Determining the page corresponding to the original page feature that passes the matching as the page similar to the target page.

2. The method according to claim 1, characterized in that, The method further includes: Extracting features from the page codes in the page code library; Performing dimensionality reduction processing on the features of each extracted page code; Recording the dimensionality-reduced page feature at the corresponding position in the bitmap; For the position in the bitmap, determining the page code having the page feature recorded at the position, and establishing a mapping relationship between the position and the page identifier corresponding to the determined page code to generate an index.

3. The method according to claim 2, wherein The establishing the mapping relationship between the position and the page identifier corresponding to the determined page code to generate an index includes: Dividing the page identifiers corresponding to the determined page code into blocks to obtain multiple page identifier blocks; Connecting the multiple page identifier blocks to generate a page identifier chain; Establishing a mapping relationship between the position and the page identifier chain to generate an index.

4. The method according to claim 3, wherein The reading the page identifiers in each page identifier block in the page identifier chain to obtain each suspected similar page identifier includes: Starting from the first page identifier block in the page identifier chain, reading the page identifiers in the page identifier block, and for the next page identifier block on the page identifier chain, iteratively executing the step of reading the page identifiers in the page identifier block until the page identifiers in the last page identifier block on the page identifier chain are read and the iteration stops; Among them, the page identifiers in each of the page identifier blocks on the page identifier chain are page identifiers corresponding to page codes with page features recorded at the located positions.

5. The method according to claim 2, wherein The method further includes: When the page codes in the page code library are updated, then perform feature extraction and feature dimensionality reduction processing on the updated page codes; update the page features after dimensionality reduction of the updated page codes to the corresponding positions in the bitmap; in the index, establish a mapping relationship between the updated position and the page identifier corresponding to the updated page code.

6. The method according to claim 1, wherein The original features of the target page have eigenvalue of multiple feature dimensions; The performing feature dimensionality reduction processing on the original features of the target page to obtain the page features after dimensionality reduction includes: respectively convert the feature dimension and the range of the eigenvalue to a preset dimension range and a preset eigenvalue range to obtain the page features after dimensionality reduction.

7. The method according to any one of claims 1 to 6, characterized in that, The target page is a sub-application page; the page identifier is the identifier of the sub-application page; the sub-application page is a page provided by the sub-application; the sub-application is a lightweight application running in the environment provided by the original parent application.

8. A page recognition device, characterized in that, The apparatus includes: a feature acquisition module, configured to acquire the original features of the target page extracted from the page code of the target page; a feature dimensionality reduction module, configured to perform a hashing operation on each eigenvalue in the original features of the target page by using at least one hashing function to obtain the hash value of each eigenvalue in each hashing operation. When the at least one hashing function is multiple hashing functions, use the eigenvalue corresponding to the minimum hash value obtained by performing the hashing operation with each hashing function as the page features after dimensionality reduction. When the at least one hashing function is a single hashing function, select a preset number of eigenvalues in ascending order of the hash values as the page features after dimensionality reduction; an index query module, configured to locate the position corresponding to the page features in a preset bitmap, where each position in the bitmap uniquely records a page feature; perform an inverted index query based on the pre-constructed index with the located position as the index item to obtain a page identifier chain corresponding to the located position, and read the page identifiers in each page identifier block in the page identifier chain to obtain each suspected similar page identifier. Among them, each page identifier chain has multiple page identifier blocks, and each page identifier block has multiple page identifiers. The index includes the mapping relationship between the position in the bitmap and the page identifier. The page identifier having a mapping relationship with the position refers to the page identifier corresponding to the page code with the page features recorded at the position; a matching module, configured to acquire the original page features corresponding to each of the suspected similar page identifiers; respectively match each of the original page features with the original features of the target page; and determine the page corresponding to the original page feature that passes the matching as the page similar to the target page.

9. The device according to claim 8, characterized in that, The apparatus is further configured to: perform feature extraction on the page codes in the page code library; Perform feature dimensionality reduction on the features of the extracted page codes; Record the dimensionality-reduced page features of each page at the corresponding positions in the bitmap; For the positions in the bitmap, determine the page codes having the page features recorded at the positions, and establish a mapping relationship between the positions and the page identifiers corresponding to the determined page codes to generate an index.

10. The device according to claim 9, characterized in that, The apparatus is further configured to: Chunk the page identifiers corresponding to the determined page codes to obtain a plurality of page identifier chunks; Connect the plurality of page identifier chunks to generate a page identifier chain; Establish a mapping relationship between the positions and the page identifier chain to generate an index.

11. The device according to claim 10, characterized in that, The apparatus is further configured to: Starting from the first page identifier chunk in the page identifier chain, read the page identifiers in the page identifier chunk, and for the next page identifier chunk on the page identifier chain, iteratively execute the step of reading the page identifiers in the page identifier chunk until the iteration stops after reading the page identifiers in the last page identifier chunk on the page identifier chain; Wherein, the page identifiers in each of the page identifier chunks on the page identifier chain are the page identifiers corresponding to the page codes having the page features recorded at the located positions.

12. The device according to claim 9, characterized in that, The apparatus is further configured to: When the page codes in the page code library are updated, then Perform feature extraction and feature dimensionality reduction on the updated page codes; Update the page features after dimensionality reduction of the updated page codes to the corresponding positions in the bitmap; In the index, establish a mapping relationship between the updated positions and the page identifiers corresponding to the updated page codes.

13. The device according to claim 8, characterized in that, The original features of the target page have feature values of multiple feature dimensions; the feature dimensionality reduction module is further configured to: Convert the feature dimensions and the ranges of the feature values of the feature values into a preset dimension range and a preset feature value range respectively to obtain the page features after dimensionality reduction.

14. The device according to any one of claims 8 to 13, characterized in that, The target page is a sub-application page; the page identifier is the identifier of the sub-application page; the sub-application page is a page provided by the sub-application; the sub-application is a lightweight application running in the environment provided by the original parent application.

15. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

16. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Counterfeit website detection method, device and system

    CN111224923A